Docs
Archer is an OpenAI-compatible endpoint. Point any OpenAI client at https://api.project-archer.online/v1, use your arch_sk_ key, and every request is routed to the model best suited for it. The model field is accepted but ignored — every response comes back as archer-auto.
Quickstart
curl
shell
curl https://api.project-archer.online/v1/chat/completions \
-H "Authorization: Bearer arch_sk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "archer-auto",
"messages": [{"role": "user", "content": "Write a binary search in Rust"}]
}'Python (OpenAI SDK)
python
from openai import OpenAI
client = OpenAI(
api_key="arch_sk_...",
base_url="https://api.project-archer.online/v1",
)
resp = client.chat.completions.create(
model="archer-auto", # accepted but ignored — Archer routes for you
messages=[{"role": "user", "content": "Write a binary search in Rust"}],
)
print(resp.choices[0].message.content)JavaScript / TypeScript (OpenAI SDK)
javascript
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "arch_sk_...",
baseURL: "https://api.project-archer.online/v1",
});
const resp = await client.chat.completions.create({
model: "archer-auto", // accepted but ignored — Archer routes for you
messages: [{ role: "user", content: "Write a binary search in Rust" }],
});
console.log(resp.choices[0].message.content);Streaming
Pass stream: true to receive OpenAI-format server-sent events, terminated by data: [DONE].
python
stream = client.chat.completions.create(
model="archer-auto",
messages=[{"role": "user", "content": "Explain quicksort"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Rate limits & quotas
Every /v1/* response carries rate-limit headers:
X-RateLimit-Limit-Requests— per-minute request limitX-RateLimit-Remaining-Requests— requests left in the current windowX-RateLimit-Reset-Requests— seconds until the window resets
Exceed a limit and you get a 429 with a Retry-After header (seconds) and this body:
json
{
"error": {
"message": "Rate limit exceeded. Try again in 12s.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}Model catalog
You never pick from these — Archer draws the right one per request. Listed for reference only; the API always presents a single archer-auto model.
| Model | Speed | Context |
|---|---|---|
| GPT-OSS 120B | fast | 131K |
| Qwen3.6 27B | fast | 131K |
| GPT-OSS 20B | very fast | 131K |
| Nemotron 3 Nano 30B | fast | 131K |
| Nemotron 3 Ultra | medium | 262K |
| Nemotron 3 Super | medium | 262K |
| MiniMax M3 | medium | 524K |
| Gemma 4 31B | slow | 131K |