Docs

Archer is an OpenAI-compatible endpoint. Point any OpenAI client at https://api.project-archer.online/v1, use your arch_sk_ key, and every request is routed to the model best suited for it. The model field is accepted but ignored — every response comes back as archer-auto.

Quickstart

  1. Register for an account.
  2. Create an API key on the API Keys page — copy it once; it starts with arch_sk_.
  3. Make your first call with the examples below.

curl

shell
curl https://api.project-archer.online/v1/chat/completions \
  -H "Authorization: Bearer arch_sk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "archer-auto",
    "messages": [{"role": "user", "content": "Write a binary search in Rust"}]
  }'

Python (OpenAI SDK)

python
from openai import OpenAI

client = OpenAI(
    api_key="arch_sk_...",
    base_url="https://api.project-archer.online/v1",
)

resp = client.chat.completions.create(
    model="archer-auto",  # accepted but ignored — Archer routes for you
    messages=[{"role": "user", "content": "Write a binary search in Rust"}],
)
print(resp.choices[0].message.content)

JavaScript / TypeScript (OpenAI SDK)

javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "arch_sk_...",
  baseURL: "https://api.project-archer.online/v1",
});

const resp = await client.chat.completions.create({
  model: "archer-auto", // accepted but ignored — Archer routes for you
  messages: [{ role: "user", content: "Write a binary search in Rust" }],
});
console.log(resp.choices[0].message.content);

Streaming

Pass stream: true to receive OpenAI-format server-sent events, terminated by data: [DONE].

python
stream = client.chat.completions.create(
    model="archer-auto",
    messages=[{"role": "user", "content": "Explain quicksort"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Rate limits & quotas

Every /v1/* response carries rate-limit headers:

  • X-RateLimit-Limit-Requests — per-minute request limit
  • X-RateLimit-Remaining-Requests — requests left in the current window
  • X-RateLimit-Reset-Requests — seconds until the window resets

Exceed a limit and you get a 429 with a Retry-After header (seconds) and this body:

json
{
  "error": {
    "message": "Rate limit exceeded. Try again in 12s.",
    "type": "rate_limit_error",
    "code": "rate_limit_exceeded"
  }
}

Model catalog

You never pick from these — Archer draws the right one per request. Listed for reference only; the API always presents a single archer-auto model.

ModelSpeedContext
GPT-OSS 120Bfast131K
Qwen3.6 27Bfast131K
GPT-OSS 20Bvery fast131K
Nemotron 3 Nano 30Bfast131K
Nemotron 3 Ultramedium262K
Nemotron 3 Supermedium262K
MiniMax M3medium524K
Gemma 4 31Bslow131K