One OpenAI-compatible endpoint that reads each request and sends it to the model best suited for it. One key, no model-picking — every answer comes back as archer-auto.
This runs Archer's real keyword rules — the same sets and the same priority order as app/core/router.py. In production an embedding pass gets first say, and these rules catch whatever it isn't sure about.
Rate limit, server error or timeout, and Archer walks the chain until one model lands. Client errors stop it dead — a bad request fails the same way on every model.
Order is the catalog's fallback_priority. The faster Groq models hold the head of the chain; the larger, slower Ollama Cloud models sit behind them.
A curated pool sits behind your key. You never pick from it — Archer draws the right one for every request.
Whoever answers, your response always comes back as archer-auto. The client never learns which provider served it.
If you've called OpenAI, you've already written Archer. Point the client at our base URL, use your key, and route.
from openai import OpenAI
client = OpenAI(
api_key="arch_sk_...",
base_url="https://api.project-archer.online/v1",
)
resp = client.chat.completions.create(
model="archer-auto", # accepted, then ignored
messages=[{
"role": "user",
"content": "Write a binary search in Rust",
}],
)
print(resp.choices[0].message.content)Each prompt is read and sent to the model that fits it — code, math, analysis, or a quick reply.
Rate-limited or down? Archer retries down a fixed chain until a model lands the shot.
A drop-in /v1/chat/completions endpoint. Keep your SDK — just swap the base URL.
A single arch_sk_ key stands in for the wallet of provider keys you'd otherwise juggle.
See which model answered, why it routed there, plus tokens and latency, on your dashboard.
Whoever answers, the response always comes back in the same archer-auto shape.
Only the base URL and the key. Archer speaks the OpenAI chat-completions protocol, streaming included, so any OpenAI SDK works unchanged. The model field is accepted and then ignored.
An embedding pass compares your last user message to six precomputed domain centroids. Above the similarity threshold it decides; below it — or if the embedding provider is unreachable — the keyword rules above take over. Both engines pick a domain, and the domain-to-model mapping lives in the catalog.
Retryable failures (rate limit, server error, timeout) walk the fallback chain until a model answers. Client errors are never retried — a bad request fails identically everywhere. If every model fails you get a 503.
Yes, standard SSE chunks. Fallback applies until the first byte reaches you; after that the response is committed, so a mid-stream provider failure ends the stream rather than silently swapping models under a client that is already rendering tokens.
Not from the API — every response reports archer-auto by design. The dashboard's Logs page shows the real model, the routing reason, whether it fell back, time-to-first-token, tokens and latency.
30 requests per minute and 10,000 requests per month on the free tier, with X-RateLimit headers on every /v1 response. If the limiter's Redis is unavailable, requests are allowed through rather than rejected.
Sign up, generate a key, and point your first request at Archer. The right model is already waiting on the line.