Orchestration-as-a-Service

Every prompt
finds its mark.

One OpenAI-compatible endpoint that reads each request and sends it to the model best suited for it. One key, no model-picking — every answer comes back as archer-auto.

8 models · 2 providersPOST /v1/chat/completionsSSE streaming
router · live view
The router, in the open

Type a prompt. Watch it get aimed.

This runs Archer's real keyword rules — the same sets and the same priority order as app/core/router.py. In production an embedding pass gets first say, and these rules catch whatever it isn't sure about.

messages[-1].content
⌘↵ to route
01
Tokenize the last user message
waiting for a prompt
02
Check the rules in priority order
coding → math → writing → simple → analysis → default
03
Resolve the domain to a model
domain → model comes from the DB catalog
04
Normalize and log
every response is stamped archer-auto
response.model
archer-auto
served by
—
Fallback chain

A miss never reaches you.

Rate limit, server error or timeout, and Archer walks the chain until one model lands. Client errors stop it dead — a bad request fails the same way on every model.

01
gpt-oss-120b
Groq
02
qwen3.6-27b
Groq
03
gpt-oss-20b
Groq
04
nemotron-3-nano
Ollama Cloud
05
nemotron-3-ultra
Ollama Cloud
06
nemotron-3-super
Ollama Cloud
07
minimax-m3
Ollama Cloud
08
gemma4-31b
Ollama Cloud

Order is the catalog's fallback_priority. The faster Groq models hold the head of the chain; the larger, slower Ollama Cloud models sit behind them.

The quiver

8 models. One target.

A curated pool sits behind your key. You never pick from it — Archer draws the right one for every request.

GET /v1/models
ModelProviderContextDrawn for
gpt-oss-120bGroq131KCode, math, and the default shot
qwen3.6-27bGroq131KFast replies to short, simple asks
gpt-oss-20bGroq131KQuick conversational turns
nemotron-3-nanoOllama131KLight general-purpose work
nemotron-3-ultraOllama262KAnalysis and step-by-step reasoning
nemotron-3-superOllama262KDeep reasoning, held in reserve
minimax-m3Ollama524KVery long-context analysis
gemma4-31bOllama131KWriting and long-form drafting

Whoever answers, your response always comes back as archer-auto. The client never learns which provider served it.

Drop-in

Same code. New aim.

If you've called OpenAI, you've already written Archer. Point the client at our base URL, use your key, and route.

One endpoint
Swap the base URL. Nothing else in your call changes.
One key
A single arch_sk_ key replaces the wallet of provider keys.
Zero model-picking
The model field is accepted, then ignored. Routing and fallback are automatic.
quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="arch_sk_...",
    base_url="https://api.project-archer.online/v1",
)

resp = client.chat.completions.create(
    model="archer-auto",   # accepted, then ignored
    messages=[{
        "role": "user",
        "content": "Write a binary search in Rust",
    }],
)
print(resp.choices[0].message.content)
Built to never miss

What the layer actually does.

Intent routing

Each prompt is read and sent to the model that fits it — code, math, analysis, or a quick reply.

Automatic fallback

Rate-limited or down? Archer retries down a fixed chain until a model lands the shot.

OpenAI-compatible

A drop-in /v1/chat/completions endpoint. Keep your SDK — just swap the base URL.

One key for all

A single arch_sk_ key stands in for the wallet of provider keys you'd otherwise juggle.

Every shot logged

See which model answered, why it routed there, plus tokens and latency, on your dashboard.

Normalized replies

Whoever answers, the response always comes back in the same archer-auto shape.

FAQ

The questions developers actually ask.

Do I have to change my code?

Only the base URL and the key. Archer speaks the OpenAI chat-completions protocol, streaming included, so any OpenAI SDK works unchanged. The model field is accepted and then ignored.

How does it decide which model answers?

An embedding pass compares your last user message to six precomputed domain centroids. Above the similarity threshold it decides; below it — or if the embedding provider is unreachable — the keyword rules above take over. Both engines pick a domain, and the domain-to-model mapping lives in the catalog.

What happens when a provider fails?

Retryable failures (rate limit, server error, timeout) walk the fallback chain until a model answers. Client errors are never retried — a bad request fails identically everywhere. If every model fails you get a 503.

Does streaming still work?

Yes, standard SSE chunks. Fallback applies until the first byte reaches you; after that the response is committed, so a mid-stream provider failure ends the stream rather than silently swapping models under a client that is already rendering tokens.

Can I see which model actually answered?

Not from the API — every response reports archer-auto by design. The dashboard's Logs page shows the real model, the routing reason, whether it fell back, time-to-first-token, tokens and latency.

What are the limits?

30 requests per minute and 10,000 requests per month on the free tier, with X-RateLimit headers on every /v1 response. If the limiter's Redis is unavailable, requests are allowed through rather than rejected.

Ready to let it fly?

Sign up, generate a key, and point your first request at Archer. The right model is already waiting on the line.