Request access
Menu
← All posts · Product

Introducing EVO Router

Frontier model performance at a fraction of the cost, on the workloads you already run.

Most of what you send to a frontier model does not need one. Almost nobody checks, because checking means building evals for every workload and testing every alternative against them, and getting it wrong means shipping worse output to users.

So the model gets picked once, at the start, and never revisited. New models ship every month. The prompt gets tuned, the scaffold gets rewritten, and the model stays whatever it was on day one. You keep paying frontier prices for a model that stopped being frontier months ago. Newer ones ship every few weeks that do the job better and cost less.

EVO Router does that checking. It is a drop-in endpoint: point your gateway at it and it starts by serving the model you already use, returning the output you already get. From there it learns each workload from live traffic, searches offline for a more efficient way to serve it, and moves traffic only once that route clears your quality bar on requests it has not seen.

The quality bar is set from your own workloads, not a leaderboard. EVO builds evals out of your production outputs, so clearing the bar means matching the quality your traffic actually got.

what evo optimizes

Routing is one lever. EVO searches six.

  • Model and provider. The right model for each kind of request, re-checked as new ones ship.
  • Prompts. The same prompt does not behave the same on every model, so EVO rewrites it to fit.
  • Fusion and voting. Several small models answering together, with a tiebreak.
  • Cascades. A cheap first pass, escalating only when the answer needs it.
  • Harness and tools. Tool definitions, planning loops and retries, tuned per model, since a scaffold built for one model rarely suits another.
  • Context and caching. What goes in the window, in what order, and what gets reused.

Every candidate is scored offline against your baseline before anything ships, so nothing moves on a vendor claim or a public benchmark number.

frontier model performance at a fraction of the cost

To show what that search produces, we pointed it at a public benchmark. GPQA Diamond is 198 PhD-level science questions, and the models at the top of it cost five to twelve cents each. EVO ran unattended overnight on $47 of API spend.

GPQA Diamond: accuracy against cost per question Log-scale cost on the horizontal axis against GPQA Diamond accuracy. EVO Router scores 93.2 percent for under one cent a question, left of every model at comparable accuracy. 86% 88% 90% 92% 94% $.003 $.01 $.03 $.1 $.3 cost per question (log scale) GPQA Diamond accuracy cost-quality frontier Gemini 3.1 Pro Preview: 94.1% at $0.0715 Claude Opus 5 (Adaptive Reasoning, High Effort): 93.7% at $0.0452 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort): 93.7% at $0.0635 Kimi K3 (max): 93.5% at $0.1512 GPT-5.5 (xhigh): 93.5% at $0.1540 Claude Opus 5 (Adaptive Reasoning, Max Effort): 93.2% at $0.0921 GPT-5.5 (high): 93.2% at $0.0898 Grok 4.5 (high): 93.1% at $0.0316 MiniMax-M3: 92.9% at $0.0099 Gemini 3.6 Flash (high): 92.8% at $0.0300 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback): 92.6% at $0.2504 GPT-5.5 (medium): 92.6% at $0.0524 GPT-5.6 Terra (max): 92.5% at $0.0730 Qwen3.7 Max: 92.3% at $0.0819 Gemini 3.5 Flash (high): 92.2% at $0.0503 Gemini 3.5 Flash (medium): 92.1% at $0.0352 GPT-5.4 (xhigh): 92.0% at $0.0917 Claude Opus 4.8 (Adaptive Reasoning, Max Effort): 92.0% at $0.3123 Claude Opus 5 (Adaptive Reasoning, Medium Effort): 91.9% at $0.0241 GPT-5.3 Codex (xhigh): 91.5% at $0.0849 Claude Opus 4.7 (Adaptive Reasoning, Max Effort): 91.4% at $0.3281 GPT-5.6 Luna (max): 91.1% at $0.0099 Claude Sonnet 5 (Adaptive Reasoning, Max Effort): 91.1% at $0.3663 Grok 4.20 0309 v2 (Reasoning): 91.1% at $0.0405 Kimi K2.6: 91.1% at $0.0858 GPT-5.5 (low): 91.0% at $0.0214 GPT-5.6 Terra (xhigh): 90.8% at $0.0257 DeepSeek V4 Flash 0731 (Reasoning, Max Effort): 90.8% at $0.0027 Gemini 3 Pro Preview (high): 90.8% at $0.0936 DeepSeek V4 Pro (Reasoning, High Effort): 90.5% at $0.0092 GPT-5.2 (xhigh): 90.3% at $0.1112 Grok 4.3 (high): 90.1% at $0.0238 Qwen3.7 Plus: 90.0% at $0.0208 GPT-5.2 Codex (xhigh): 89.9% at $0.3435 Muse Spark 1.1 (xhigh): 89.8% at $0.0316 Gemini 3 Flash Preview (Reasoning): 89.8% at $0.0433 Hy3: 89.7% at $0.0127 GPT-5.6 Terra (high): 89.6% at $0.0184 Kimi K2.7 Code: 89.6% at $0.0396 Claude Opus 4.6 (Adaptive Reasoning, Max Effort): 89.6% at $0.5184 GPT-5.6 Luna (xhigh): 89.5% at $0.0059 Inkling Small: 89.5% at $0.0212 GLM-5.2 (max): 89.5% at $0.0876 Grok Build 0.1 0616: 89.5% at $0.0230 DeepSeek V4 Flash (Reasoning, Max Effort): 89.4% at $0.0083 Qwen3.5 397B A17B (Reasoning): 89.3% at $0.0252 GPT-5.6 Luna (high): 89.2% at $0.0034 Nex-N2-Pro: 89.2% at $0.0141 Grok 4.3 (medium): 89.0% at $0.0087 Claude Opus 5 (Adaptive Reasoning, Low Effort): 88.9% at $0.0076 DeepSeek V4 Pro (Reasoning, Max Effort): 88.8% at $0.0252 Qwen3.6 Max Preview: 88.8% at $0.0902 Gemini 3 Pro Preview (low): 88.7% at $0.0217 Claude Opus 4.7 (Non-reasoning, High Effort): 88.5% at $0.0273 Grok 4.20 0309 (Reasoning): 88.5% at $0.0351 Qwen3.6 Plus: 88.2% at $0.0306 Kimi K2.5 (Reasoning): 87.9% at $0.0445 Grok 4: 87.7% at $0.1383 Agnes 2.5 Pro Alpha: 87.6% at $0.0039 GPT-5.4 mini (xhigh): 87.5% at $0.0771 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort): 87.5% at $0.4138 MiniMax-M2.7: 87.4% at $0.0130 GPT-5.1 (high): 87.3% at $0.0745 GPT-5.6 Terra (medium): 87.2% at $0.0093 Inkling (xhigh): 87.2% at $0.0885 GPT-5.4 (low): 87.1% at $0.0150 GLM-5.1 (Reasoning): 86.8% at $0.0783 Nemotron 3 Ultra 550B A55B (Reasoning): 86.7% at $0.0396 Hy3-preview (Reasoning): 86.7% at $0.0050 DeepSeek V4 Flash (Reasoning, High Effort): 86.7% at $0.0030 MiMo-V2.5-Pro: 86.6% at $0.0120 Claude Opus 4.5 (Reasoning): 86.6% at $0.3443 GPT-5.2 (medium): 86.4% at $0.0242 Qwen3.5 397B A17B (Non-reasoning): 86.1% at $0.0095 Gemini 3.1 Pro Preview: 94.1% at $0.0715 Gemini 3.1 Pro Claude Opus 5 (Adaptive Reasoning, High Effort): 93.7% at $0.0452 Claude Opus 5 MiniMax-M3: 92.9% at $0.0099 MiniMax-M3 Grok 4.5 (high): 93.1% at $0.0316 Grok 4.5 GPT-5.6 Sol (low): 89.8% at $0.0150 GPT-5.6 Sol (medium): 92.6% at $0.0233 GPT-5.6 Sol (high): 92.8% at $0.0362 GPT-5.6 Sol (xhigh): 93.1% at $0.0552 GPT-5.6 Sol (max): 94.1% at $0.1094 GPT-5.6 Sol at max effort low effort Alternate routing layer: 89.56% at $0.00852 alternate router 89.6% · $.0085 EVO Router: 93.18% at $0.00825 (average of 2 trials) EVO 93.2% · $.0083 EVO (our run) GPT-5.6 Sol, by effort Alternate router Other models (Artificial Analysis)

Accuracy against cost per question, log scale. Field and effort-ladder data: Artificial Analysis, retrieved 3 August 2026, costed from their published token counts at list prices. EVO's point is our own run at OpenRouter billed rates, overlaid on their chart. The router point is that vendor's own published figure under their own harness. In the scorecard below, GPT-5.6 Sol on GPQA Diamond is Artificial Analysis's run at maximum effort; on IFEval both the router and Sol figures are that vendor's published numbers, since no independent IFEval measurement was available. EVO's bars are our own runs throughout.

The staircase is the cost-quality frontier, the cheapest way anyone has found to reach each level of accuracy. EVO sits on it. Of the 61 models in the competitive band, 49 cost more and score no better, and the cheapest one that scores higher costs 5× as much.

The dark line is a single frontier model across its reasoning-effort settings. EVO beats every setting but the most expensive, at a tenth of the price.

Solve rate and cost per task, by benchmark On GPQA Diamond and IFEval, EVO Router matches or beats both a alternate routing layer and the GPT-5.6 Sol reference model on solve rate, at between four and twelve times lower cost per task. EVO Router Alternate router GPT-5.6 Sol 0% 25% 50% 75% 100% solve rate (higher is better) $0 $0.025 $0.05 $0.075 $0.1 cost per task (lower is better) EVO Router on GPQA Diamond: 93.18% 93.2% EVO Router on GPQA Diamond: $0.00825 per task $0.0083 Alternate router on GPQA Diamond: 89.56% 89.6% Alternate router on GPQA Diamond: $0.00852 per task $0.0085 GPT-5.6 Sol on GPQA Diamond: 94.14% 94.1% GPT-5.6 Sol on GPQA Diamond: $0.10941 per task $0.1094 GPQA Diamond 198 questions EVO Router on IFEval: 97.91% 97.9% EVO Router on IFEval: $0.00185 per task $0.0019 Alternate router on IFEval: 96.73% 96.7% Alternate router on IFEval: $0.00730 per task $0.0073 GPT-5.6 Sol on IFEval: 96.23% 96.2% GPT-5.6 Sol on IFEval: $0.01469 per task $0.0147 IFEval 541 prompts

Solve rate above, cost per task below, on a shared axis. GPQA Diamond and IFEval are the two benchmarks EVO has been run against end to end. Sources differ by bar and are listed under the frontier chart above.

On GPQA Diamond EVO beats the alternate router on both axes, and the reference model scores a point higher while costing 13 times as much. On IFEval EVO scores above both and costs 8 times less than the reference. That is the shape the search keeps producing: parity or better on the axis you care about, and a different order of magnitude on the one you pay for.

93.18% on GPQA Diamond at $0.00825 a question. The nearest model that scores higher costs 5× more.

the system it found

Two commodity models. No frontier model is called anywhere in the pipeline.

The two-tier cascade A question goes to three parallel samples of one cheap open-weight model. If all three agree, about 91 percent of the time, that is the answer. If they disagree, a single call to a second open-weight model from a different lab overrides it. question open-weight model 13 parallel samples · t=0.7 answer model 21 call · overrides answer all three agree · ~91% they disagree · ~9%

A cheap open-weight model answers three times. If all three agree, which happens on about 91% of questions, that is the answer, and the question cost about two tenths of a cent. If they disagree, one call to a second model settles it. The second model was not chosen for being strong. It was chosen because it comes from a different lab, so its errors are decorrelated from the first model's.

That is the kind of solution EVO's search tends to find: an optimized path for most of the traffic, with spend concentrated on the few requests where that path is uncertain. Here that concentration is steep. The 9% of questions that escalate consume 74% of the total bill. That is what the search arrived at for this workload, with the models available this month. Next month there are new models and the answer changes, which is the reason to have something that keeps checking.

Reported GPQA figures are the average of two scored trials, 91.9% and 94.4%; a single 198-question run carries roughly plus or minus 3.4 points at 95% confidence. This is a cross-harness comparison on the same 198 questions: Artificial Analysis ran theirs under their protocol at maximum-effort settings and list pricing, and ours is our own run overlaid on their chart rather than measured by them. Prices are a snapshot of 3 August 2026.

get started

EVO Router is in private beta. Request access and we will send you an API key and a base URL. It sits next to the gateway you already run, as a provider or a proxy in front of it, and it does not replace it.

Request access or read the docs first

Once you have a key, integration is a base URL and a model prefix:

Python · OpenAI SDK two lines change
from openai import OpenAI

client = OpenAI(    base_url="https://api.evo-hq.com/v1",    api_key=os.environ["EVO_API_KEY"],
)

client.chat.completions.create(    model="evo-auto/gpt-5.6",    messages=[{"role": "user", "content": "..."}],
)

Your SDK, messages and tool definitions stay as they are. The model you name after evo-auto/ is whatever you run today, and it stays the fallback, so on day one every request passes straight through to it. Workloads start moving only once EVO has evidence they should, and you can pin any workload back to the base model at any time.