The evo blog
Notes on autoresearch, harness optimization, and putting coding agents to work on things you can measure.
-
Hillclimbing a smart router on agent workloads
Beats every single model on ITSMBench (Vibrant Labs) at 21× lower cost than the nearest one — 71.4% at $0.0224 a task.
-
Introducing EVO Router
A drop-in endpoint that serves the traffic you already have at frontier quality, for a fraction of the cost. On GPQA Diamond it reaches 93.43% at $0.00896 a question, against $0.045 for the nearest model that scores higher.
-
The meta-loop that rewrites the optimization loop, live
How evo 0.5 runs the optimizer and a meta-controller as two threads on one event loop, reshaping the search while experiments are still running.
-
Autoresearch your way into improving your models + harness
evo can now optimize both axes — model weights and the harness — in one loop. On a benchmark a recent launch hit 0.701 with RL on a 120B model, evo reached 0.7766 — and the winning solution used no LLM at all.
-
Using Autoresearch to improve harness skillsAgents are now improving themselves. You point one at a problem, leave it alone, and come back to a version that scores higher — the model never changed, the wrapper did. A detailed SealQA run: 5/20 → 11/20.