Research
Notes from building Agent Foundation Models.
Method, benchmarks, and the engineering behind small models that act. Written by the team building them, evidence over adjectives. Figures that aren’t measured yet are marked as such.
More is on the way, including the Execution-Reliability Benchmark methodology and results. Want a specific topic covered? Tell us.
Evaluation sandbox
Own your intelligence, don’t rent it.
The evaluation sandbox converts one high-volume workflow to a model you own, with cost and reliability measured against your current stack.