Section · 1 articles

Benchmarks

Reproducible measurements with frozen, versioned protocols.

AgentsEvaluation
How we benchmark
Every benchmark has a frozen, versioned protocol (models, hardware, dataset, metrics, decoding, repo commit). Changing anything creates a new version and a new article — old results stay online. Raw data ships with the repo.
Benchmarks · Agents

Eleven open-weight models, one tool-calling harness, 2,200 tasks

The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.

V. Levy dit Vehel, A. Martin · 2 Sept · 3 min