Every benchmark has a frozen, versioned protocol (models, hardware, dataset, metrics, decoding, repo commit). Changing anything creates a new version and a new article — old results stay online. Raw data ships with the repo.
The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.