Across all formats — watch, learn, research, benchmarks and demos.
The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.
LLM-as-a-judge is everywhere. We asked whether a judge is systematically kinder to models from its own vendor — and got an answer we did not like.
On three client corpora, swapping the embedder moved recall@10 by 2 points; swapping the reranker moved answer accuracy by 11.