Theme · 3 articles

Evaluation

Across all formats — watch, learn, research, benchmarks and demos.

AgentsLLMRAGEvaluationMLOpsClassical MLDataToolingSafetyPolicy
Benchmarks · Agents

Eleven open-weight models, one tool-calling harness, 2,200 tasks

The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.

V. Levy dit Vehel, A. Martin · 2 Sept · 3 min
Research · Evaluation

Can a judge model grade its own family? Preliminary results

LLM-as-a-judge is everywhere. We asked whether a judge is systematically kinder to models from its own vendor — and got an answer we did not like.

V. Levy dit Vehel · 25 Aug · 1 min
Research · RAG

Retrieval is not the bottleneck anymore. Reranking is.

On three client corpora, swapping the embedder moved recall@10 by 2 points; swapping the reranker moved answer accuracy by 11.

K. Nguyen · 18 Aug · 1 min