Theme · 2 articles

Agents

Across all formats — watch, learn, research, benchmarks and demos.

AgentsLLMRAGEvaluationMLOpsClassical MLDataToolingSafetyPolicy
Benchmarks · Agents

Eleven open-weight models, one tool-calling harness, 2,200 tasks

The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.

V. Levy dit Vehel, A. Martin · 2 Sept · 3 min
Watch · Agents

MCP one year on: what actually got adopted

The protocol won; most of the servers did not. A look at which patterns survived contact with production.

K. Nguyen · 22 Aug · 1 min