Section · 2 articles

Research

Internal R&D: questions, results, and what we still do not know.

LLMRAGEvaluation
Research · Evaluation

Can a judge model grade its own family? Preliminary results

LLM-as-a-judge is everywhere. We asked whether a judge is systematically kinder to models from its own vendor — and got an answer we did not like.

V. Levy dit Vehel · 25 Aug · 1 min
Research · RAG

Retrieval is not the bottleneck anymore. Reranking is.

On three client corpora, swapping the embedder moved recall@10 by 2 points; swapping the reranker moved answer accuracy by 11.

K. Nguyen · 18 Aug · 1 min