Internal R&D: questions, results, and what we still do not know.
LLM-as-a-judge is everywhere. We asked whether a judge is systematically kinder to models from its own vendor — and got an answer we did not like.
On three client corpora, swapping the embedder moved recall@10 by 2 points; swapping the reranker moved answer accuracy by 11.