TL;DR

Stop A/B-testing embedding models. Put the same energy into the reranker and the prompt that consumes its output.

Setup

Three anonymised client corpora (legal, industrial maintenance, HR policy), 90 hand-written questions each, five embedders, four rerankers.

Results

Embedder choice: recall@10 spread of 2.1 points. Reranker choice: end-to-end answer accuracy spread of 11.4 points. The interaction term is small.

Next

Cost. The best reranker is also the slowest; we are measuring where the accuracy/latency curve bends.