TL;DR
Stop A/B-testing embedding models. Put the same energy into the reranker and the prompt that consumes its output.
Setup
Three anonymised client corpora (legal, industrial maintenance, HR policy), 90 hand-written questions each, five embedders, four rerankers.
Results
Embedder choice: recall@10 spread of 2.1 points. Reranker choice: end-to-end answer accuracy spread of 11.4 points. The interaction term is small.
Next
Cost. The best reranker is also the slowest; we are measuring where the accuracy/latency curve bends.