We build, measure and explain applied AI — and publish what we learn.
Field notes, teaching material, internal research, benchmarks and live demos from the Magellan Partners AI team. Written by the people who ran the experiments.
The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.
Tool calling, planning loops, memory and failure modes — implemented in 200 lines, then broken on purpose.
Chunking, embeddings, hybrid search, reranking, evaluation, and the part nobody talks about: updating the index.
Variance, confidence intervals and why your 2-point improvement might be noise.