<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Principia — Applied AI notes from Magellan Partners</title><description>Field notes, teaching material, research, reproducible benchmarks and live demos from the Magellan Partners AI team.</description><link>https://principia.vldv.ovh/</link><language>en</language><item><title>Eleven open-weight models, one tool-calling harness, 2,200 tasks</title><link>https://principia.vldv.ovh/benchmarks/toolcalling-open-weight-2026-09/</link><guid isPermaLink="true">https://principia.vldv.ovh/benchmarks/toolcalling-open-weight-2026-09/</guid><description>The ranking is not the one you expect. Smaller models with a strict JSON schema beat larger ones with free-form tool calls — and retries, not raw capability, explain most of the gap.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate><category>Benchmarks</category><category>Agents</category><category>Evaluation</category></item><item><title>The AI Act&apos;s GPAI obligations, read by people who have to comply with them</title><link>https://principia.vldv.ovh/watch/ai-act-gpai-obligations/</link><guid isPermaLink="true">https://principia.vldv.ovh/watch/ai-act-gpai-obligations/</guid><description>The general-purpose AI provisions entered into application in August 2025. A year on, here is what actually changed in how we ship models for clients — and what did not.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate><category>Watch</category><category>Policy</category><category>Safety</category></item><item><title>Speculative decoding is finally boring — and that&apos;s good news</title><link>https://principia.vldv.ovh/watch/speculative-decoding-boring/</link><guid isPermaLink="true">https://principia.vldv.ovh/watch/speculative-decoding-boring/</guid><description>Every serving engine now ships it, defaults are sane, and the 2× is real on the workloads that matter. Here is what to switch on.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><category>Watch</category><category>LLM</category><category>MLOps</category></item><item><title>Chunking, explained with a ruler and a pair of scissors</title><link>https://principia.vldv.ovh/learn/chunking-ruler-scissors/</link><guid isPermaLink="true">https://principia.vldv.ovh/learn/chunking-ruler-scissors/</guid><description>Before embeddings, before vector databases, there is a much dumber question: where do you cut the document? It turns out the answer decides most of your retrieval quality.</description><pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate><category>Learn</category><category>RAG</category></item><item><title>Can a judge model grade its own family? Preliminary results</title><link>https://principia.vldv.ovh/research/judge-model-own-family/</link><guid isPermaLink="true">https://principia.vldv.ovh/research/judge-model-own-family/</guid><description>LLM-as-a-judge is everywhere. We asked whether a judge is systematically kinder to models from its own vendor — and got an answer we did not like.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><category>Research</category><category>Evaluation</category><category>LLM</category></item><item><title>MCP one year on: what actually got adopted</title><link>https://principia.vldv.ovh/watch/mcp-one-year-on/</link><guid isPermaLink="true">https://principia.vldv.ovh/watch/mcp-one-year-on/</guid><description>The protocol won; most of the servers did not. A look at which patterns survived contact with production.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate><category>Watch</category><category>Agents</category><category>Tooling</category></item><item><title>New: the chunking playground now supports scanned PDFs</title><link>https://principia.vldv.ovh/demos/chunking-playground-scanned-pdfs/</link><guid isPermaLink="true">https://principia.vldv.ovh/demos/chunking-playground-scanned-pdfs/</guid><description>OCR runs before chunking, so you can compare strategies on the documents you actually have — the ugly ones.</description><pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate><category>Demos</category><category>RAG</category><category>Tooling</category></item><item><title>Retrieval is not the bottleneck anymore. Reranking is.</title><link>https://principia.vldv.ovh/research/reranking-is-the-bottleneck/</link><guid isPermaLink="true">https://principia.vldv.ovh/research/reranking-is-the-bottleneck/</guid><description>On three client corpora, swapping the embedder moved recall@10 by 2 points; swapping the reranker moved answer accuracy by 11.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate><category>Research</category><category>RAG</category><category>Evaluation</category></item></channel></rss>