TL;DR

On 600 paired answers, three judges each rated their own family 4–7 points higher than the two others did. The effect survives prompt shuffling. We do not yet know whether it is style preference or something worse.

Question

Does a judge model favour answers produced by models trained by the same organisation?

Setup

Three families, one judge per family, 600 question-answer pairs, blinded and shuffled. Each judge grades all 600 answers on a 1–10 rubric.

Results so far

Every judge grades its own family higher. The self-preference is 4–7 points on a 100-point normalised scale and does not disappear when we strip stylistic markers.

What we do not know

Whether this is stylistic (judges like what they would have written) or semantic (judges accept claims they would have made). The next run separates the two with paraphrased answers.