On 600 paired answers, three judges each rated their own family 4–7 points higher than the two others did. The effect survives prompt shuffling. We do not yet know whether it is style preference or something worse.
Question
Does a judge model favour answers produced by models trained by the same organisation?
Setup
Three families, one judge per family, 600 question-answer pairs, blinded and shuffled. Each judge grades all 600 answers on a 1–10 rubric.
Results so far
Every judge grades its own family higher. The self-preference is 4–7 points on a 100-point normalised scale and does not disappear when we strip stylistic markers.
What we do not know
Whether this is stylistic (judges like what they would have written) or semantic (judges accept claims they would have made). The next run separates the two with paraphrased answers.