Skip to main content
1 min readKnowledge Resource

Knowledge Resource

Research Summary: Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Summary & Analysis prepared by
Aziz Shuaib Ausi
Resource type
Research Summary / Knowledge Resource
Resource published on AZIZ OS
17 September 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
About this Summary & Analysis

AZIZ OS provides independently prepared summaries and analytical interpretations of externally published research and knowledge sources. The underlying works remain attributable to their original authors and rights holders. This resource is intended to improve accessibility and understanding and does not replace the original publication.

Checking access…

Research into LLM-as-judge panels indicates that the identity of the judge LLM significantly influences evaluation outcomes, even when controlling for candidate quality. The study found a measurable 'same-family lift,' where LLMs tend to favor candidates from their own developmental family, suggesting a potential bias in automated evaluation systems.

Why it matters

This research reveals an inherent bias in LLM-as-judge systems, where evaluators may favor candidates from their own developmental family. Organizations relying on or developing AI for automated evaluation, content generation, or decision-making processes must understand and mitigate such biases to ensure fairness, objectivity, and reliability. Failure to address this could lead to skewed results, suboptimal decisions, and potentially erode trust in AI systems.

Key insights

  • The identity of an LLM judge significantly affects the judgment outcome, making it challenging to isolate this effect from the quality of the candidate being judged.
  • A common per-family statistic for evaluating LLM judges is highly confounded with candidate quality and strongly correlates with Bradley-Terry ability (r = 0.95).
  • A corrected estimator was derived to specifically compare judges while holding candidate family constant.
  • Using this corrected estimator, all four studied open-weight LLM families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) exhibited a positive 'same-family lift' ranging from 3.4 to 8.4 percentage points.
  • The global 'same-family lift' was statistically significant at 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002).
  • This effect persists even when applying panel-based quality controls and an independent human-consensus anchor.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2609.17857

Citation

Cite the original work (APA 7)

The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.

Verification

This is an authenticated AZIZ OS resource record.

Verification ID
ASA-EXE-2026-00678
Version
v1.0 · r0
Issued
17 September 2026
Resource prepared by
Aziz Shuaib Ausi
Resource status
Research Summary / Knowledge Resource
Underlying work
Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels
Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Provenance status
Attribution requires verification
Rights
Underlying publication rights remain with the respective copyright holder(s). Refer to the original source for authoritative publication and licensing information.

This verification confirms the AZIZ OS resource record and its documented provenance. It does not establish authorship of the underlying external work.

Verify this resource