1 min readExecutive Guide

Executive Guide

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

Research indicates that Large Language Models (LLMs) demonstrate uncalibrated reliance on proxy attributes when making decisions, specifically failing to align their reliance with the actual predictive evidence. This issue is identified in a clinical-ranking task, raising concerns about LLMs' deployment in sensitive areas such as triage and lending where distinguishing between legitimate inference and impermissible proxy use is critical.

Checking access…

Research indicates that Large Language Models (LLMs) demonstrate uncalibrated reliance on proxy attributes when making decisions, specifically failing to align their reliance with the actual predictive evidence. This issue is identified in a clinical-ranking task, raising concerns about LLMs' deployment in sensitive areas such as triage and lending where distinguishing between legitimate inference and impermissible proxy use is critical.

Why it matters

The uncalibrated reliance of Large Language Models on proxy attributes poses significant risks for deploying AI in critical decision-making processes. This highlights a fundamental challenge in ensuring fairness and accuracy, particularly where decisions impact individuals or allocate resources. Addressing this issue is crucial for maintaining trust and avoiding unintended biases in AI-driven systems across various sectors.

Key insights

  • LLMs are increasingly used in decision-making contexts like triage and lending, necessitating a clear distinction between task-relevant inference and impermissible proxy use.
  • Traditional auditing methods, which assess decision changes based on demographic shifts, are insufficient because attributes correlated with protected groups can possess predictive value.
  • This study measures causal proxy effects in four LLMs using a clinical-ranking task with known ground truth, allowing for exact computation of warranted reliance.
  • The audit generated three verdicts: over-reliance, warranted reliance, and under-reliance on proxies.
  • All tested LLMs relied on proxies with no information when presented with neutral labels.
  • For informative proxies, the models exhibited all three reliance verdicts, suggesting inconsistent and uncalibrated behavior.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.22887

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). Proxy reliance in large language model decisions is uncalibrated to predictive evidence. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00604

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00604
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication