Intelligence

ai

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Source
OpenAI Research
Published
Last verified
6 Aug 2026
Confidence
High
Evidence
Original document retained
Reading time
1 min
Country
United States
Relevant to
Operations & Delivery, Technology & Data

Executive summary

What happened, and why should leadership care?

OpenAI Research indicates that the implementation of two specific API settings significantly enhanced GPT-5.6's performance on the ARC-AGI-3 benchmark. These settings, which include retaining reasoning and enabling compaction, led to a tripling of scores and improved efficiency. This development highlights advancements in AI model optimization and performance.

Why this matters

Why is this strategically important?

This development is strategically important as it demonstrates a clear methodology for enhancing AI model performance and efficiency through specific configuration adjustments. It indicates that significant gains can be achieved not just through fundamental model architecture changes but also through optimized operational parameters, which can lead to more robust and capable AI systems across various applications.

Key insights

What should be noted from the evidence?

  • Two API settings were instrumental in improving GPT-5.6 performance.
  • Scores on the ARC-AGI-3 benchmark tripled due to these changes.
  • The settings involved retaining reasoning capabilities.
  • Compaction was another key feature enabled.
  • Efficiency of GPT-5.6 was also boosted by these settings.

Evidence and confidence

How far can this assessment be trusted?

High confidence. Named institution, original document retained and analysis corroborated.

Analysis is prepared by the AZIZ OS Intelligence Engine. The original publication remains the authoritative record, and executive judgement remains entirely human.

Source

Where does this originate?

Reported by OpenAI Research · United States. This briefing summarises the publication for executive use; the document itself is not reproduced here.

Read the original publication