tech
How GPT-5.6 Fuses Frontier Intelligence with Frontier Efficiency
We designed the GPT‑5.6 model family to balance capability and cost across the spectrum of tasks people use our models for. Our flagship model, GPT‑5.6 Sol, with max reasoning outperforms Claude Fable 5 on the Artificial Analysis Coding Agent Index at less than half of the cost. Terra performs as well as GPT‑5.5 on intelligence benchmarks at half the price, and Luna is our fastest and most affordable model, priced 80% less than the cost of Sol. To deliver these efficiencies, our research and technical teams have made significant optimizations at every major layer of our stack. These improvements span our models, inference (how we run models to generate output), and our agentic harness, which is used by both Codex and ChatGPT Work.

TL;DR
- The GPT-5.6 model family is designed for a balance of capability and cost across various tasks.
- GPT-5.6 Sol offers superior reasoning performance compared to Claude Fable 5 at less than half the cost.
- Terra matches GPT-5.5 performance on intelligence benchmarks at half the price, while Luna is the fastest and most affordable.
- Efficiency improvements have been made across models, inference processes, and the agentic harness used by Codex and ChatGPT Work.
- Optimizations in inference include load balancing, speculative decoding, caching, and kernel optimization to maximize output from hardware.
- The agentic harness streamlines repeated work by avoiding context bloat, preserving prefixes for prompt caching, and loading tools efficiently.
- GPT-5.6 Sol played a key role in autonomously optimizing production kernels and improving inference stack performance.
- These continuous, compounding improvements aim to deliver more cost-efficient intelligence to users and customers.