Claude Code and Codex bet on different harnesses. Your team is compounding one of them every week + 2 prompts to audit which.

On February 5th, Anthropic dropped Claude Opus 4.6, and within the hour, OpenAI responded with GPT-5.3-Codex. The developer community spent the next three weeks running benchmarks, writing comparisons, and arguing about which model won.

Claude Code and Codex bet on different harnesses. Your team is compounding one of them every week + 2 prompts to audit which.

TL;DR

  • AI model performance can vary drastically (78% vs 42%) depending on the harness, a difference often missed in standard evaluations.
  • Five architectural decisions contribute to team lock-in, with cumulative compounding effects.
  • A $2 billion company is spending 100% of its revenue on API costs, highlighting ignored economic factors.
  • A harness audit tool scores lock-in across five dimensions and suggests appropriate tools.
  • An executive brief generator translates the audit into financial terms for leadership alignment.