GPT-5.4 beat human performance on desktop tasks and missed a question a child would get right. Both are true. Here's what to do with that.

I asked ChatGPT 5.4 a question today. A simple one: “I need to wash my car. The carwash is 100 meters away. Should I walk or drive?”

GPT-5.4 beat human performance on desktop tasks and missed a question a child would get right. Both are true. Here's what to do with that.

TL;DR

  • GPT-5.4 failed a simple logic question about walking or driving 100 meters to a carwash, while Claude Opus 4.6 and Gemini 3.1 Pro answered correctly.
  • Despite this failure, GPT-5.4 demonstrates genuine strengths in quantitative modeling, file processing, and self-knowledge, according to blind evaluations.
  • The article suggests that OpenAI's focus might be on building 'agentic infrastructure' rather than solely a chatbot, with promoted benchmarks reflecting this goal.
  • Key findings also include a 'toggle' that splits one model into two products, areas where GPT-5.4 falls apart, and the significance of a major AI hire.