Four words. No context. 169,400 views. On August 15, 2026, OpenAI president Greg Brockman posted "hedonic treadmill of model expectations" on X — and the internet immediately understood exactly what he meant, even without an explanation.
TL;DR
| Question | Answer |
|---|---|
| Who said it? | Greg Brockman, OpenAI president and co-founder, August 15, 2026 |
| What does it mean? | A real psychology concept — people adapt to gains and reset to a baseline — applied to how AI model launches feel |
| Why did it land? | It names something practitioners already felt: each new model resets what "impressive" means within days |
| Did he explain it? | No — the tweet is four words with zero elaboration, which is part of why it spread |
| What's the irony people flagged? | OpenAI, as a frontier lab, both rides and drives the release cadence the tweet describes |
| Is this specific to one launch? | Unclear — it landed the same week as OpenAI's own Astra research release and rumored Fable 5.1 reports |
The tweet, and the concept behind it
The hedonic treadmill — also called hedonic adaptation — comes from a real, decades-old body of psychology research, most associated with Brickman and Campbell's 1971 paper on adaptation-level theory. The core finding: people experiencing a major positive change — a raise, a new relationship, a lottery win — get a real mood boost, but it fades faster than expected, and satisfaction settles back toward a fairly stable personal baseline. The metaphor is a treadmill because you can run as fast as you want and the ground underneath doesn't actually move — effort (or gain) doesn't translate into lasting position.
Applied to AI model releases, the mechanism is almost too clean: a new frontier model ships, it's genuinely better on every axis that matters, users are amazed for a period measured in days — and then that new capability becomes the assumed floor. The next release has to clear an already-elevated bar just to register as exciting at all, and the cycle resets immediately. Nobody stays impressed at the new baseline; they start asking what's next.
Why the replies prove the point
The reply thread reads like a live demonstration of the concept it's describing. Alan Blair: "Nothing's matched the thrill of Fable coming out. I really hope you guys have your Fable moment soon" — nostalgia for a specific launch high that's already worn off, directed at a competitor's model, less than a year removed from Fable 5's original release. David Stark, more pointed: "Yeah, maybe just give us back 4o the only good thing you've got to offer. #4oForAll" — a direct callback to the #Keep4o movement from earlier in 2026, when OpenAI retired GPT-4o in February and brought it back after sustained user backlash — a case study in exactly the adaptation-and-attachment cycle the tweet names, just running in reverse (users adapting to an old model's absence rather than a new model's presence).
DK's reply reframes the metaphor sharply: "We need to make a hedonic treadmill that just acts like a regular treadmill and makes you healthier the more you do it, that would be good for hedonists I think" — noting that the AI-release version of the treadmill produces no lasting gain in user satisfaction despite real, measurable capability gains, unlike an actual treadmill which does produce fitness gains proportional to effort.
The irony nobody let slide
The sharpest replies target Brockman's own position. Elian: "you built the treadmill, don't act surprised everyone's running on it." Ehsan Azish: "funny to see the guy shipping the treadmill also naming it." Both land on the same real tension: Brockman is president of one of the two or three labs setting the actual release cadence the tweet is describing. OpenAI's own August 2026 output — the Astra research release on August 1, which solved ten previously unsolved problems in mathematics and theoretical computer science — is itself a data point in the cycle he's naming. Whether that undercuts the observation or just makes it a more credible one, coming from someone with the clearest possible view of the mechanism, is genuinely a matter of read. Hari Seldon's reply pushes back from a different angle: "is it really hedonic if its improving people's lives and civilization at large?" — a fair challenge to the metaphor, since the psychological literature's hedonic treadmill describes subjective satisfaction resetting, not that the underlying gains are illusory or reversible.
What this actually explains, for people evaluating models
Set aside the tweet-reply theater — there's a genuinely useful framing underneath for anyone whose job involves picking which model to use or reporting on new releases:
- "This feels underwhelming" is not the same as "this isn't better." A model that's a clear step up on benchmarks can still land flat with users purely because their baseline already moved. Evaluate against the actual measured capability, not against your gut reaction to the announcement.
- The launch-week reaction and the six-month verdict are different signals. Immediate hype and immediate disappointment are both adaptation artifacts; what a model is actually good for tends to become clearer once the initial baseline-reset settles.
- Nostalgia for an old model is a real signal worth taking seriously, not just sentimentality — the #Keep4o precedent shows it can represent a genuine capability gap (tone, personality, specific workflow fit) that a newer, more capable-on-paper model doesn't cover, which is exactly why OpenAI reversed course and brought GPT-4o back.
- Labs shipping the cycle and labs critiquing the cycle are often the same labs — that's not hypocrisy so much as the position everyone in this industry is actually in. It's worth reading vendor commentary about "hype" or "expectations" with that in mind, from any lab.
Separate emotional adaptation from a useful baseline
The metaphor becomes practical when you keep two records: how a release feels and what it changes in your work. A model can feel ordinary because you already expect fluent answers, while still reducing the number of corrections a particular task needs. Conversely, an impressive demo can produce little improvement on your own documents, code, or support queue.
Choose a few recurring tasks before trying the next release. Preserve the prompts, permitted sources, expected deliverables, and known failures. This does not need to become a large public benchmark. It gives you a stable reference that cannot quietly move every time your expectations do.
For a document assistant, one task might require extracting a deadline and citing the exact paragraph. Another might require admitting that the document does not establish an answer. Better prose alone should not pass either task. The relevant change is whether the model reliably follows the source and handles uncertainty.
How to compare without chasing the announcement
Review outputs with model names hidden where practical. Ask a reviewer to identify factual errors, missing requirements, and unnecessary work before choosing a favorite. If the reviewer recognizes a distinctive voice, record that limitation rather than calling the comparison blind.
Include the model already serving the task. Launch comparisons often focus on two new releases while forgetting the existing baseline. The real product decision is whether a replacement improves enough to justify integration, verification, and any changed operating cost.
Measure correction effort, not only first-response appeal. Suppose the new model writes a stronger opening but misses the same required detail in every run. A team may enjoy its voice and still decide the current model is more dependable for that workflow. That is a coherent decision; the adaptation metaphor does not invalidate it.
Do not explain away real regressions
Hedonic adaptation is an interpretation of subjective reaction, not a universal explanation for dissatisfaction. Users can notice real losses in instruction following, language coverage, latency, or interaction style. Asking for concrete examples helps distinguish a changed expectation from a changed capability.
When someone prefers an older model, preserve their task and compare the outputs instead of labeling the preference nostalgia. Tone can matter in a coaching or writing tool, while strict formatting may matter more in an extraction pipeline. A single overall intelligence score does not resolve those preferences.
For product teams, write release notes around observable behavior. Explain which workflow changed and what remains uncertain. Avoid claiming that every task improved because an aggregate benchmark rose. Specific notes give users a way to test the change and report a reproducible regression.
Set a review cadence that fits your work
If daily model news interrupts delivery, evaluate on a planned cadence and reserve urgent reviews for a material failure or dependency change. Keep a short list of reasons to switch: improved success on a blocking task, a meaningful operating-cost change, or a capability the application genuinely needs.
The useful response to the treadmill is not permanent skepticism or permanent excitement. It is a stable decision process that makes room for both subjective experience and measured task outcomes. You can appreciate a new release without rebuilding your workflow immediately, and you can keep an older choice when the evidence still supports it.
Related reading
- Has AI Reached Superintelligence? The Astra Debate, Defined
- What Is AI Slop? A Practical Definition
- Claude Fable 5 and Mythos 5: SOTA Autonomy and Safeguards
- What Is Mermaid Slop? The Generic AI Diagram Problem, Explained
- Hedonic Treadmill dictionary entry
External: Greg Brockman on X · OpenAI retiring GPT-4o and older models
The tweet and its replies are quoted as publicly posted on X on August 15, 2026. explainx.ai has not independently confirmed what specific event, if any, prompted Brockman's post.
