explainx / blog
Grok 4.5: VulcanBench Lead and Graffiti Conjecture 284
Grok 4.5 hits 91.3% on VulcanBench v3 at $0.32/solved task and a Capy agent on Grok 4.5 Medium reportedly refutes Graffiti Conjecture 284. What holds.
explainx / blog
Grok 4.5 hits 91.3% on VulcanBench v3 at $0.32/solved task and a Capy agent on Grok 4.5 Medium reportedly refutes Graffiti Conjecture 284. What holds.

Jul 9, 2026
SpaceXAI's Grok 4.5 lands July 8–9 at $2/$6 — Musk says Opus 4.7-class, faster and cheaper. Opus 4.8 still leads SWE-Bench Pro and independent DeepSWE 1.1. Snorkel GDPval+ favors Grok on professional work. Full decision tree.
Jul 8, 2026
Cursor's July 8 blog confirms Grok 4.5 live in the IDE — joint SpaceXAI training, broad STEM mix beyond Composer 2.5's coding focus, double usage week one. Benchmarks vs Opus 4.8, GPT-5.5, Fable 5 — plus CursorBench contamination note.
Jun 28, 2026
xAI's Grok 4.5 has entered private beta at SpaceX and Tesla, built on the 1.5T V9 foundation model with supplemental Cursor training data. Elon Musk says early evaluations show performance close to — and possibly exceeding — Anthropic's Claude Opus.
Two July stories collided into one Grok headline: consumer rollout of Grok 4.5 across Grok.com, X, and mobile apps, and a louder claim that the same model family is winning coding cost benchmarks while an agent on Grok 4.5 Medium helped poke a hole in a decades-old graph theory conjecture. explainx.ai separates launch marketing, VulcanBench numbers, and the Graffiti 284 claim so you can decide what is shipped, what is scored, and what still needs a mathematician’s pen.
We already covered the Cursor/SpaceXAI public launch and the Opus comparison. This piece is the late-July update: wide consumer access, VulcanBench Report No. 4, and the math thread.
| Question | Short answer |
|---|---|
| What’s new this week? | Grok 4.5 called out as live on Grok.com, X, iOS, Android |
| Coding claim? | 91.3% on VulcanBench v3 (21/23) at medium effort |
| Cost claim? | ~$0.32 per solved task in that report (suite ~$6.67) |
| Speed claim? | Vendor: ~80 TPS; efficiency pitch is steps + $/task |
| Math claim? | Agent on Grok 4.5 Medium → counterexample to Graffiti 284 |
| Peer reviewed? | No public math paper yet — traces + checks reported |
| Which model in the app? | UI often opaque — use API IDs for serious work |
| Replace Claude/GPT? | Suite-specific win; not a universal crown |
xAI’s Grok account framed Grok 4.5 as “our most capable model yet,” available across the consumer surfaces people already use: the web app, X, and the mobile apps. That matters because earlier July coverage was skewed toward Cursor, Grok Build, and API/console access — developer channels — not every free or Premium X user.
The practical effect:
grok-4.5 next to the reply (a recurring X complaint: “it said 4.5 last week; today it says 4.0”).For builders, treat consumer Grok like any other chat surface: great for exploration, weak for reproducible agent loops. If you need deterministic model selection, stay on the console/API path documented in our launch guide and pair it with a proper agent harness.
VulcanBench is not a vibes leaderboard. The July Technical Report No. 4 run compared Grok 4.5, Claude Fable 5, and GPT-5.6 Sol across 23 software tasks at low/medium/high effort — 207 graded runs with hidden tests. Tasks are drawn from real merged PRs across multiple languages.
Headline row that went viral:
| Config | Score | pass@1 | Approx $/solved |
|---|---|---|---|
| Grok 4.5 (medium) | 21/23 | 91.3% | ~$0.32 |
| Grok 4.5 (high) | 21/23 | 91.3% | ~$0.42 |
| Fable 5 (low)† | 20/23 | 87.0% | ~$0.61 |
| GPT-5.6 Sol (high) | 20/23 | 87.0% | ~$0.80 |
| Grok 4.5 (low) | 19/23 | 82.6% | ~$0.18 |
† Fable 5’s card notes safety refusals on some runs with Opus 4.8 fallback — read the report’s methodology before declaring a three-way knockout.
The report’s own takeaway: Grok plateaus at medium. High effort burned ~33% more tokens and ~31% more money for the same 21/23, failing the same two tasks. That is the efficiency narrative in one sentence — not “Grok is smarter at every temperature,” but “the Pareto front for this suite sits at medium.”
Two frontier-hard tasks went 0-for-27 across models and efforts. So 91.3% is a ceiling of the suite, not proof Grok clears every repo on Earth.
Cross-check against:
VulcanBench rewards cheap, correct PR-style fixes. It does not measure long-horizon research agents, design taste, or cyber refusal behavior. If your workload is closer to Grok Build office agents or multi-hour math search, borrow the cost mindset — do not copy the percentage onto a slide as “best model.”
Separate from coding benches, July chatter claimed a ~30-year open graph-theory statement — Graffiti Conjecture 284, from Siemion Fajtlowicz’s automated conjecture generator — fell to a counterexample found with help from a Capy agent running Grok 4.5 Medium.
Reported shape of the claim:
Famous named graphs are exactly where automated conjecture systems and AI search collide: they are small enough to compute invariants, weird enough that hand proofs get sticky, and well documented enough that an agent can retrieve the right object. Finding that δ* (or whichever dual-degree quantity the conjecture bounds) violates Graffiti 284 on this graph is the kind of search + check win AI already demonstrated elsewhere in 2026.
| Claim | Who | Verification bar |
|---|---|---|
| Planar unit distance / Erdős line | OpenAI internal model | External mathematicians + papers — see our coverage |
| Jacobian conjecture counterexample | Human + Fable 5 assist | Explainer: what the conjecture is |
| Expert ChatGPT session | Terence Tao | Shared conversation analysis |
| Graffiti 284 | Capy + Grok 4.5 Medium | Developer traces; await formal write-up |
The honest product lesson matches “will AI replace mathematicians?”: models are getting better at finding counterexamples; humans still own definitions, publication, and social verification.
X threads around the wide release cluster into three jobs:
If you are wiring Grok into team workflows, steal identity and audit ideas from Block Buzz / Nostr agent rooms rather than dumping an unlabeled consumer chat into production.
For agent product managers, the combined lesson is effort ladders beat brand loyalty. Medium Grok winning a PR suite does not cancel Opus on multilingual SWE or Fable on long research loops — it means your router should price the rung, not the logo. Pair that with the math-thread humility: a Slack agent finding Hoffman–Singleton is exciting; a refereed note is shipping.
Primary sources to verify: VulcanBench Technical Report No. 4 model card (July 2026 3-way run) · xAI / Grok product posts on X · developer write-ups on Graffiti 284 / Hoffman–Singleton · Artificial Analysis Grok 4.5 notes.
Benchmarks, pricing, and math claims reflect public materials and press summaries as of July 24, 2026. VulcanBench figures are suite-specific; Graffiti Conjecture 284 status awaits formal mathematical publication. Consumer availability and model labeling can change without notice — verify in your own console before committing workloads.