explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • Gemini 3.7 Flash: what actually shipped
  • The benchmark data — and its honesty caveat, upfront
  • Cross-referencing Grok 4.6 honestly
  • Pricing snapshot
  • When each model actually makes sense
  • Bottom line
  • Related on explainx.ai
← Back to blog

explainx / blog

Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6: The Real Numbers

Google's own launch charts put Gemini 3.7 Flash ahead of Sonnet 5 and GPT-5.6 Terra on three of four benchmarks. Grok 4.6 isn't even on the chart — here's why that matters.

Aug 14, 2026·16 min read·Yash Thakker
Google GeminiGemini 3.7 FlashGrok AIClaude Sonnet 5GPT-5.6AI Benchmarks
go deep
Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6: The Real Numbers

Google shipped Gemini 3.7 Flash on August 14, 2026 — just three weeks after Gemini 3.6 Flash — and called it, in Sundar Pichai's words, "a workhorse for performance at great value." Per Google's official announcement (Tulsee Doshi, blog.google, August 13, 2026) and its model card, the price is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50/$7.50 on January 1, 2027 — with a 1 million token context window and multimodal input, live day one in the Gemini API, AI Studio, and Antigravity. Google simultaneously cut Gemini 3.6 Flash's price to match ($0.75/$3.75) as of August 13, so the two Flash generations run at identical intro pricing right now. explainx.ai flagged this pricing the day before as an unverified X leak — the rumored input number turned out correct, though the leak never mentioned the output price, the price step-up in 2027, or the benchmark charts Google actually shipped with.

Those charts are the interesting part. Google's own launch data puts Gemini 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra on three of four published benchmarks — but the charts don't include Grok 4.6 at all, which fresh off its own August 12 launch is very much part of the conversation people are having about this tier of model. This piece lays out Google's numbers exactly as published, cross-references Grok 4.6 honestly from separate sources where it exists, and is explicit about the one place two different benchmarks share a name but not a scale — the same discipline explainx.ai applied comparing Fable 5, Grok 4.6, GPT-5.6 Sol, and Qwen3.8-Max and Sonnet 5 against GPT-5.6 Luna Max.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

QuestionDirect answer
What did Google actually ship?Gemini 3.7 Flash — $0.75/M input tokens through end of 2026, 1M context, multimodal, live in Gemini API / AI Studio / Antigravity
Does it beat Sonnet 5?On Google's own chart, yes on 3 of 4 benchmarks — AutomationBench, Code Arena, and a near-tie on FrontierCode
Does it beat GPT-5.6 Terra?Mixed — ahead on AutomationBench and Code Arena, roughly tied on FrontierCode, clearly behind on DeepSWE V1.1
Where's Grok 4.6 on these charts?Nowhere. Google's four benchmark charts don't include it at all — say so, don't guess
Can we compare Grok 4.6 anyway?Only by cross-referencing separate sources: Grok 4.6 High scored 1618 Elo on a different Code Arena WebDev leaderboard, not Google's 1588
Are the two "Code Arena" numbers the same test?No — different panels (Google DeepMind's own vs. LMArena/Arena.ai), different absolute scales. Don't average them
Is this an independent benchmark?No — it's Google's own vendor-published launch chart, same caveat as any lab's first-party numbers
Who actually wins DeepSWE V1.1?GPT-5.6 Terra, at 69.6% — the one row where Gemini 3.7 Flash trails clearly

Gemini 3.7 Flash: what actually shipped

Google's August 13-14, 2026 announcement — via official posts from @Google, @OfficialLoganK (Logan Kilpatrick), @GoogleDeepMind, and @GoogleAIStudio — pitches Gemini 3.7 Flash as "the most intelligent workhorse model yet for coding and agents." The headline numbers:

SpecGemini 3.7 Flash
Input price$0.75 per million tokens, held through end of 2026
Context window1 million tokens
ModalityMultimodal
AvailabilityGemini API, AI Studio, Antigravity — live at launch
Prior generationGemini 3.6 Flash, launched July 21, 2026 at $1.50/M input — roughly 3 weeks earlier

A 50% input-price cut three weeks after the last Flash release is an aggressive cadence even by Google's own recent pace — 3.5 Flash (June) → 3.5 Flash-Lite and 3.6 Flash (July 21) → 3.7 Flash (August 14) is four Flash-tier releases inside nine weeks. That's the same rapid-iteration pattern explainx.ai flagged as a signal Google is defending the high-volume, cost-sensitive segment against GLM and DeepSeek pricing, not chasing frontier-reasoning parity outright.

The benchmark data — and its honesty caveat, upfront

Every number in the four tables below comes from Google DeepMind's own launch charts, methodology published at deepmind.google/models/evals-methodology/gemini-3-7-flash. That means this is first-party, vendor-published data — the same category as any lab's own launch chart, including SpaceXAI's Grok 4.6 table or Anthropic's own Sonnet 5 comparisons. Treat it the way explainx.ai treats any self-reported benchmark: directionally informative, not an independent audit. Nobody ran all four models in one neutral harness for this specific chart set.

The second, equally important caveat: Grok 4.6 is not on any of these four charts. Google's comparison set is its own generations plus Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Where this piece brings in a Grok 4.6 number, it's flagged explicitly as cross-referenced from a separate source — not read off the same chart.

AutomationBench — enterprise workflow automation

ModelScore
Gemini 3.7 Flash30.4%
GPT-5.6 Terra23.6%
Gemini 3.6 Flash17.0%
Claude Sonnet 510.7%
Grok 4.6Not tested / no published score on this chart

This is the largest single gap in Google's chart set. Gemini 3.7 Flash nearly doubles Gemini 3.6 Flash's own prior score (17.0% → 30.4%) in three weeks, and posts nearly 3x Claude Sonnet 5's 10.7%. AutomationBench measures enterprise workflow automation specifically — multi-step business-process tasks, not general coding — so this gap says more about agentic task-completion on structured business work than it does about raw coding ability.

Code Arena — web development (Google's chart)

ModelElo
Gemini 3.7 Flash1588
Claude Sonnet 51541
Muse Spark 1.21535
Gemini 3.6 Flash1538
GPT-5.6 Terra1523
Grok 4.6Not on this chart

Read this table carefully — it is not the same leaderboard as the Grok 4.6 comparison piece. explainx.ai's Fable 5 vs Grok 4.6 comparison cites a different "Code Arena WebDev" leaderboard, run by LMArena/Arena.ai, with entirely different absolute numbers: Fable 5 at 1627, GPT-5.6 Sol xHigh at 1622, Grok 4.6 High at 1618, Grok 4.5 at 1553. Both charts use the name "Code Arena" for a web-development coding evaluation, and both are Elo-style scores — but they are two different panels, judged by different raters, on a different scale. Gemini 3.7 Flash's 1588 on Google's own chart and Grok 4.6's 1618 on LMArena's chart are not directly comparable numbers, even though they look like they belong on the same axis. This is exactly the trap explainx.ai's guide to reading AI benchmarks warns about: a shared benchmark name is not a shared benchmark.

DeepSWE V1.1 — long-horizon software engineering

ModelScore
GPT-5.6 Terra69.6%
Gemini 3.7 Flash65.3%
Muse Spark 1.254.9%
Claude Sonnet 553.8%
Gemini 3.6 Flash48.6%
Grok 4.6Not on this chart

This is the one row where Gemini 3.7 Flash does not lead — GPT-5.6 Terra wins DeepSWE V1.1 outright, by more than 4 points. It's worth naming plainly rather than burying: Google's own launch chart shows its newest Flash model losing a benchmark to a competitor, and that's a more credible chart for including it. For color, SpaceXAI's own Grok 4.6 launch table separately reported Grok 4.6 High at 65.9% and GPT-5.6 Sol Max at 73% on the same-named DeepSWE V1.1 — a different GPT-5.6 tier (Sol, not Terra) scored on a different vendor's chart, so treat that as directional confirmation that DeepSWE V1.1 favors GPT-5.6's family generally, not a number you can merge into this table.

FrontierCode 1.1 Main — production code quality

ModelScore
Gemini 3.7 Flash43.6%
Claude Sonnet 542.7%
GPT-5.6 Terra41.3%
Gemini 3.6 Flash34.4%
Grok 4.6Not on this chart

The closest race on the whole chart set — Gemini 3.7 Flash, Sonnet 5, and GPT-5.6 Terra sit inside a 2.3-point band. Call this a statistical three-way tie on production code quality rather than a clear win; a 0.9-point edge over Sonnet 5 is well inside the noise you'd expect from any single-run vendor chart.

Beyond the four headline benchmarks: where Gemini 3.7 Flash does not lead

The four tables above are the ones the task at hand centers on, but Google's own launch chart runs wider than that — and the fuller table undercuts any temptation to read this as a clean Gemini sweep. Same source (Google DeepMind's evals methodology page), same five-model set:

BenchmarkGemini 3.7 FlashGemini 3.6 FlashClaude Sonnet 5GPT-5.6 TerraMuse Spark 1.2
Artificial Analysis Intelligence Index5652555757
Terminal-bench 2.1 (agentic terminal coding)85.8%78.0%80.4%87.4%82.9%
Terminal-bench 3.0 (general agent capabilities)14.9%5.4%14.6%20.8%—
GDPVal-AA v2 Elo (knowledge work)15251422159815781628
Harvey LAB-AA (complex legal workflows)90.7%85.1%90.1%85.2%—
OSWorld-2.0 (agentic computer use)38.1%33.8%39.6%50.2%—
HLE-Verified (multidisciplinary expert reasoning)53.6%51.2%31.0%51.1%—
LVBench (long video understanding)85.4%84.2%68.5%78.9%—

Four separate models lead at least one row here, which is the honest picture: GPT-5.6 Terra actually leads the widest set — the AA Intelligence Index (tied with Muse Spark 1.2 at 57), both Terminal-bench versions, and OSWorld-2.0 by a wide 12-point margin. Muse Spark 1.2 leads GDPVal-AA v2 outright at 1628 Elo. Claude Sonnet 5 beats Gemini 3.7 Flash on GDPVal-AA v2 (1598 vs 1525) and OSWorld-2.0 (39.6% vs 38.1%) — real counterpoints to the AutomationBench and Code Arena leads covered above, and worth knowing if your workload leans toward knowledge-work or computer-use agents rather than coding. Gemini 3.7 Flash's clearest genuine strengths in this fuller table are HLE-Verified and LVBench — the latter unsurprising given Gemini's video-native multimodal design, the former a real reasoning-benchmark win where Sonnet 5 notably lags at 31.0%.

Terminal-bench 3.0's absolute scores (5.4%-20.8%) are worth a separate note: this is a hard, newer suite where every model in the set scores under 21%. Don't read GPT-5.6 Terra's win there the same way you'd read a 70%+ benchmark lead — it's a lead within a range where all five models are still mostly failing the task.

Cross-referencing Grok 4.6 honestly

Since Google's charts leave Grok 4.6 out entirely, the only responsible way to place it in this four-way conversation is to pull real numbers from where Grok 4.6 actually has been tested — and say clearly that none of it is the same chart as the numbers above.

SourceMetricGrok 4.6 result
SpaceXAI's own Aug 12 launch tableDeepSWE V1.165.9% (Grok 4.6 High)
LMArena/Arena.ai Code Arena WebDevElo1618 (Grok 4.6 High) — different leaderboard than Google's Code Arena above
Product Compass Bug Hunt Bench v10Bugs fixed / cost / time42 of 105 bugs, $22.73, 34 minutes
RuneScape BenchComposite (⟨ln⟩)6.28 — highest of that panel

None of these four numbers were measured against Gemini 3.7 Flash, Sonnet 5, or GPT-5.6 Terra in the same run. The DeepSWE V1.1 row is the closest thing to a genuine cross-check — both SpaceXAI and Google DeepMind published scores under the same benchmark name — and even there, Grok 4.6's 65.9% sits between Gemini 3.7 Flash's 65.3% and GPT-5.6 Terra's 69.6% on Google's chart, which is suggestive but not proof, since it's two different labs running the same-named suite independently. AutomationBench and FrontierCode 1.1 Main simply have no published Grok 4.6 number anywhere at the time of writing — not a gap this piece is going to paper over with an estimate.

One independent anecdote worth naming, with the same single-run caveat explainx.ai applied to the banana-drawing and cost-comparison threads in the Fable-5-vs-Grok-4.6 piece: a Hacker News post from developer jjcm (662 points, 376 comments) ran an image-to-HTML fidelity test — reproduce a source image as working HTML/CSS — across Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6. Opus 5 won on visual polish, but several commenters flagged it as notable that Gemini 3.7 Flash outperformed Grok 4.6 on the same task, given Grok 4.6's momentum off its own recent launch. Treat that as one prompt, one judge, no scoring rubric — color that complicates a simple "Grok 4.6 is the value pick" narrative, not a benchmark result to weigh against the tables above.

Pricing snapshot

ModelList price (per million tokens)Notes
Gemini 3.7 Flash$0.75 input / $3.75 output, through Dec 31, 2026Rises to $1.50/$7.50 on Jan 1, 2027; Google cut Gemini 3.6 Flash to match ($0.75/$3.75) the same week
Gemini 3.6 Flash$0.75 input / $3.75 output (price-matched, Aug 13, 2026)Was $1.50/$7.50 at its own July 21 launch — this is a mid-cycle cut, not the original price
Claude Sonnet 5$2 input / $10 outputAnthropic's permanent pricing as of August 11, 2026
GPT-5.6 Terra$2 input (per Google's own comparison chart)Sits between Luna (cheapest) and Sol (priciest) in OpenAI's own GPT-5.6 tier structure
Muse Spark 1.2$1.25 input / $4.25 output (per Google's own comparison chart)Included in Google's launch table as a fourth competitor, not covered elsewhere on explainx.ai yet
Grok 4.6$2 input / $6 outputSame as Grok 4.5; fast variant is 2x ($4/$12) — not part of Google's own pricing chart

Gemini 3.7 Flash's $0.75 input rate undercuts Sonnet 5's $2 by more than 60% and Grok 4.6's $2 by the same margin — on sticker price alone, it's the cheapest model in this comparison during the intro window, before even accounting for the AutomationBench and Code Arena leads on Google's own chart. That gap narrows on January 1, 2027, when Gemini 3.7 Flash's own price doubles to $1.50/$7.50. As explainx.ai found comparing Sonnet 5 against GPT-5.6 Luna Max, sticker price and cost-per-completed-task diverge once you factor in tokens burned per task — nobody has published a cost-per-task figure for Gemini 3.7 Flash yet, so treat the pricing table as list price, not a verified cheapest-per-task claim.

When each model actually makes sense

PriorityBest pick on this dataWhy
Cheap, high-volume agentic codingGemini 3.7 FlashLeads AutomationBench and Code Arena on Google's own chart, at roughly a third of Sonnet 5's or Grok 4.6's input price
Long-horizon software engineeringGPT-5.6 TerraOnly model to clearly win DeepSWE V1.1 on Google's chart (69.6%)
Production code quality at the marginAny of the top threeGemini 3.7 Flash, Sonnet 5, and GPT-5.6 Terra sit within 2.3 points on FrontierCode 1.1 Main
Independently-verified long-horizon agent executionGrok 4.6, with caveatsHighest RuneScape Bench composite and fast, cheap Bug Hunt Bench run — but zero overlap with Google's own chart set
Claude Code / Anthropic ecosystem lock-inClaude Sonnet 5Still the only option here with native Claude Code OAuth subscription billing in CI

Bottom line

No single model wins this comparison cleanly, and the honest version of that sentence has three layers. First, on the four core benchmarks this piece leads with, Gemini 3.7 Flash tops three — AutomationBench by a wide margin, Code Arena web development, and a near-tie on FrontierCode — while GPT-5.6 Terra wins the fourth outright (DeepSWE V1.1, 69.6% to Gemini's 65.3%). Second, once you widen to Google's fuller benchmark set, GPT-5.6 Terra actually leads the most rows overall — both Terminal-bench versions and OSWorld-2.0 by a wide margin — and Muse Spark 1.2 and Claude Sonnet 5 each post real wins of their own (GDPVal-AA v2, and Sonnet 5 edges Gemini on OSWorld-2.0 too). This is not a Gemini sweep by any honest reading of Google's own data. Third, and just as important: this is Google's own vendor-published data, and Grok 4.6 isn't on any of it. Every Grok 4.6 number in this piece came from a different source, measuring a similarly-named benchmark on a different scale — treat the "Code Arena" comparison especially carefully, since Google's 1588 and LMArena's 1618 look like they belong on the same axis and don't. Gemini 3.7 Flash's price is real and aggressive at $0.75/M input tokens through the end of 2026; its benchmark lead is real on several of Google's own rows and needs independent confirmation before it's treated as settled. Run your own workload before switching a default.

Related on explainx.ai

  • GPT-5.6 Sol Ultrafast mode: 750 TPS on Cerebras, same-day answer to this launch's speed positioning — HN's jjcm ties the two launches together directly
  • Gemini 3.7 Flash pricing leak: the rumor before the confirmed launch — what was speculative on August 13 that Google confirmed a day later
  • Grok 4.6 launch: official evals, pricing, Cursor access — the source for every Grok 4.6 number cross-referenced above
  • Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max: full comparison — the source of the other Code Arena WebDev leaderboard cited above
  • Claude Sonnet 5 vs GPT-5.6 Luna Max: cost comparison
  • Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: what actually changed
  • How to read AI benchmarks
  • AI benchmarks complete guide
  • Anthropic makes Claude Sonnet 5 pricing permanent

Sources: Introducing Gemini 3.7 Flash, Tulsee Doshi, blog.google, August 13, 2026 · Gemini 3.7 Flash model card, Google DeepMind · evals methodology at deepmind.google/models/evals-methodology/gemini-3-7-flash · Official X posts from @Google, @OfficialLoganK, @GoogleDeepMind, @GoogleAIStudio, August 13-14, 2026 · SpaceXAI Grok 4.6 launch, August 12, 2026 · Product Compass Bug Hunt Bench v10 · LMArena/Arena.ai Code Arena WebDev leaderboard · Hacker News, jjcm image-to-HTML fidelity thread (662 points, 376 comments)


Benchmark figures and pricing reflect August 14, 2026. AutomationBench, Code Arena, DeepSWE V1.1, and FrontierCode 1.1 Main figures for Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2 come from Google DeepMind's own first-party launch chart, not an independent audit. Grok 4.6 does not appear on that chart; every Grok 4.6 figure in this piece is cross-referenced from a separate, independently-sourced benchmark and should not be read as measured on the same scale. Re-run your own workload before switching a production default.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 13, 2026

Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max: Who Actually Wins?

Grok 4.6's August 12 launch set off a fresh round of four-way frontier comparisons on X. explainx.ai pulls together three independent benchmarks — a 105-bug hunt across two real repos, a long-horizon RuneScape XP test, and LMArena's Code Arena WebDev leaderboard — plus the viral cost and creativity threads, to see how Fable 5, Grok 4.6, GPT-5.6 Sol, and Qwen3.8-Max actually compare.

Jul 13, 2026

Gemini 3.5 Pro Benchmark Leak — Beating Fable 5 and GPT-5.6? (July 17 Target)

@EntelligenceAI claims Gemini 3.5 Pro tops Fable 5 and GPT-5.6 in internal tests with a July 17 launch window. X replies say wait for real evals — explainx.ai separates leak hype from what Google must prove.

Aug 14, 2026

GPT-5.6 Sol Ultrafast Mode: 750 Tokens/Sec via Cerebras, No Pricing Yet

OpenAI's August 13 preview of Ultrafast mode runs GPT-5.6 Sol at up to 750 tokens per second on Cerebras silicon — 14x the model's normal speed. It ships first to a select group of API customers, with no pricing and no Codex or ChatGPT access confirmed, drawing pointed criticism from paying subscribers and independent commentary tying it to competitive pressure from Gemini 3.7 Flash.