explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What Google has actually said
  • The demos, in the order you can actually check them
  • Why these clips look like a generation jump
  • What people are asking
  • A skeptical checklist before you forward the clip
  • What to do this week
  • Related reading
← Back to blog

explainx / blog

Gemini 4 Arena Demos: The 3D Jump, and What Is Still a Rumor

Google DeepMind, Gemini, AI Models, Coding Agents, Model Leaks

Arena clips claim a rumored Gemini 4 beats Opus 5.5 and Astra at 3D. Google has not confirmed the model. Here is what the demos show.

Sep 27, 2026·13 min read·Yash Thakker
add explainx.ai
go deep
Gemini 4 Arena Demos: The 3D Jump, and What Is Still a Rumor

A September 27, 2026 feed summary says suspected Gemini 4 Pro tests show a large 3D jump: horse motion ahead of GPT-6 Astra, a floatplane ahead of Claude Opus 5.5, generation two to three times faster, plus a voxel lighthouse and a detailed 3D Wii remote. The same summary says Google DeepMind chief Koray Kavukcuoglu confirmed an early version before year's end, and that the tests are unverified.

The second sentence is the one to fix first. Kavukcuoglu did confirm early post-training. He did not set a deadline of December 31. On September 24 he said Google wants an early post-training build as soon as possible, hoping to ship much earlier than the end of 2026, with no calendar date. explainx.ai's Gemini 4 fast-track post has the quote. "Before year's end" is a weaker line than the one he actually used.

The demos are still the reason the thread is large. People are not arguing about a press release. They are watching one-shot 3D programs that look like a different ceiling than last month's checkpoints. This post catalogs the clips we can actually point at, says what the jump looks like if the labels are right, and marks the lines that are still a summary with no primary video.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
Is Gemini 4 released?No. Post-training only, per Kavukcuoglu on September 24, 2026
Ship timing he statedEarly build ASAP, much earlier than end of 2026. No date
Best-documented clipFloatplane vs Opus 5.5, Pranav Reddy, September 26
Astra clipReddy also posted an arena comparison vs GPT-6 Astra and called the margin huge. Prompt text was not in the post we reviewed
Lighthouse you can openA public one-shot scene labeled Gemini 3.8 Flash, not Gemini 4
"2–3x faster"In the September 27 summary. Not a published bake-off
Wii remoteNamed in that summary. No stable primary URL found
What to doPin prompts. Do not route production on an arena nickname

What Google has actually said

The confirmed facts are short, and they are easy to inflate.

Koray Kavukcuoglu, speaking at The Information's AI Agenda Live Summit on September 24, 2026, said Gemini 4 is in post-training. Google's intention is to release an early post-training output as soon as possible because the team is excited by the results, and to keep iterating quickly. He hoped that build would arrive much earlier than the end of 2026. He did not say "Gemini 4 Pro," did not confirm Chatbot Arena, and did not attach the model to any video.

That sits on top of a skipped generation. Gemini 3.5 Pro was teased and never broadly shipped. The last flagship was Gemini 3 Pro in November 2025. Through the summer Google shipped Flash revisions, including the official Gemini 3.8 Flash on September 2. Rivals already have public frontier models: GPT-6 Astra, and Claude Opus 5.5 as of September 22. A visually strong stealth checkpoint would be on-message for a comeback. It is not the same thing as a model card.

Aggregator posts have already turned "much earlier than year's end" into an October launch, a 10 million token window, and robotics motor control. None of that is in Kavukcuoglu's remarks. Treat those rows as fan fiction until a Google doc says them.

The demos, in the order you can actually check them

1. Floatplane versus Opus 5.5

This is the clip with a durable URL. On September 26, Pranav Reddy posted a side-by-side labeled Gemini 4 Pro in arena against Claude Opus 5.5 on a realistic 3D floatplane takeoff. The video is the primary source. Viewers describe the Gemini-labeled side as smoother acceleration, cleaner water reflections, and a steadier flight path. Reddy's caption says Gemini 4 Pro completely outperforms Opus 5.5.

explainx.ai's floatplane write-up already covers why one WebGL takeoff is not a coding benchmark. The short version, so this post does not repeat that piece: "physics" in these reels usually means believable motion in a browser scene, not a fluids solver. Opus 5.5's job, in the launch benchmarks and in the five browser games built with it, is long-horizon code with a human in the loop. A 30-second takeoff can look like a blowout and still say nothing about a multi-file bugfix.

The jump people are reacting to is real as a genre of output. One prompt produces an interactive vehicle, a water surface, and a camera that stays coherent for the length of a clip. That is a harder visual program than a still image. It is also exactly the genre where a model trained on shiny demos can look a generation ahead on the first try.

Replies on the same thread already warn that the arena label might be a Flash checkpoint, not Pro. Until Google or the arena operator publishes the model ID, the honest name is "rumored Gemini 4-series arena model."

2. The GPT-6 Astra comparison

About fourteen hours before this post's sources were collected, Reddy posted again: Gemini 4 Pro in arena versus GPT-6 Astra max, with the claim that Gemini 4 Pro beats Astra by a huge margin. The September 27 summary says community comparisons show the same rumored model ahead on horse motion, and that generation was two to three times faster.

We can confirm the existence of Reddy's Astra post and the wording of the summary. We cannot confirm a horse-motion prompt, a frame count, or a stopwatch. The post text we reviewed does not include the prompt. "Huge margin" is the author's verdict on a video, not a score.

If you are choosing between Astra and a future Gemini 4 for motion-heavy codegen, this clip is a reason to add a motion prompt to your eval, not a reason to cancel an Astra route today. Astra is a shipping model. The other side is a nickname.

3. Voxel lighthouse: open it, and read the label

The playable scene people are passing around as proof of taste is Moon Island Light Station, hosted as a Gemini 3.8 Flash one-shot demo. The page is not shy about the name. The hostname says gemini-3-8-flash.

What is on the page is why the clip feels like a jump, whoever trained the weights. It is not a lighthouse sprite. It is a small simulation with state:

  • A granite outpost you orbit, with a lantern room, catwalk, and spyglass.
  • A nor'easter: barometer, dog watch, vessels in peril, a schooner off a reef.
  • A Fresnel lens on clockwork, a kerosene mantle with PSI and tank level, pane salt you can polish, a steam foghorn you can blast.
  • Camera presets and a keeper's standing orders that explain the machinery.

A September 2 post from the person sharing that build said Gemini 3.8 Flash produced it in about five minutes through an agent harness, and that the output was genuinely good rather than merely fast. That is a shipping Flash model doing one-shot interactive fiction plus 3D, weeks before the Gemini 4 Pro rumor peaked.

Folding this lighthouse into "Gemini 4 Pro" is how the summary gets its "precise voxel lighthouse" line. The scene is worth opening. The label on the scene is 3.8 Flash.

4. Voxel pagoda, pelican, and the alias problem

On September 17, Lumina posted a voxel pagoda and wrote that it was Gemini 4 Pro being tested in arena under the name Gemini 3.8 Flash, made in about five minutes. The same day, Harshith posted an SVG of a pelican on a bicycle under that same arena name, and another account claimed a side-by-side of "3.8 Flash" versus "4.0 Pro." Office Chai collected those posts the same day.

This is the identity trap. Gemini 3.8 Flash is a real, documented model with a price and benchmarks. Community testers also say a stronger checkpoint is being served behind that string. Both can be discussed. They cannot be the same sentence. A five-minute pagoda labeled as a stealth Pro and a five-minute lighthouse labeled as shipping Flash are evidence of a visual-codegen moment in September, not a single SKU's scoreboard.

A separate user claimed an internal codename, Argon, and that a long JSON generation kept going until the client crashed, which they read as a higher output ceiling than 3.8 Flash. That is one stress test from one account. It is not a context-window spec. Do not copy "10 million tokens" from later roundups. That figure is not in the posts above.

5. Wii remote and the two-to-three-times line

The September 27 summary praises a detailed 3D Wii remote and says generation was two to three times faster, with better design taste. Poonam Soni posted the same morning that if Gemini 4's benchmarks are real, this is Google's biggest comeback, and pointed at "10 wild examples" people are building with a rumored Arena checkpoint.

We did not find a stable URL for a Wii remote prompt, a list of those ten examples, or a timing sheet that shows a 2x or 3x ratio against Astra or Opus on the same machine. Report them as summary claims. The demos that do have URLs already show the taste argument without the multiplier: dense geometry, working controls, and a scene that holds together after the first prompt.

Speed anecdotes that do exist are absolute, not relative. About five minutes for the pagoda. About five minutes for the lighthouse. Five minutes of agent time for a playable diorama is fast next to a human modeling pass. It is not a measured speedup over GPT-6 Astra until someone publishes both clocks.

Why these clips look like a generation jump

Strip the model name off and the artifacts are still ahead of a typical "draw me a 3D object" prompt from early 2026. Three things show up together.

The output is a program, not a picture. The lighthouse has pressure, a rotating lens, and a foghorn. The floatplane is a motion sequence with a camera. A model that only matches a still would fail these prompts even if the mesh looked fine in frame one.

Motion stays legible. The floatplane commentary is about acceleration and a stable flight path, not a single pretty splash. Reddy's Astra post, whatever the subject, is being received as the same kind of win: the rumored checkpoint keeps a moving subject coherent where the comparison model does not. That is the claim. It is not a published Elo.

The taste is in the props. Keeper's orders, a mercury bath under a Fresnel lens, salt on the glass. Whether that came from 3.8 Flash or a stealth Pro, viewers are reacting to specificity. Generic "AI 3D" looks like a gray asset. These scenes look like someone art-directed them. For a one-shot, that is the part that is hard to fake with a longer prompt on an older model, and it is also the part a cherry-picked prompt can exaggerate.

None of that is a Terminal-Bench score. Opus 5.5 versus GPT-6 Sol is the comparison that has prices and task types. Use the arena reels to decide whether visual codegen belongs in your eval. Use the priced models to decide what serves traffic on Monday.

What people are asking

Is this Gemini 4 Pro or Gemini 3.8 Flash?

Both names are in play, and that is the point. Shipping 3.8 Flash can already one-shot a dense interactive scene, as the lighthouse page shows. Testers also insist a stronger weight is riding under the Flash name in arena. Google has not mapped either claim to Gemini 4 Pro. If your note to the team says "Pro beats Opus," you are ahead of the evidence. If it says "September visual codegen, including shipping Flash, jumped," you are on the clips.

Did Kavukcuoglu promise a 2026 release?

He said he hopes for much earlier than the end of 2026, and he said as soon as possible. A feed summary that stops at "before year's end" drops the urgency he actually signaled and invents a deadline he did not set. There is still no date. October windows on blog roundups are not his sentence.

If the floatplane is real, should we rebuild our 3D stack on Gemini?

Not on a leak. The practical move is a three-prompt folder you can rerun the morning a versioned ID appears:

  1. A vehicle on water with a follow camera, one shot, no follow-up.
  2. A small machine with at least four controls that change state, in the lighthouse style.
  3. The same prompts again with "match this reference" plus a screenshot, because one-shot taste and edit-following are different skills.

Save the HTML. When Google ships a documented Gemini 4 endpoint, run those files against Gemini 3.8 Flash, Opus 5.5, and Astra with the same system instructions. The winner is the one that still works on prompt three, not the one that won a tweet.

Are the benchmarks real?

Soni's post is conditional on purpose: if the benchmarks are real. The public record right now is videos and captions. There is no Gemini 4 row on a lab scoreboard. Sergey Brin's hands-on role and the post-training comments explain why people expect a flagship. They do not grade the floatplane.

A skeptical checklist before you forward the clip

  1. Model ID from the session, not the caption. "Gemini 4 Pro" typed by the poster is not an API name.
  2. Byte-identical prompt and system message on both sides.
  3. One shot versus hidden retries. Five minutes of agent time can hide a lot of repair.
  4. Whether the page says 3.8 Flash in the URL. If it does, do not title the Slack message Gemini 4.
  5. Your actual job. Shader demos do not predict tool-use reliability.

What to do this week

Keep production on the endpoint you can name in a bill. Gemini 3.8 Flash is that model on the Google side. Opus 5.5 and GPT-6 Astra are the other two people are using as the right-hand pane.

Build the three-prompt folder above. Do not block a release on a rumor date. When a Gemini 4 model card shows up, the folder is the only way to know whether the jump in the videos survived contact with your prompts.

The clips are allowed to be impressive. A coherent floatplane and a lighthouse with a working foghorn are a real change in what one-shot visual code looks like in September 2026. The model name on the caption is the part that is still a rumor.

Related reading

  • Gemini 4 post-training: what Kavukcuoglu said
  • Floatplane leak versus Opus 5.5
  • Gemini 3.8 Flash official benchmarks and pricing
  • Five browser games with Opus 5.5
  • Opus 5.5 launch benchmarks
  • GPT-6 Sol versus Opus 5.5
  • Sergey Brin and Gemini 4
  • Primary clip: Pranav Reddy, floatplane comparison
  • Playable scene labeled Gemini 3.8 Flash: Moon Island Light Station

Kavukcuoglu's remarks are from September 24, 2026. Demo labels reflect public posts through September 27, 2026. explainx.ai did not rerun the arena prompts. Feed summaries that add a Wii remote, a horse, or a 2–3x speedup without a primary URL are marked as summaries, not measurements.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 26, 2026

Leaked Gemini 4 Pro arena tests: floatplane 3D physics vs Claude Opus 5.5

On September 26, 2026, Pranav Reddy's side-by-side video — Gemini 4 Pro in arena versus Claude Opus 5.5 on a realistic floatplane physics prompt — hit tens of thousands of views and Grok's trending summary. Google has not confirmed Gemini 4 Pro. This post unpacks the leak, the identity uncertainty (Pro vs Flash checkpoint), and why one flashy WebGL demo is not a benchmark.

Sep 25, 2026

Google DeepMind Fast-Tracks Gemini 4 After Skipping Gemini 3.5 Pro

At The Information AI Agenda Live Summit on September 24, 2026, new Google DeepMind leader Koray Kavukcuoglu said Gemini 4 entered post-training and that Google intends to ship an early post-training build as soon as possible — potentially well before end of 2026 — after Gemini 3.5 Pro never launched. Here is what post-training means, why Google skipped 3.5 Pro, and how builders should prepare API and agent routes.

Sep 16, 2026

Google Launches Gemini 3.8 Live and 3.8 Live Extended Thinking

Gemini 3.8 Live and 3.8 Live Extended Thinking are Google's newest voice models, built for near real-time dialogue with background tool calls and live progress narration. One leads a speech-quality index outright; the other trades some of that fluency for deeper multi-step reasoning.