SWE-Game is a new benchmark that asks coding agents to build, finish, repair and port real video games, and the headline is humbling: across six models, the best overall score on the three construction tasks stays below 60 out of 100. Opus5 leads every one of the five task types, but even it reaches only 50.38 on Brief-to-Game, where the agent starts from a short description. The paper, "SWE-Game: Can Coding Agents Build the Games We Want?" (arXiv 2609.33678), was submitted on September 27, 2026, revised on October 7, and is trending on Hugging Face Papers.

SWE-Game at a glance
| Item | Detail (from the paper) |
|---|---|
| Tasks | 247 |
| Reference games | 41 executable Godot games (24 in 2D, 17 in 3D) |
| Gameplay categories | 13 |
| Task types | Brief-to-Game, GDD-to-Game, skeleton completion, fault repair (83 cases), Godot-to-Unity porting |
| Models evaluated | Six |
| Top model | Opus5, best overall in all five task types |
| Best construction score | Below 60 out of 100; Brief-to-Game best is 50.38 |
| Executable-check accuracy | 92.59 percent balanced accuracy vs human labels |
| Video VLM judge accuracy | 78.41 percent on the same labels |
| Visual rubric vs humans | Spearman 0.829 on 200 gameplay clips |
Why a game benchmark at all?
Most coding benchmarks ask whether a patch makes a test suite pass. Games are different: the right answer is an experience. A platformer can compile, show a character on screen and still be broken because the spike does not hurt, the jump arc is wrong, or the level cannot be finished. The authors argue that this gap between "code that runs" and "a game that plays" is exactly what current software-engineering benchmarks miss. If you have followed the debate over SWE-style scores, for example how reward hacking and contamination can inflate SWE-bench numbers, you will see why a harder-to-fake, behavior-based test is attractive.
The paper positions itself against earlier game-development evaluations such as GameDevBench, OpenGame, WebGameBench and V-GameGym, which it says largely rely on multimodal models to judge the output. SWE-Game instead builds an evaluator that can drive a game and observe what happens.
The five tasks
Each task starts from reference materials that describe the intended gameplay: videos, assets, requirements and project code, depending on the task.
- Development from a brief. The agent gets a short description and must build a playable game. This is the most open-ended task and, in the abstract, the one reported at 50.38 for the best model.
- Implementation from a game design document. A fuller specification, closer to what a studio hands to a developer.
- Skeleton completion. A partly built project with missing pieces to fill in.
- Repair. 83 cases where faults were injected into a working game. Scoring looks at whether the broken behavior is restored and whether the rest is preserved.
- Godot-to-Unity porting. Move a game across engines while keeping the behavior.
The first three are the construction tasks, the ones capped below 60. The paper says repair and porting are scored too, with Opus5 on top there as well, but the abstract gives the sub-60 ceiling only for construction.
How a game gets graded
This is the most interesting design choice. Agent-built games are implemented independently, so the evaluator cannot assume anything about their internals. SWE-Game uses a shared instrumentation interface: submissions provide bindings that map their own objects to standard roles, actions and observable values such as health, progress and score. The evaluator then owns the drivers and probes. In the paper's example, a spike-contact test checks that touching a hazard is followed by the required health change or death transition. The submission supplies only bindings; the checks and expected outcomes belong to the evaluator, so an agent cannot quietly write easier tests for itself.
Evaluation then combines three kinds of evidence:
- Engine-state checks on what the running game actually does.
- Certified reference-input replay, where a validated sequence of inputs is replayed against the agent's game to see if the same mechanics hold.
- Agent-authored feature demonstrations, where the agent shows that a feature works.
Separately, game-specific vision-language rubrics score presentation: how the game looks, not whether it functions. Splitting function from looks is sensible, because a polished game with broken physics and a bare-bones game with perfect mechanics are different failures.
What went wrong most often
The authors reviewed submissions and say the predominant problems were requirement omissions and gameplay logic errors. In plain language: agents often skip things the spec asked for, and when they implement a mechanic they get its rules subtly wrong. That matches what many developers see when an agent builds a feature that looks right in a screenshot but fails once you play it. It also lines up with reliability work elsewhere, such as the Microsoft ThinkingBox study of database-state failures in agents, where surface success hid wrong underlying state.
The judge result matters as much as the leaderboard
The second finding may be more reusable than the model scores. On human-labeled behaviors from 100 agent-built games, the executable checks reached 92.59 percent balanced accuracy, while a video-based VLM judge reached 78.41 percent. Rubric-based visual scores, though, tracked human taste well: a Spearman correlation of 0.829 with human ratings of 200 gameplay clips.
The practical reading is that you should not grade game behavior by asking a model to watch a video, but you can use a model to grade how a game looks if you give it a game-specific rubric. The authors conclude that runtime evidence plus visual assessment is the sensible combination. Anyone building their own agent evaluation, in games or in any interactive app, can borrow that pattern. Our complete guide to AI benchmarks covers why the judge itself needs validating.
How to read the Opus5 result
Opus5 is the best of six on every task type. That fits other recent results where the same model family leads, such as the TasteVal research-taste benchmark and the Epoch Capabilities Index ranking. Caveats apply. The abstract does not list the other five models, and the paper's "Opus5" label should not be assumed to map to any specific product tier without checking the full text. The study was run by the benchmark authors, with their own harness, on Godot-centered reference games, and we found no independent replication. Rankings of other models "vary across development activities", the authors say, which is a polite way of saying the leaderboard is not one-dimensional.
For model choice in general, see how current models compare in our Opus 5.5 versus Sonnet 5.5 comparison and the small-model lineup that includes Haiku 5.5.
What this means if you build with agents
- Do not expect one-shot games. A sub-60 ceiling on construction means a human still needs to review mechanics, not just visuals.
- Give agents a spec, not just a sentence. The paper reports a Brief-to-Game best of 50.38 and the best of the design-document task is separate; richer inputs are the point of having several task types. We have not seen the per-task numbers beyond the abstract, so check the paper's tables.
- Make behavior testable. The instrumentation idea, exposing health, score and state through a standard interface, is what lets a test fail for the right reason. If you want agents to write games or any stateful app, build probes first.
- Repair is its own skill. Injected-fault repair measures restoring and preserving behavior. That is closer to day-to-day maintenance than greenfield generation, and worth testing on your own codebase. See our overview of agent harnesses for how the scaffold around the model affects scores.
- Context for the game industry. AI game creation is already a funded category, for example Astrocade's $56M round. A benchmark like this gives a way to check the claims.
Limits and open questions
- The abstract does not say whether the dataset and evaluator are public; check the paper and its linked resources before relying on them. The Hugging Face page lists one dataset and one Space citing the paper.
- Six models is a small field, and the model naming in the abstract is terse.
- Sub-60 scores depend on how the rubric weights each requirement; a different weighting could shift conclusions.
- The 92.59 and 78.41 percent figures come from 100 games labeled by humans; the sample is modest.
Bottom line
SWE-Game is a useful reality check: agents can already repair and port games reasonably, but building a game that matches what you asked for is still hit-or-miss, with the best result on construction under 60 out of 100. The most transferable lesson is methodological: grade behavior with executable probes, and use vision models for looks, not for truth.
Related reading
- Cursor, reward hacking and SWE-bench contamination
- Microsoft ThinkingBox agent benchmark
- Agent Lightning and SWE-bench
- Gemini 4 Argon launch and benchmarks
- AI benchmarks: a complete guide
Primary: Chen et al., "SWE-Game: Can Coding Agents Build the Games We Want?", arXiv 2609.33678, and its Hugging Face papers page.
Details are accurate as of October 9, 2026 and are based on the paper's abstract and the portion of the HTML version we could read, not a full reading of every table.
