explainx / blog / topics
AI Coding Tools
AI coding tools range from autocomplete in the editor to agents that take a ticket and come back with a pull request. The market moves weekly: new models, new pricing, new agent modes.
This page tracks that movement across Cursor, GitHub Copilot, OpenCode, Antigravity, Devin and others, with our guides on vibe coding and agentic development.
110 stories · latest Oct 9, 2026
Start here
Taste vs Judgment: A Two-Filter Guide to Picking One AI Draft
When a model produces many acceptable options at once, taste alone cannot choose between them. This guide explains the two-filter method from Addy Osmani's framing: taste cuts the pile to a few, judgment, the price of quality, picks the one to ship. Includes prompts, a decision rubric, failure modes and an interactive lab.
We Built 5 Friendslop Browser Games With Claude Opus 5.5 — Play Them Free
bunpav.com/play now hosts five no-download multiplayer party games — Bonk Club, Hexfall, Turbo Trolley, Splat Attack and Clang! — all built with Claude Opus 5.5 in Claude Code. Here are the real gameplay clips, what each game plays like, and how the stack actually works: ~30k lines of TypeScript, zero image assets, 122 AI-generated sound effects and a server-authoritative multiplayer backend.
How to Become an AI-Native Builder: What the Term Actually Means
"Become an AI-native builder" has become a common course pitch in 2026, but the phrase gets used loosely enough that it's worth pinning down what it actually means, concretely, versus what "using AI to code sometimes" means. Here's the real distinction, the skills that separate the two, and how to actually build the habit.
Meat Proxy: Don't Forward AI Output You Haven't Read
"Meat proxy" is the 2026 slang for a person who pastes model output into Slack, a PR, or a group chat without reading it. Niklas Gruhn coined it on August 3. Here is the definition, the code-review failure mode, and the difference between a relay and a colleague.
Cursor for Product Managers: Can You Build a Real AI Prototype?
Product managers, founders, and marketers can use Cursor to turn a precise brief into a working prototype without pretending the result is production software. This guide covers the build loop, review gates, project ideas, and the parts of a Claude Code-first workflow that transfer directly to Cursor.
How to Run Loops in Cursor (Agent, Cloud Agents, Automations)
Cursor already runs an inner tool loop in Agent. The built-in /loop skill repeats a prompt on an interval while your session is open. Cloud Agents and Automations are the layer that keeps going after you close the laptop. This how-to maps each surface, with copy-paste kickoffs from official docs and Cursor's own /loop skill.
Timeline
October 2026
Oct 9
AgentPlane: Run Claude Code, Codex and Cursor Side by Side, FreeAgentPlane, announced on October 8, 2026, is a local web app that runs many coding agents on your own machine, with one approval inbox, per-turn revert and optional git worktrees. We read the README. Here is how it works, and how it differs from AgentsInTheCloud.
Oct 9
ts-rust: An LLM-Written Rust Port of the TypeScript Compiler, TestedThe pingdotgg team ported the Go-based TypeScript 7 compiler to Rust almost entirely with LLMs. The README claims 181,711 ported tests pass and about half the type-check time of Go. It also says nobody read the code. We read the repo and the Hacker News thread.
Oct 7
Taste vs Judgment: A Two-Filter Guide to Picking One AI DraftOct 5
Devin Dreaming Memory and the Agent Memory Repo Open Spec, ExplainedCognition gave Devin persistent memory stored as a git repository of markdown files, plus a daily Dreaming pass that merges and prunes it. The format is an open spec that other agents can use too. Here is how it works, how it compares, and where it can go wrong.
Oct 2
Karpathy: Ask for Discardable Software Artifacts Now That Code Is AbundantOn October 2, 2026, Andrej Karpathy argued that as intelligence and code become abundant, you can ask for large, custom, discardable software artifacts, such as web apps and video explainers, that would never have made sense to build before. The shift is economic, not a new model capability.
September 2026
Sep 27
Programming Languages in the AI Era: What to Expose to AgentsJosé Valim's September 24, 2026 essay asks what programming languages should optimize for once coding agents are users. The practical answer is three surfaces: explicit types, a queryable program database, and runtime state an agent can inspect.
Sep 27
What a Prince of Persia Fan Port Shows About Coding AgentsPriyan R spent months handing Prince of Persia to frontier coding agents and only playing the result. Opus 5.5 got a level-1 screen from 8,429 differing pixels down to 2. The useful part for anyone grading agents is the oracle and the diff, not a leaderboard of model names.
Sep 26
Google Antigravity /plan: dedicated planning mode before the agent executesOn September 25–26, 2026, Google Antigravity shipped a dedicated planning mode: type /plan and the agent researches, drafts an implementation plan for your review, and waits for approval before editing code. The same flow exists on the Antigravity CLI. You can also ask for a plan in plain English for a lighter variant — explainx.ai maps how /plan fits next to /boost, Teamwork, and harness design patterns builders already use elsewhere.
Sep 25
We Built 5 Friendslop Browser Games With Claude Opus 5.5 — Play Them FreeSep 24
Cursor Rollouts and Security Reviewer: Agents That Watch Your Deploys, Plus How They Fit Cursor ProjectsCursor launched Rollouts, which writes a monitoring plan when a PR opens and verifies the change after it deploys, alongside a faster Security Reviewer. Both are on Teams and Enterprise with 10 days of included credits. Here is how they work, where they fit with Cursor Projects and cloud agents, and what to check first.
Sep 22
Linear Reworked CI for Agentic Coding: What Changed and What to StealOn September 21, 2026, Linear engineer Mufeez Amjad published how Linear rebuilt CI after agent-assisted development quadrupled its test suite but left every pull request waiting on the same gates. The post is a concrete playbook — infrastructure, critical-path jobs, setup amortization, and Vitest sharding — not just "buy faster runners." explainx.ai breaks down the numbers, the risks (including agent-written tests), and how it pairs with Anthropic's parallel story about test impact analysis at 25x CI volume.
Sep 19
How to Become an AI-Native Builder: What the Term Actually MeansSep 19
MiniMax Open-Sources Its Code CLI With a Top FrontierHarness Eval ScoreMiniMax open-sourced its Code CLI around September 18, 2026, and it scored 76.7% (23 of 30 tasks) on the FrontierHarness Eval benchmark running Kimi K3 — the highest recorded pass rate on that benchmark to date, achieved at $1.83 per pass and the fastest median solve time among tested tools.
Sep 18
Bend: A Language That Blocks AI Coding Mistakes With Math ProofsVictor Taelin, creator of HVM, released Bend on September 17, 2026: a language that compiles to native CPU and GPU code and lets an AI coding agent write LAWS.bend files — formal statements a compiler mathematically verifies can never be broken, no matter what code an agent writes afterward. It drew 299 Hacker News points, real technical pushback about under-specification, and a separate controversy over a squashed commit history that briefly overshadowed the language itself.
Sep 18
GitHub Rewrote Copilot's Runtime to Rust — One Engineer, 832K LinesGitHub published a detailed account on September 16, 2026 of migrating the runtime behind Copilot CLI, the Copilot app, and the Copilot SDK from TypeScript/Node.js to Rust — 832,378 lines of production code plus 468,689 lines of tests, built primarily by one engineer with Copilot itself handling 61% of the 1.13 million tool calls involved, across 128 merged pull requests in roughly 14.5 weeks.
Sep 16
Cognition and AWS Sign Multi-Year Deal to Deploy Devin at EnterprisesCognition and AWS announced a multi-year Strategic Collaboration Agreement on September 15, 2026, aimed at using Devin to modernize enterprise legacy systems — analyzing old codebases, recreating their behavior, and testing replacements against known outputs. Here's what the deal actually covers.
Sep 15
Cline Desktop Brings Open-Weight Coding Agents Beyond VS CodeCline has moved beyond its VS Code extension with a standalone desktop beta for macOS and Windows. The real shift is a persistent workspace for model choice, scheduled runs, imported sessions, plugins, skills, and MCP servers—not a new coding model.
Sep 14
Notch: "Programming Is a Little Bit Solved" — and Why That Worries HimMarkus "Notch" Persson, Minecraft's creator, posted that programming is "a little bit solved" thanks to AI — but that his only regret is the "mega corporation owned AI" trajectory pulling toward dystopia. The reply section split into two camps: people arguing coding is nowhere near solved, and people pointing him toward local, open-weight models as the actual answer to his complaint.
Sep 12
Devin Fusion Hits CLI: Cognition Cuts Coding Costs 39%One day after launching SWE-2, Cognition rolled its "Fusion" dual-model harness into Devin CLI, pairing a frontier planning model with a cost-effective execution model. Cognition and a third-party benchmark both put the savings in the 35-39% range — with real quality tradeoffs on judgment-heavy tasks.
Sep 11
SWE-2: Cognition's Frontier-Near Coding Model Lands in DevinOn September 10, 2026, Cognition shipped SWE-2 — post-trained from Kimi K3 at multi-trillion-parameter RL scale — scoring 50.0% on FrontierCode 1.1 Main within one point of Fable 5.1 at 64% lower cost. SWE-2 medium beats SWE-1.7 with 58% fewer turns and 81% lower cost, but long-horizon Terminal-Bench 4 still trails frontier labs by a wide margin.
Sep 11
Cursor Projects: Persistent Agents With Shared Context Across Months of WorkCursor shipped Projects on September 10, 2026: a project-scoped coordinator agent that maintains shared context files across cloud and local machines, delegates implementation to parallel subagents, and can subscribe to Slack, PRs, or schedules so work continues without re-onboarding the model every session. explainx.ai breaks down how it differs from last week's self-hosted cloud agents update — and from the MCP memory hacks teams have been using to patch the same gap.
Sep 11
Real-SWE Results: Fable 5.1 Wins, No Model Clears 40%Specific Labs co-founder janak launched Real-SWE on September 10, 2026 — a coding agent benchmark built entirely from real, private company codebases rather than public repos. The full results are now out: Fable 5.1 tops the leaderboard at 38.8%, no model clears 40%, and "missed requirements" is the single most common way every frontier model fails.
Sep 10
Marc Andreessen: Devin Writes 90% of Cognition's Production CodeOn September 9, 2026, Marc Andreessen published "Investing in Cognition" on a16z — arguing software will eat the world at compute speed now that agents write most code. Cognition says Devin produces 90%+ of its own production commits, up from 13% in a year, with enterprise proof points at Mercedes, Rivian, and Itau. explainx.ai maps the claims, the harness, and the caveats.
Sep 9
FrogNano: Microsoft's 4B Coding Agent With Zero Frontier DistillationOn September 9, 2026, Microsoft Research's FrogNano report landed with a provocative claim: a competitive 4B software-engineering agent trained purely with reinforcement learning on synthetic tasks, with zero distillation from frontier models. The mechanism is online task synthesis calibrated to what the checkpoint can just barely solve — not raw task volume.
Sep 8
Google Antigravity Gets an AlphaGenome Atlas Skill for GenomicsOn September 8, 2026, Google DeepMind launched AlphaGenome Atlas — a 1-petabyte, precomputed map of predicted molecular effects for every one of the 9 billion possible single-letter changes in the human genome. The same day, Google Antigravity shipped it as an agent skill, turning variant prioritization from a manual search into an agent-run workflow.
Sep 5
GitHub Copilot HydraFusion: Model Orchestration Over Model SelectionSatya Nadella tweeted about Project HydraFusion on September 4, 2026 — a GitHub Copilot research preview that routes coding tasks across drafting, critique, and escalation models instead of running one model end to end. Here's what the official post actually says, how the orchestration works, and why "model orchestration" is becoming the next competitive axis for agent harnesses.
Sep 5
Omen Alpha: OpenCode's New Stealth Model at $0.20 Per Million TokensOpenCode added a second anonymous stealth model, Omen Alpha, to OpenCode Go on September 4, 2026 — a 500K-context coding model with reference pricing of $0.20 per million input tokens. This is the same playbook OpenCode ran with Ox Alpha in August, right down to the unconfirmed lab speculation.
Sep 3
Cursor Cloud Agents Now Run on Your Own InfrastructureCursor shipped the ability to run cloud agents on infrastructure you manage on September 3, 2026 — your own machine pools or supported sandbox providers (AWS Lambda, Cloudflare, Coder, Daytona, E2B, Modal, Namespace, Vercel) — so agents can reach internal services and specialized hardware while Cursor still owns the orchestration.
Sep 3
Fable 5.1 Built a Minecraft Mod From Two YouTube Clips for $20A Reddit user fed Fable 5.1 two YouTube links — a Naruto lightning-dragon scene and a railgun mod short — and asked it to combine them into a Minecraft weapon. It analyzed both videos frame by frame, wrote the Fabric mod, modeled the dragon in Blender through an MCP bridge, and launched the game to test it. Total cost: $20.54.
Sep 3
Chopin: GitHub Next's Multiplayer Agent Planning PrototypeGitHub Next released Chopin on September 2, 2026 — an early, open-source prototype for getting a whole team aligned on a plan together, visually and in real time, before handing work off to coding agents. Here's what problem it's addressing and why "planning tools feel lonely" is a real complaint worth taking seriously.
Sep 1
The End of Software Engineering? Zhenfeng Cao's Agentic Paradigm PaperChinese researcher Zhenfeng Cao's June 2026 arXiv paper — resurfaced by a 395K-view X thread on August 31 — argues that LLM agents don't speed up software engineering; they replace its premise. Code stops being the product and becomes disposable tooling inside a reasoning loop. explainx.ai walks through the thesis, the benchmarks Cao cites, the EvoClaw performance cliff, and the human work that doesn't go away.
Sep 1
Google's Gemini 3.7 Flash Showcase: What Googlers Are One-ShottingGoogle's official X account spent a thread showing off what Googlers built with Gemini 3.7 Flash across Antigravity, AI Studio, and Gemini Spark — a one-shot Kerr black hole physics simulation, a motif-hunting "Art Codec" gallery, and a viral Omni video hack. Here's the honest read on a company highlight reel, and what "one-shot" actually implies for Flash-tier models.
Sep 1
Google Antigravity /boost: Deep Reasoning for Hard TasksOn September 1, 2026, Google Antigravity introduced /boost — a slash command for tasks too hard for a single fast pass. It spends more tokens on extended reasoning, routes through an orchestrator into a deep-reasoning pipeline, and runs execution-and-verification loops. Available on Antigravity 2.0 and the CLI for Pro and Ultra subscribers.
August 2026
Aug 29
Andrew Ng's AI Engineering Skills Map, Part 2: Software Engineering FundamentalsAndrew Ng's AI Engineering Skills Map named software engineering fundamentals as the second of four core skills — now he's broken it into five areas: full-stack applications, data, system architecture, security and reliability, and scaling in production. explainx.ai breaks down what each one requires and why "vibe coding without understanding" loses.
Aug 29
OpenAI Is Cutting Off Cursor: What to Do Before the November 12 ShutoffOpenAI set a November 12, 2026 shutoff for its models in Cursor after the SpaceX acquisition. Cursor CEO Michael Truell responded that OpenAI is only ~5% of traffic and that Cursor is negotiating. Anthropic, meanwhile, said it will increase compute for Claude in Cursor — the opposite move. This guide covers what GPT-5.x and Codex users should do now, and what the split response means for multi-model routing.
Aug 29
OpenCode Adds Qwen3.8-Flash 125B — A Preview of Qwen4 ArchitectureOpenCode wired Alibaba's Qwen3.8-Flash 125B into its Go backend as a preview around August 28, 2026. It is a fast, mid-size open-weight mixture-of-experts model that doubles as the first hands-on look at Qwen4 architecture — but the preview label, provider-reported benchmarks, and "architecture hint, not a release" framing all matter before you route real work to it.
Aug 28
Google Antigravity Teamwork: Multi-Agent Framework for Long-Horizon WorkGoogle Research and DeepMind used Antigravity's Teamwork multi-agent framework on frontier math, theoretical CS, and systems engineering work. Agents propose, stress-test, and build over hours or days via five patterns — powerful, token-heavy, and explicitly not for everyday tasks.
Aug 27
Hackers Talked Cursor's AI Agent Into Breaching 7 CompaniesReuters reported on August 27, 2026, that a Russian-speaking group called Aur0ra used Cursor's built-in AI coding agent to breach seven companies — by convincing the agent, nearly every time it initially refused, that the attack was an authorized security test. The model reportedly running the agent was Anthropic's Claude Sonnet 4.5. Update (Aug 31, 2026): CloudSEK and Gambit Security confirmed the same operator deployed actual Aurora ransomware — a Zig-coded Windows/Linux/ESXi encryptor — after using Cursor's agent for hands-on exploitation against 10 of 20+ victim organizations, marking one of the first documented cases of an agentic coding tool used as attack infrastructure rather than just text generation.
Aug 27
Antigravity Now Builds Interactive Generative UI Artifacts Inside Your IDEGoogle Antigravity can now render interactive HTML/CSS/JS artifacts inline and in its artifacts panel, driven by a /generative_ui slash command and exportable to a standalone HTML file. It landed in version 2.11.0 alongside Chart.js, Plotly, and KaTeX support. Here's what it actually costs, how it compares to Claude Artifacts and Mermaid, and the security question Google's docs do not yet answer.
Aug 27
The New Twitter Launched — and the Loudest Reply Was About Its HomepageOperation Bluebird's twitter.new went live, claiming Musk's X abandoned the Twitter name. The highest-engagement reply to the news wasn't about the lawsuit — it was "the most vibe coded piece of junk homepage I've ever seen." explainx.ai fetched the real site and checked the accusation against the actual copy and markup.
Aug 25
Lars Faye: AI Coding Will Prevent Expertise — What the Studies SayAug 24
AGENTS.md for Code Quality: What Belongs in the File vs CIAug 24
Meat Proxy: Don't Forward AI Output You Haven't ReadAug 23
Microsoft Says Coding Is "Worth It Now More Than Ever" — X DisagreedAug 22
He Couldn't Upload Music to X, So He Vibe-Coded a Music AppAug 22
Cursor for Product Managers: Can You Build a Real AI Prototype?Aug 22
Google Antigravity Remote Control: Browser-Based Agent Sessions ExplainedAug 22
Someone Played an AI Video Inside Google Sheets — Here's HowAug 20
Cursor Auto Pricing: Per-Model Rates and Higher Limits on August 24Aug 20
Cursor Ships Event-Driven Cloud Agents and Isolated VMs for AI Coding SwarmsAug 20
How to Run Loops in Cursor (Agent, Cloud Agents, Automations)Aug 20
Is Software Dying or Changing? Why Teams Rebuild InternallyAug 19
Cursor Built Its Own Git Host — Here's How Origin Actually WorksAug 18
Cursor Origin: New Code Hosting Platform Launches Beside GitHub OutageAug 17
GitHub Copilot Canvases: Making Agentic Workflow State VisibleAug 16
GitHub Copilot Adds Grok 4.6 Across CLI, IDE, and CloudAug 14
SpaceX Closes $60B Cursor Acquisition — Cursor Joins the Grok TeamAug 8
Databricks on Managing AI Coding Costs at Scale: 4 Cost LeversAug 5
Cursor Adds Google Workspace Plugins: Gmail, Drive, and Calendar in the IDEAug 5
Cursor Open-Sources Mixture-of-Kittens: An MoE Megakernel for NVL72sAug 5
Eight Myths About AI and Software Engineering, Backed by DataAug 4
Should You Manually Retype LLM-Generated Code? The HN DebateAug 2
Cursor Gave FFmpeg Developers Free Credits — Why It MattersAug 2
Cursor Restores Dollar Costs After Usage-Page BacklashAug 1
Software for One: Building Personal Apps with AI Coding Agents
July 2026
Jul 31
2x, Not 10x: What Coding With LLMs Actually Delivers in 2026Jul 29
Cursor Start India: ₹649 Plan with Grok 4.5 and ComposerJul 27
AI Burnout Is Real — Focus and FollowthroughJul 27
scriptc: Vercel Labs Compiles TypeScript to Native Binaries — No Node, No V8Jul 26
AI Coding Agent Evals: How They Score on Real RepositoriesJul 24
How Do We Stop Vibe Coding? Trust Beyond Prompt RitualsJul 23
The Book Prize Index: A Vibe-Coded Semantic Search Tool for Award-Winning NonfictionJul 23
Cognition Acquires Poke Maker Interaction — Devin Meets Texting AgentsJul 23
Cursor Router: Automatic Model Selection for Teams and EnterpriseJul 23
Uncle Bob Doesn’t Review AI Code. He Builds a Gauntlet InsteadJul 22
Codeberg Bans Vibe-Coded Projects: What the New ToU Actually SaysJul 22
Cursor Agent Swarms: SQLite in Rust, Planner/Worker EconomicsJul 22
Cursor Doubled Usage Limits Again? July 21 Clarifies the 2× PoolJul 20
The Zen of Parallel Programming: Why More CPUs Do Not Fix Human SyncJul 19
“Programming Is Low-Intelligence Work” — kache’s X Debate Explained (July 2026)Jul 17
Recursive Model Improvement — Lee Robinson's AI Engineer Talk (Cursor, SpaceXAI)Jul 15
OpenCode Desktop Tabs — Multi-Session Layout, Worktrees Gap, and TUI Escape HatchJul 13
Destructive Command Guard: Stop AI Agents Before They Wreck Your RepoJul 13
Should Developers Stop Reading AI-Generated Code? The Review DebateJul 9
SWE-1.7: Cognition's Frontier Coding Model at 1000 tok/s on DevinJul 5
Programmer Mental Health After AI: Flow State, Context Switching, and Finding PeaceJul 3
Kimi K2.7 Code in GitHub Copilot: First Open-Weight ModelJul 1
How to Run Open Source Models Locally and Wire Them Into OpenCode (2026)
June 2026
Jun 30
Run AI Coding Agents From Your Phone: pocketdev, Cursor iOS, and OpenClawJun 29
Cursor for iOS Launches: Cloud Agents on Your Phone — Ben Lang's Big Day (June 29)Jun 29
Is English really the hottest programming language? Karpathy's 2023 tweet, three years laterJun 27
Vibe Coding for Product Managers: Build Prototypes Without Coding (2026 Guide)Jun 27
The Biggest Vibe Coding Nightmares (And How to Avoid Them)Jun 27
What is Cursor? How to Build Your First HTML Project with AI (2026 Guide)Jun 27
What is Vibe Coding? The Complete Guide to AI-Assisted Development (2026)Jun 25
Ornith-1.0: Self-Scaffolding Open Models for Agentic CodingJun 24
Notion Meets Cursor: Assign Bugs and Features to Cloud Agents From Your Task BoardJun 23
Impeccable + GitHub Copilot: AI Design Quality Built InJun 19
Claude Code $20 vs Codex vs Gemini CLI vs GLM-5.2: Which Coding Agent Plan Is Best in 2026?Jun 19
What Is Kilo Code? Open-Source AI Agent for VS Code, JetBrains, and CLIJun 19
OpenCode: The Open Source AI Coding Agent for Terminal, Desktop, and IDE (2026)Jun 17
Cursor Origin: the agent-first git hosting platform that wants to replace GitHub (2026)Jun 16
SpaceX Is Acquiring Cursor for $60 Billion — The SEC Filing Explains EverythingJun 11
Cursor CLI Slash Commands: Complete Reference (2026)Jun 11
OpenCode Slash Commands: Complete TUI Reference (2026)Jun 8
Anthropic Engineer: Stop Prompting Claude, Build Loops That Prompt Themselves (Harness Engineering Explained)Jun 7
Y Combinator Launches Paxel: AI Coding Habits Profiler for Builder Reports and Startup School Applications