Merged timeline of 85 items — blog publish times and listing timestamps, cut at midnight . Page 1 of 2.
On September 24, 2026, ClaudeDevs said Anthropic will again charge for requests its safeguards block before Claude responds, limited to categories with low false positive rates. The API docs spell out exactly which refusal categories are billed, and how fallback credit softens the cost if you build on the API.
Google announced three Chrome features aimed squarely at students on September 24, 2026: Gemini can now analyze podcasts and non-YouTube video, generate interactive quizzes from your open tabs and Google Docs, and sync a tab's exact scroll position across devices. Here's what shipped, what's still US/India desktop-only, and how it fits the rest of Google's AI study push.
Two Claude Opus 5.5 demos went viral within days of each other in late September 2026 — a 2:16 animated sweep through Western civilization (5.8M views) and a short "GPS, explained by Claude" video. Neither is a text-to-video diffusion model at work. Both are Claude planning a storyboard, writing JavaScript renderer code per scene, and rendering that code through headless Chrome and FFmpeg — a documented pipeline, not a new video-generation modality. Here's how it actually works, and the "slop or magic" debate the Western civilization clip set off.
F-Droid shipped 2.0 on September 24, 2026, its largest update in 10 years, just as its Keep Android Open campaign warns that Google is changing how apps get installed. Here is what the release changes, what it drops, and why AI builders shipping Android apps outside the Play Store should pay attention.
bunpav.com/play now hosts five no-download multiplayer party games — Bonk Club, Hexfall, Turbo Trolley, Splat Attack and Clang! — all built with Claude Opus 5.5 in Claude Code. Here are the real gameplay clips, what each game plays like, and how the stack actually works: ~30k lines of TypeScript, zero image assets, 122 AI-generated sound effects and a server-authoritative multiplayer backend.
Fastino Labs released GLiNER2.5-Decide on September 24, 2026: a 340M-parameter encoder that answers typed questions under user-defined rules and returns probabilities. It scored 60.1% across 17 datasets and runs on CPUs, which makes it a candidate for routing, triage and LLM-as-judge steps.
At The Information AI Agenda Live Summit on September 24, 2026, new Google DeepMind leader Koray Kavukcuoglu said Gemini 4 entered post-training and that Google intends to ship an early post-training build as soon as possible — potentially well before end of 2026 — after Gemini 3.5 Pro never launched. Here is what post-training means, why Google skipped 3.5 Pro, and how builders should prepare API and agent routes.
Google published its Project Suncatcher explainer on September 24, 2026 and Sundar Pichai confirmed the Transporter-18 ride the next day. The prototype is a hardware survival test, not a data center. The most telling detail is that the chips can only run about 15 minutes before they must cool off.
On September 25, 2026, Google Research announced a unified multi-agent framework for temporally consistent long-form video. It bundles four papers into one pipeline that tracks world state, plans globally, generates segment by segment, and critiques its own output. Here is what each part does and what you can borrow today.
Anthropic published a write-up on how it made claude.ai and the desktop app 3.1x faster on average in two weeks in August 2026, using Claude itself to find and fix bottlenecks. The headline numbers are striking, but the method is the reusable part: measure deterministically, let the agent iterate, and ratchet guardrails daily.
Menlo Research released the full reinforcement-learning training pipeline behind Asimov 1 — its $499-deposit, developer-oriented humanoid robot — covering the PPO and Adversarial Motion Priors code, the Isaac Lab simulation setup, and the sim-to-real deployment path that took the robot from zero to walking. It's a genuinely open humanoid stack, not just an open hardware BOM. Here's what's in the repo and what it's actually good for.
Meta announced GitHub alongside Notion and Box at Connect 2026, and Meta's Model API GitHub agent cookbook documents a production-shaped flow: Muse Spark via OpenCode triages issues, reviews pull requests, answers repo questions with citations, and only opens fix PRs after a maintainer applies an agent-fix label. Here is how the integration is meant to work and how it compares to coding-agent harnesses you already run.
On September 24, 2026, The Verge reported that Meta Muse users could export large parts of its virtual machine, including system files and internal docs. Meta says that is intended behavior because each user gets their own Linux VM. Both sides are partly right, and the real security question is narrower than the headline.
Microsoft announced its biggest Copilot update yet on September 25, 2026: Home merges Chat and Cowork with real, editable Office documents built in; Code lets non-developers describe an app and get one, hosted in the company's own tenant; and Autopilot is a persistent agent with its own identity that keeps working while you're away. Here's what each piece actually does, the new usage-based pricing model behind it, and what's still preview-only.
On September 24, 2026, Elon Musk wrote that SpaceX will reach "pole position" in about six months and could have a Fable/GPT-6-level model in two to three months. His argument had three parts: acceleration, diminishing returns on intelligence, and hardware. Here is how each holds up against what SpaceXAI has actually shipped.
On September 21, 2026, Odyssey shipped a playable research preview of Agora-2, a multi-agent world model that supports up to 20 humans and agents in one shared simulation with streaming pixels and explicit shared state — five times Agora-1 capacity and multiple environments. AWS demoed it live with Matt Wood. Here is how it differs from single-player world models and what robotics and agent trainers should watch.
On September 24, 2026, TestingCatalog reported references to a ChatGPT Pro Max subscription at $500 per month with fastest Work and Codex access, while OpenAI official pricing still tops out at $200 Pro and new $200 sign-ups remain paused. Here is what is confirmed, what is leak-only, and how to decide whether ultra tiers are worth it for agent workloads.
With Washington rejecting new slowdown rules and a public-private safety partnership stalled, The Information reported September 23-24, 2026 that OpenAI, Google, and Anthropic are advancing a self-governed standards body tentatively called SAFA. Here is what it would do, who might run it, and what builders should expect from voluntary audits versus law.
Peter Steinberger says OpenClaw deleted roughly 400,000 lines of AI-generated tests without much change in code coverage, using a test-audit skill that is now public in the repo. The skill is a reusable authoring gate plus an audit workflow, and its rules apply to any codebase where agents write tests.
Perplexity announced Photon on September 24, 2026, a Rust-based retrieval and ranking service that now powers its Search API. Its new Fast Search preset returns 95% of results within 230 ms and cuts the cost of agent tasks by 68% against the default preset, at a small relevance cost. Here is when to switch.
Robert O'Callahan, known for the rr record-and-replay debugger, resigned from Google DeepMind on September 24, 2026 — not from a safety team, but from a chip-design group building the next generation of faster, cheaper AI hardware. He says the pace itself is the problem, plans to keep building rr and Pernosco, and pursue "unambiguously pro-human" work like AI-assisted debugging. Here's what's verified, how it differs from September's other resignations, and what reactions split on.
Axios reported on September 24, 2026 that a political memo circulating inside the White House ecosystem frames effective altruism as a fringe movement that built the AI-doom pipeline and places Anthropic CEO Dario Amodei at its center. The document arrives as Amodei pushes pacing the frontier, Trump calls AI risk a hoax, and Anthropic faces a reported IPO window — here is what the memo claims, what is verifiable, and what it changes for builders choosing a frontier lab.
On September 24, 2026, Politico reported that the White House asked OpenAI and Anthropic to withhold new frontier models from the UK AI Security Institute until US government review completes — extending the US-first access pattern from consumer export controls into allied testing pipelines. Here is what was requested, what Anthropic already did with Mythos 5.1, and how builders should plan around split testing regimes.
On September 24, 2026, Wiz announced Scan for Good with Google DeepMind — free authorized scanning of public-facing critical infrastructure, nonprofits, and public services using Wiz Red Agent and Gemini 3.8 Flash Cyber, human validation, and CISA collaboration. Early results included hundreds of high-severity findings across rail, hospitals, archives, and open-source repos.
Anthropic moved Claude Code cloud sessions out of research preview and is handing existing Pro and Max subscribers a one-time credit, $100 or $250, to try them. Here is how to start a session, claim before October 7, what the credit covers, and the honest objections from the replies.
Meta used Connect 2026 to push Muse, its personal agent, off the phone and onto your face and keychain. Here is every announcement with prices, ship dates, what was left vague, and what it changes if you build on Muse or teach people to use AI.
The Information reports Microsoft will discount Copilot by 30% for customers with 1,000 to 10,000 seats and 50% at 10,000 or more, starting around October, as it merges chat, Cowork, Autopilot and Code into a super app and moves to seat-plus-usage billing. Here is what is confirmed, what is reported, and how to negotiate.
Anthropic's first release since calling for "pacing the frontier" claims Fable 5.1-level performance at 40% lower cost, a rewritten communication style, and the strongest safety scores of any Claude model to date. Here is every number from the announcement, plus what developers who switched from Opus 5 are actually reporting in the first hours of real usage.
Anthropic published a developer playbook the same day Opus 5.5 launched, and it contains some genuinely counter-intuitive advice — stop telling the model to "think carefully" (it always does now), hand over entire tasks instead of micromanaging steps, and when a design comes out generic, list the specific patterns you don't want rather than asking for something vaguely "not generic." Here's the full guide, condensed.
Claude Opus 5.5 beats Fable 5.1 on every benchmark Anthropic published — Terminal-Bench 4.0, GDPval-AA, Humanity's Last Exam — at a fraction of the cost. And yet the loudest developer reaction to Opus 5.5's launch was a Reddit thread titled "What's the point of Fable if Opus 5.5 is stronger in every category?" Here's the honest answer, benchmark table and all.
OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 launched within hours of each other, and the instinct is to treat them as direct rivals. The pricing tells a different story — Sol is a mid-tier, cost-optimized model at half Opus 5.5's price, not a flagship competing on raw capability. Here's what actually overlaps, and where the comparison breaks down.
Every prompting habit built around Opus 5's quirks — the hedging, the small-step supervision, the vague design requests — is now dead weight with Opus 5.5. Reading Anthropic's own playbook gets you halfway there; actually practicing the new workflow on real work tasks, with feedback, is the other half. That's what explainx.ai's live Claude for Work workshop on October 3-4, 2026 is built for.
Anthropic's own demo thread showed off a napkin-styled coding UI and a physics-accurate pencil sketch. Independent builders went further — formally verifying a production SDK, benchmarking vibe-coded Minecraft clones against three other frontier models, and building a CAPTCHA that works backwards. Here are 10 real, sourced things people built with Opus 5.5 in its first 24 hours, with links to every one.
SpaceXAI released Grok 4.7 on September 21, 2026 — a larger base model with a longer reinforcement-learning run on harder, longer-horizon tasks, served at the same $2/$6 price and 2x speed of Grok 4.6. Official evals show it leading electrical engineering and legal-work benchmarks while trailing Fable 5.1 on coding and terminal work. Here's the full table, the new safeguard stack, and where it's live today.
Before Jev, teams needing fast structured classification typically reached for XGBoost (fast, cheap, lower ceiling on accuracy) or a fine-tuned BERT model (higher accuracy, more setup, still not free-text generation). Jev sits in a genuinely different spot on that spectrum — not because typed-output classification is new, but because of how it's trained and how it reports confidence. Here's an honest comparison.
Security researchers disclosed RatHat, a China-linked Android malware family distributed via smishing and malvertising. It abuses Accessibility-service permissions to self-enable Developer Options and pair ADB for shell access outside the app sandbox, then calls a mainstream generative AI assistant to interpret the screen and navigate the device autonomously. It intercepts uninstall attempts, fakes a Play Store error, and auto-reinstalls to retain shell access.
"System One Model" entered the AI vocabulary in September 2026 when TypeSafe AI used it to describe Jev, a model that returns a choice, a score, or a probability instead of generating text. The term borrows Daniel Kahneman's System 1/System 2 psychology and is likely to outlast the specific product that coined it. Here's what it actually means as a category, and how it differs from a reasoning LLM.
Anthropic engineer Sachin Malhotra's September 14, 2026 post on the Claude blog traces how agentic coding moved the SDLC bottleneck from writing code, to reviewing it, to running CI on it — and details three failed patches before a full redesign of the test impact analysis service that decides which tests run on which pull request.
Three days after Dario Amodei's "Pace the Frontier" essay, the debate jumped from AI labs into markets and national politics. Investor Michael Burry called the pacing warnings "self-serving hype tied to IPOs," AI researcher Gary Marcus voiced similar doubt, Palantir CTO Shyam Sankar cast AI safety as ideological overreach, President Trump dismissed the whole premise as a "hoax" — first in a post, then live on a call to the All-In Summit with Jensen Huang on stage — and Kamala Harris called for Congress to pass a law slowing frontier AI down. Here's who said what, and why it now matters more than the essay itself.
A viral post claimed China opened the world's first mass-production plant making a humanoid robot every 10 minutes. That's real — it's UBTECH's new factory in Liuzhou, Guangxi — but "one every 10 minutes" describes the line's designed pace, not its planned annual output, which is a fraction of what that rate implies if sustained year-round.
Y Combinator CEO Garry Tan told CNBC and TechCrunch he wants regulators to leave AI distillation alone — and floated the idea of an "American distillation regime" letting domestic open-weight labs train on frontier models the same way Anthropic accuses Chinese labs of doing. His argument: frontier labs didn''t ask permission to scrape the internet, so they shouldn''t get to dictate what customers do with model outputs either.
On September 14, 2026, Elon Musk described Grok 4.8 as a 2.5 trillion-parameter model trained on xAI's new C++ software stack, with pretraining wrapping the same week and reinforcement learning starting immediately after. He also tempered Grok 4.7 expectations versus Opus 5.0, framed 4.9 as likely Astra/Fable-class, and echoed industry talk of a capability slowdown — here's what that means if you ship on frontier APIs.
Dario Amodei's "We Must Pace the Frontier" essay drew reactions fast — Elon Musk posted support within roughly an hour, Sam Altman committed OpenAI to match Anthropic's embedded-evaluator program, and Google DeepMind CEO Demis Hassabis called the essay's direction "correct," pointing to DeepMind's own proposal for an industry-wide AI standards body. Not everyone agreed: Chamath Palihapitiya called it a power grab that threatens open-source AI, and one reply called for Anthropic to be nationalized outright. Here's the full reaction, and what it means that industry coordination — Amodei's Step 2 — may already be starting.
Microsoft announced on September 12, 2026 that Grok models are now available as a preview option inside Copilot for Word, Excel, and PowerPoint — rolled out through Microsoft's Frontier Program, off by default, and requiring a separate admin setting. It's the clearest sign yet that Microsoft is treating Copilot as a multi-model surface rather than a Microsoft/OpenAI-only product.
A single viral tweet from OpenClaw creator Peter Steinberger — "I saw their Soul.md file and now i'm curious" — turned a Meta Muse invite code into a case study on the fastest-growing convention in AI tooling: giving an agent its identity in a plain markdown file. Here's what's actually confirmed about Soul.md, and how it fits alongside CLAUDE.md and AGENTS.md.
Joe Benton, who led a safety research team at Anthropic, and Josh Engels, an AI safety researcher at Google DeepMind, both resigned within days of each other in September 2026, telling NBC News "there are no adults in the room." They join METR. Here's what's confirmed, how it differs from the Jacob Coxon resignation days earlier, and what it does and doesn't mean for anyone building on these labs' models.
Perplexity Developers announced membership in the Rust Foundation, saying the goal is to improve how "people and agents" build with Rust. The announcement leans on SPACE, Perplexity's Rust-built sandbox runtime that already powers Perplexity Computer and the Agent API's sandbox tool. Here's what membership actually gets Perplexity, and why more AI companies are making the same move.
Anthropic's September 10, 2026 threat intelligence report disclosed that Moonshot AI and DeepSeek silently rerouted user requests to Claude and displayed its responses as their own models' output — while Alibaba ran the largest distillation attack Anthropic has ever measured, at 151 million exchanges.
For the first time, Anthropic reportedly declined to give the UK AI Security Institute pre-release testing access to a frontier model — in this case, Mythos 5.1 — breaking a pattern of voluntary pre-deployment evaluation access that UK AISI has relied on with major labs. explainx.ai covers what pre-release testing access actually involves, why a lab might restrict it, and what it signals about the evaluator-lab relationship heading into 2027.
On September 10, 2026, Perplexity published Q2D-Web (Query2Doc-Web) — a large-scale benchmark for first-stage retrieval in agentic RAG systems. Built from 23,000 PII-free production searches over nine months, it pairs 190 million web documents with 69,721 agent-reformulated queries, ten languages, and an average of 99.6 positive relevance judgments per query. explainx.ai breaks down why the benchmark exists, how it differs from MS MARCO Web, and what it means if you ship embedding models or agent search stacks.