Merged timeline of 52 items — blog publish times and listing timestamps, cut at midnight . Page 1 of 2.
CREEM 2.0 empowers users to effectively sell and expand their AI-driven products, enhancing market reach.
Bitrise Remote Dev Environments provide cloud-based Mac environments for developers to build applications seamlessly.
Modaal for Android enables developers to ship applications on both iOS and Android from a single project.
Text Agent Store is a marketplace where users can interact with AI agents via text messaging.
NovaSynth by Noveum allows users to test voice agents in real-world scenarios for improved performance.
On September 17, 2026, Anthropic opened public beta applications for its Life Sciences Verification Program, giving vetted biotechs and academic labs access to models including Mythos under new safeguards. Alongside it, Anthropic published results showing Claude accelerated more than 30 open-source biomolecular models by roughly 4x in under four weeks, and it's co-sponsoring a $1M protein design competition with Adaptyv Bio.
Anthropic published measurements on September 17, 2026 meant to make the pace of AI development legible to outsiders — how much of its own AI R&D Claude now performs, how tightly ~30,000 internal agents are monitored, and what share of compute goes to safety work. The headline number: Claude "leads" 26% of Anthropic's model R&D tasks, up from under 1% in February 2026.
Victor Taelin, creator of HVM, released Bend on September 17, 2026: a language that compiles to native CPU and GPU code and lets an AI coding agent write LAWS.bend files — formal statements a compiler mathematically verifies can never be broken, no matter what code an agent writes afterward. It drew 299 Hacker News points, real technical pushback about under-specification, and a separate controversy over a squashed commit history that briefly overshadowed the language itself.
Yoshua Bengio told AFP on September 16, 2026 that humanity is "losing control" of AI and needs guardrails "similar to nuclear arms controls." He's one name in an unusually cross-ideological coalition — Geoffrey Hinton, Steve Wozniak, and Richard Branson alongside Steve Bannon and Glenn Beck — publicly backing the Sanders-Casar bill to permanently ban superintelligent AI. Here's what Bengio actually said, versus the vaguer "recursive self-improvement should be illegal" framing circulating online.
Anthropic rolled out redesigned Projects in Claude Code on September 17, 2026 — one long-running conversation that acts as a coordinator, starting a cloud-session thread for each piece of work, passing shared memory between them, and reporting back what needs your attention. Here's what a project actually is, how it's organized, and when it beats a single cloud session.
Anthropic quietly shipped a governance feature worth an admin's attention: members of a Claude Team or Enterprise organization can now publish a skill or plugin they've built straight into an org-wide library everyone on the team can use — and for Team organizations, the default policy ships as Open, meaning no review step, unless an admin changes it first.
Exa introduced Snapshot on September 18, 2026: a search API that returns results from the web as it existed on a date you specify, backed by more than 400 billion historical webpage snapshots spanning two decades. The pitch isn't browsing old pages for nostalgia — it's preventing web leakage in RL training and reproducible evals, and enabling point-in-time backtesting that previously took months of manual data collection.
On September 18, 2026, Figure released Helix 2.5 and rented 30 Bay Area homes to prove a specific, falsifiable claim: a single foundation model, pretrained on human video, can walk into a home it has never seen and tidy a room, fold towels, or make a bed with no fine-tuning in that home. The headline ablation is the real story — Index pretraining alone took zero-shot success from 9% to 56%.
GitHub published a detailed account on September 16, 2026 of migrating the runtime behind Copilot CLI, the Copilot app, and the Copilot SDK from TypeScript/Node.js to Rust — 832,378 lines of production code plus 468,689 lines of tests, built primarily by one engineer with Copilot itself handling 61% of the 1.13 million tool calls involved, across 128 merged pull requests in roughly 14.5 weeks.
Google updated its Gemini API managed agents with a new harness, antigravity-preview-09-2026, bringing the Antigravity coding agent's tools and behavior into AI Studio and the Interactions API on Gemini 3.8 Flash. Two new APIs ship alongside it: Files, for moving data in and out of an agent's sandbox, and Credentials, which lets an agent call GitHub or Slack without the model ever seeing the actual token.
Google Labs introduced CC on September 18, 2026 — an AI agent built around household logistics rather than individual productivity: a shared "Your Day Ahead" morning brief, autosynced Google Calendar and Tasks across up to five family members, meal-plan drafting in Google Chat, and delegated paperwork like school permission slips. It's US-only, 18+, and waitlist-gated.
Google Research introduced a generative UI system built specifically for classroom use on September 18, 2026 — teachers describe a topic and get a guided, interactive simulation back, constrained by learning-design guardrails rather than an open-ended UI generator. A 30+ item STEM sample library ships alongside it, and a developer publicly noted he'd built a similar tool months earlier with off-the-shelf LLMs and custom UI tooling.
Two separate builders reported GPT-6 Astra decoding historical ciphers that had sat unsolved for decades — an 82-character 1941 German Army Enigma message (MVUEH) and a 1918 WWI German naval radio transmission from a public list of 50 unsolved ciphers. One result got direct sign-off from a working Enigma historian; the other has an honest, unresolved question about why a message using an already-known key sat unsolved for so long. Here's what actually happened, and what's still unverified.
The Institute of Foundation Models introduced Uno on September 17, 2026: a small, cheap-to-train diffusion adapter bolted onto an existing autoregressive LLM's weights that generates multiple tokens per step while provably reproducing the exact same output distribution as one-token-at-a- time decoding. On K2-Horizon-7B, Uno delivers up to 2.2x higher throughput at batch size 1, with no separately trained draft model required.
Researchers from Cambridge published "Infinite-Parameter LLMs" on arXiv on September 16, 2026 — a proposal to stop storing live, run-time context (facts a user supplies, corrections they give) in the prompt, where it's re-read every request and discarded when the session ends, and instead compile it directly into a model's weights via a compact hypernetwork updated online as a session proceeds. Here's the actual mechanism, what it would change if it works, and why it's a research proposal, not a product.
Mark Zuckerberg announced Muse for Mac on September 18, 2026 — extending Meta's personal agent, previously mobile, web, and WhatsApp only, onto the desktop where it can act across apps, files, calendar, notes, and messages on your computer. Here's what's actually different from the phone version, and why "you control what it can access" is the sentence worth reading twice given Muse's existing Sentinel VM security model.
Meta's Muse personal agent reached #1 on the US App Store's Top Free Apps chart roughly one week after launch — ahead of ChatGPT at #2, per a screenshot circulating alongside Alexandr Wang's own announcement — with a 4.9-star rating across more than 13,000 reviews. The milestone landed the same week Meta shipped Muse for Mac, though the chart position and rating reflect the original mobile app's first week, not the desktop client specifically.
OpenAI introduced Astra for Law on September 17, 2026: GPT-6 Astra paired with a legal search index spanning more than 230 million U.S. legal sources, plus instructions tuned for legal analysis and writing. Early access goes to select firms through Trusted Access in ChatGPT and Codex, with partners like Harvey and Sullivan & Cromwell already building on it. Here's what changed, the benchmark behind the 40% claim, and what a vendor-run comparison against Claude does and doesn't prove.
PrismML released Ternary Bonsai 2 27B on September 17, 2026: a 1.76-bit compression of Qwen3.8 27B that fits in 5.9GB while retaining 98.2% of the full-precision model's aggregate benchmark score — up from 95% in the first Bonsai release two months earlier. It runs coding agents locally at 143 tokens/second on an RTX 5090, but Hacker News users report inconsistent real-world throughput and at least one clear reasoning failure in testing.
Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026 — an omnimodal model that moves past describing audio and video toward acting on them: editing footage, translating dubbed dialogue while preserving voice, and building deep-research reports from a video's content. Audio input pricing drops more than 98%, and Agentic Understanding cuts token consumption by roughly 46% versus processing a whole video statically.
Security researcher s1r1us and team disclosed a nine-step exploit chain that took over OpenAI employee ChatGPT and Codex accounts, reaching connected Slack, GitHub, and email access — all found and responsibly disclosed in under 72 hours. The most striking detail: Claude Opus 4.8 found the underlying libheif vulnerability, and Opus 5, released mid- investigation, built a working exploit from scratch in about three hours.
Claude Code, pi, and Hermes look different on the surface but solve the same ten underlying problems. These are the concepts that separate a demo that falls over after ten turns from an agent you can trust to run unattended — each with a concrete example and a way to build it yourself.
GPT-6 Astra's launch-week coverage produced dozens of demos, but most builders don't need a maze-solving CAPTCHA gauntlet — they need to know what's actually worth building with it today. These are ten concrete, buildable project ideas GPT-6 Astra is well-suited for, each grounded in a real demo or benchmark explainx.ai has already covered.
Claude Code, pi, and Hermes all call the same model APIs. What separates a working coding agent from a demo that falls over after ten turns is everything wrapped around the model: the agent loop, the tool contracts, the context and memory system, and the recovery logic that keeps a session alive across failures. That layer now has a name — harness engineering. Here's what it actually covers.
Brett Adcock posted eleven words and 444,800 people looked. The post itself says nothing, but three verifiable things published in the weeks before it narrow the space of what Figure can plausibly be showing, and they all point in the same direction.
NVIDIA HPC Developer announced CUDA Rust on September 16, 2026 — two paths, cuda-oxide for SIMT kernels compiled to PTX and cutile-rs for tile-based programming on stable Rust, both designed to catch aliasing errors at compile time that CUDA C++ leaves to runtime debugging.
A model calling itself "Union Alpha," with no publicly confirmed creator, posted a 74% score on the DeepSWE software-engineering benchmark this week — enough to edge out GPT-5.6 Sol. It's the latest in a recurring 2026 pattern of anonymously-branded stealth models appearing on public leaderboards before their developer is revealed.
AgentBeam has moved from a pure observability tool to an active local enforcement layer: one npm install, one beam setup command, native hooks for six agent clients, and a dashboard for org-wide policy. This guide covers the new setup flow, what it protects, and how it differs from the earlier observation-only Beam CLI.
Most people use Claude like a search engine — type a question, read the answer, close the tab. Claude Projects changes this entirely. Here is how to set one up for your specific role, with real examples for marketing, sales, HR, operations, and product management.
Google Antigravity CLI (agy) introduces advanced terminal paradigms: nsjail/sandbox-exec containment, directory plugins, parallel subagents, and a rich slash command set. Here is how they work.
TimesFM 2.5 packs a 16k context window and continuous quantile forecasting into 200M parameters. Available on PyPI, Hugging Face, BigQuery ML, and Google Sheets. Here's how it works.
Unreal Engine 5.8 (June 17, 2026) connects LLM agents to the Editor via MCP. Grummz showed Claude and Codex beside the engine controlling Blueprints, PCG, and lighting—Epic's first-party Toolset plus a growing plugin ecosystem.
A plump cartoon cat with 100 trillion parameters and a score that demolished Fable 5 on something called VoltaireBench. Le Chaton Fat was never real — but for 72 hours in June 2026, a sizeable chunk of AI Twitter wasn't sure. Here is how the hoax started, why it spread, and what it reveals about the state of AI benchmark hype.
At Google I/O 2026, two tiny bipedal robot ducks showcased Gemma 4 E2B running fully on-device—one on a Raspberry Pi 5, one on a Jetson Orin Nano—using multimodal inputs to see, hear, and speak in real time.
Research shows 26.1% of agent skills contain vulnerabilities and 5.2% show likely malicious intent. NVIDIA's SkillSpector is an open-source scanner that catches them before they reach your agent.
OKF formalizes the LLM wiki pattern: folders of markdown files with YAML frontmatter, linked into a knowledge graph. No SDK required—just files in git. Google ships BigQuery enrichment agent, HTML visualizer, and three sample bundles.
Don't re-discover knowledge every query—let the LLM compile and maintain a wiki. Karpathy's gist defines raw sources, an LLM-owned wiki layer, and a CLAUDE.md schema. Here is the full pattern, when to use it vs RAG, and 20+ implementations.
Prompt engineering optimizes a single instruction you type by hand. Loop engineering optimizes the autonomous system that decides what to prompt, when to prompt it, and whether the result is acceptable. Here's what it means and why it matters.
Claude Code has two billing paths: a subscription plan (Pro at $20/month, Max at $100–$200/month) or bring-your-own API key with pay-per-token rates. This guide breaks down every option, compares them against Cursor and Copilot, and shows you exactly when each makes sense.
Settings passed on the CLI vanish when the session ends. settings.json makes them permanent. This reference covers every key you can put in .claude/settings.json—from model selection and permission globs to MCP server definitions, PostToolUse hooks, and custom agents—plus a clear project vs. user vs. local split so your team shares the right config.
Claude Code sessions don't have to start from scratch every time. With --continue and --resume, you can instantly reload your last conversation or pick any past session from an interactive list—preserving all the context you built up.
Single prompts are dead for serious software development. Anthropic engineers run iterative loops where Claude observes, plans, acts, and reflects over hours or days, shipping 8x more code with 80%+ authored by AI by May 2026.
From SKILL.md to CLAUDE.md, a comprehensive guide to every type of markdown file used to configure, instruct, and extend AI agents in 2026. Includes file structure, best practices, and real-world examples.
DESIGN.md is the visual source of truth for AI agents. This guide ranks the top 10 directories for finding and generating agent-native design blueprints.
Codex pets look whimsical; operationally they are a status surface for long agent runs. This guide goes settings-deep: Appearance & Pets, composer commands, hatch-pet packaging, art direction, and how to choose top built-in vs custom mascots without drowning in sprite tech debt.