explainx / blog / topics
Open-Weight Models
Open-weight models publish their weights so anyone can run, fine-tune, and study them. Releases from DeepSeek, Qwen, Zhipu GLM, Moonshot Kimi, Mistral and others often close the gap with closed models within months.
This page tracks each release we covered, with benchmark results and practical notes on running them.
195 stories · latest Oct 9, 2026
Start here
Small Models Compared: Haiku 5.5 vs GPT-6 Luna vs DeepSeek vs Gemini Flash vs GLM
Claude Haiku 5.5 landed at $0.10 per million input tokens, the same sticker as GPT-6 Luna. Does that make the small-model tier a tie? We compare five cheap models on list price, cache pricing, context, licensing and the benchmarks that can honestly be compared, and say which to use for subagents, coding, long context and self-hosting.
Hugging Face Puts RL Environments on the Hub: What OpenEnv Means for Agent Training
On October 5, 2026, Hugging Face said RL environments can be shared and loaded on the Hub like datasets, bundling tasks, tests, containers and reward functions, to end framework fragmentation. This guide explains what an RL environment is, what OpenEnv standardizes, who is behind it, and how a builder can try it.
DeepSeek DSec: The Sandbox Platform Running 380,000 Concurrent Agent Environments, and Why Agent Execution Is "Untrustworthy"
DeepSeek's DSec paper describes the sandbox layer behind agent training: function calls, containers, microVMs and full VMs under one API, 380,000 concurrent sandboxes and over 5,000 creations per second. Its most useful line is an admission: agent execution is untrustworthy, and no single mechanism prevents all misbehavior. Here is the design and the lessons.
A Developer Put All 3.1M arXiv Papers Into One 16TB Hugging Face Dataset
Instead of scraping arXiv paper-by-paper, builders can reportedly pull the whole preprint archive — 3.1 million papers, 16TB — from one Hugging Face dataset. explainx.ai walks through what's plausible about the claim, how to stream a dataset this size without downloading it whole, and the licensing and deduplication traps that come with any single-source scientific corpus.
Calacanis vs Musk: Is the Open–Frontier Gap Already Negligible?
After a week of cheap capable open releases, Calacanis called the open–frontier gap negligible. Musk replied it is a world of difference. The useful answer is task-conditional — and it reshapes how you route agents.
What Are LLM Parameters? Top 10 Model Sizes (July 2026)
Parameters measure how many learned numbers sit in a model checkpoint — not tokens, not context length. explainx.ai explains total vs active MoE counts, why closed frontiers hide size, and ranks the top disclosed LLM sizes as of July 2026, led by Kimi K3 at 2.8 trillion.
Timeline
October 2026
Oct 9
Ecosia Drops Mistral for Open-Weight Models: What Politico Reported and What It MeansPolitico reported on October 6, 2026 that Ecosia is dropping Mistral as its AI model supplier and moving to open-weight models, some of them Chinese, hosted by Melious. The headline many people saw was a new Ecosia and Mistral partnership. The reporting says the opposite. This post separates the facts from the framing.
Oct 9
LightOnOCR-3: Apache 2.0 OCR in 0.8B, 1B and 4B, Ranked on ParseBenchLightOn released LightOnOCR-3 on October 8, 2026. Three Apache 2.0 sizes read pages, return labeled bounding boxes, describe images and turn charts into tables. The company says the 4B and 0.8B lead open models on ParseBench. The public leaderboard tells a more mixed story.
Oct 9
Sakana Namazu Powers Aillis Evidence Finder: A Japanese LLM Scores 96.4% on the Physician ExamOn October 9, 2026, Sakana AI announced that its Japanese-focused LLM Sakana Namazu was adopted by Evidence Finder, a literature search tool for physicians from Aillis Inc. The Namazu-based evaluation model scored 96.4% on the 120th national physician exam, which Aillis says is the highest published score among GENIAC-designated domestic models. Sakana itself says an exam score is only one facet of usefulness.
Oct 9
Tencent Youtu-Parsing-Omni: A 5B Open Model That Parses Documents, Audio and VideoTencent has published the weights of Youtu-Parsing-Omni, a compact 5B model that reads document pages, charts, geometry figures, audio and video and answers in one structured JSON envelope. explainx.ai walks through the benchmark table, how to run it, and the license clause to check first.
Oct 8
China Has 24 GW of Data Center Capacity vs 56 GW in the US: What SemiAnalysis FoundSemiAnalysis, as relayed by the Financial Times, says China has more than 24 GW of operating data center capacity with roughly 50 GW more planned or announced, against about 56 GW in the US by the end of 2026. Gigawatts are not the same as compute, so here is how to read the figures.
Oct 8
Hugging Face ML Intern: How One Developer Built Six Custom Models for About $103 (and the Prompt Playbook Behind It)On October 8, 2026, Hugging Face published a hands-on account of ML Intern, an agent in HuggingChat that plans, trains, evaluates and publishes models. Six projects, about USD 103 in compute, and a repeatable prompt structure. Here is what was built, what it cost, and how to copy the approach.
Oct 8
Netflix Instadocs: AI Gone Wild vs the Hugging Face RecordNetflix will release Instadocs: AI Gone Wild on October 12, 2026, a documentary on the July breach of Hugging Face by OpenAI's own agents. Before it lands, here is what the published record says, which numbers in the promotion differ from it, and what builders should take from the story regardless of the film.
Oct 8
Small Models Compared: Haiku 5.5 vs GPT-6 Luna vs DeepSeek vs Gemini Flash vs GLMOct 7
GLM 5.3 Lands on Amazon Bedrock: What Enterprise Teams GetAWS announced on October 5, 2026 that Z.ai's GLM 5.3 is on Amazon Bedrock, a 753B mixture-of-experts coding model with a 1M-token context window. This post covers the model IDs, the OpenAI-compatible call, caching and service tiers, and the caveats about eligibility and vendor-reported benchmarks.
Oct 6
Hugging Face Puts RL Environments on the Hub: What OpenEnv Means for Agent TrainingOct 6
Interfaze-1-Lite: An Apache 2.0 Model for OCR, Speech and Structured ExtractionInterfaze, a Y Combinator company, released interfaze-1-lite on October 5, 2026: an open-weight model aimed at deterministic backend jobs such as document OCR, speech-to-text, classification and extraction, returning confidence scores and bounding boxes. It runs on one 80 GB GPU. Here is what it is, the numbers, and how to test it against your current pipeline.
Oct 6
Mistral Large 4 "Le Chonk": 1T Parameters, 49B Active, Open Weights Due October 27Mistral released a preview of Large 4, nicknamed Le Chonk, on October 6, 2026: a natively multimodal mixture-of-experts model with about 1 trillion parameters and 49 billion active per token. The API is live today, the open weights are promised for October 27, and the lab claims the strongest open-weight results from the US or Europe. Here is what the numbers say, what is still unverified, and what to do before the weights land.
Oct 5
Reflection AI Beam: 501B Open-Weight MoE, Apache 2.0 Weights Due This MonthOn October 5, 2026, Reflection introduced Beam, its first open-weight model: 501 billion parameters, 23 billion active, built for coding and agents. Weights are promised under Apache 2.0 later this month. Here are the numbers, how it compares with GLM, Kimi and Qwen, and what is still unverified.
Oct 5
Strata Runs Qwen3.8-Flash-Next 125B on a 12 GB Gaming GPU: Speed vs QualityStrata, an MIT-licensed open-source engine, claims to run the 125B-parameter Qwen3.8-Flash-Next on a normal gaming PC with 12 GB of VRAM and 32-64 GB of RAM, at 44-124 tokens per second. A 617-point Hacker News thread tested the claim. Here is the setup, the speed table, the quant guide, and the honest quality caveats.
Oct 3
Aleph Alpha Kolibri: A 78B Open-Weight German-English MoE With 3.5B Active ParametersOn German Reunification Day, Aleph Alpha released Kolibri, an Apache 2.0 mixture-of-experts model for German and English with 78 billion parameters but only 3.5 billion active per token. It leads open models of its size in the company's own evaluation, trails a dense Qwen3.8 27B, and is built around a German tokenizer, German reasoning and trained abstention. Here is what is solid, what is not, and how to run it.
Oct 3
Ai2 Opens AstaBrief 8B for Cited Research ReportsOn October 2, 2026, Ai2 open-sourced AstaBrief 8B, the Fast-mode model behind Asta's cited research reports. It turns a question plus retrieved excerpts into a grounded synthesis — Apache 2.0 weights, self-hostable, with clear quality and license caveats for builders.
Oct 3
Ling-3.1-flash: Ant Group's 560B Open-Weights Model Is Free on OpenCodeAntLingAGI's Ling-3.1-flash is a 560B-parameter mixture-of-experts model with 25B active parameters, and it is free to use on OpenCode. It ranks second among open-weights assistants on Design Arena's Mobile App Arena. The headline numbers are worth knowing, and so are the gaps.
Oct 3
Nathan Lambert Launches Trillium Labs for Open Frontier AIOn October 1–2, 2026, Nathan Lambert and Tom Zick unveiled Trillium Labs, a nonprofit built to reopen frontier post-training as science: recipes, data, code, evals, and failed runs. No public model shipped on day one. This is what Interconnects readers and open-weights builders should actually do next.
Oct 3
Qwen Censorship Audit: Hirundo Says the 3-Billion-Download Model Embeds China-Friendly AnswersCBS News reports that Hirundo, an Israeli cybersecurity startup, found China-aligned censorship in Alibaba's Qwen, the most downloaded open model of 2026. The startup claims it can edit the weights to remove it. The finding matters for builders, and so does the fact that the auditor sells the fix.
Oct 2
DeepSeek Harness Desktop: Should You Switch in 2026?DeepSeek opened a worldwide public preview of DeepSeek Harness with an installable desktop app, the existing Web UI, and the same Cordis "everything is a plugin" runtime. This is the durable explainer: what DSH is, what the October 2026 desktop surface changes, and whether you should leave Pi, OpenCode, or Claude Code.
Oct 1
Cohere Embed 5 Pro vs Fast: Shared Space, ViDoRe V3On September 30, 2026 Cohere launched Embed 5 in Pro and Fast tiers that share an embedding space, so you can index with Pro and query with Fast. List price is $0.12 vs $0.08 per million text tokens ($0.40 for images). ViDoRe V3 averages 85.8 and 84.5 are Cohere's own RCP-nDCG@10 numbers.
September 2026
Sep 30
Hugging Face: MCP Agents Must Verify the Source, Not Just the FactOn September 29, 2026, Hugging Face published Multiverse Computing’s ProvenanceGuard write-up: source-aware verification for Model Context Protocol agents. The failure is not a made-up fact. It is a true fact assigned to the wrong tool output. This post is the practitioner checklist for traces, pooling, and release gates.
Sep 29
China AI Chip Executives: Family Travel Now Needs Pre-ApprovalEnglish relays of a September 28, 2026 Bloomberg story say China has widened overseas travel pre-approval to spouses and children of some senior AI and semiconductor executives. The published September 15 exit-entry decree is a separate, broader rule. Here is what is sourced, what is not, and what it changes for hiring, M&A, and talent mobility.
Sep 29
MicroLLM Lab: Tiny WebGPU Models in the BrowserA Hacker News front-page demo lets you load Q4 small language models in Chrome, Safari, or Edge, cache them in IndexedDB, and chat with no server after the download. explainx.ai covers what 100M-class models are actually for, the vibe-coded UI fight, Firefox WebGPU failures, and how this lab differs from transformers.js and in-browser fine-tuning.
Sep 27
Fireworks Ember-1 Trims Kimi K3 Reasoning Tokens Without Raising PriceFireworks AI released Ember-1 on September 23, 2026 — a post-trained variant of Kimi K3 that learns shorter reasoning traces while holding benchmark quality. Published A/B work shows up to 71.3% fewer reasoning tokens and roughly 39% lower total tokens at the same per-million pricing as base K3. Here is what changed, how long the research preview lasts, and when agent coding teams should switch.
Sep 24
DeepSeek DSec: The Sandbox Platform Running 380,000 Concurrent Agent Environments, and Why Agent Execution Is "Untrustworthy"Sep 23
China Investigates DeepSeek and Moonshot Over 35 Million Requests Sent to AnthropicIf you've ever used a router, proxy, or fallback service that quietly sends your requests to a different model provider than the one you chose, this is the story that shows exactly what's at stake when that routing isn't disclosed. Chinese regulators are investigating DeepSeek and Moonshot over reports that 35 million user requests were routed to Anthropic's infrastructure without clear user knowledge.
Sep 23
Hugging Face Transformers Now Matches llama.cpp on GGUF PerformanceFor years, the practical advice for anyone running quantized GGUF models locally has been simple: use llama.cpp for raw speed, use Transformers for the wider Python ecosystem and model support. Hugging Face's Transformers library reportedly closed that speed gap, which changes a tradeoff a lot of local-AI tooling has been built around for years.
Sep 22
Hugging Face Tokenizers v1 RC: Up to ~30× Faster, Same Token IDsHugging Face published the first v1 release candidates of its tokenizers library on September 21, 2026, after v0.23.2 marked the final v0 line. The team reports up to roughly 30× faster single-threaded encoding versus v0.23 on an Apple M4 Max, 5.4–8.8× faster decoding, a crate about six times smaller, and lower peak memory — while keeping the same API and the same token IDs. For anyone training or serving on the Hugging Face stack, this is the default tokenizer path getting faster, not a side experiment.
Sep 22
Startups Are Switching to Open-Weight Models to Save MoneyA pattern surfaced in Brex's own customer payments data on September 21, 2026: early-stage companies test with closed models like OpenAI's, then switch to fine-tuned open-weight versions as they scale, recovering margin and getting better performance SLAs. Harvey, the legal AI company, went from negative margins to profitability this way. Here's the pattern, the named examples, and what it means if you're deciding where to run production inference today.
Sep 22
Xiaomi MiMo-V2.6 Launches: Open Weights, Radical Training TransparencyXiaomi's MiMo team released MiMo-V2.6 on September 22, 2026 — Flash (309B total, 15B activated) and Pro (1.02T total, 42B activated) open-weight models, plus a 9B Qwen3.5 distill. The launch followed the same unusually transparent process the team used for training, publishing a real-time RL dashboard, disclosing dropped datasets and failed experiments, and shipping detailed benchmark tables including scores where the model didn't win. It topped 558 points on Hacker News. Here's what shipped, why the transparency stood out, and where the skepticism in the discussion actually landed.
Sep 21
GLM-5.3 FlashX Lands on OpenRouter and Nous Portal at 200 Tokens/SecA September 20, 2026 news digest reports OpenRouter and Nous Portal now serve "GLM-5.3 FlashX" at roughly 200 tokens per second. That name is one syllable away from GLM-5.3-Flash, the MIT-licensed sibling SKU explainx.ai already covered — and the two are easy to conflate. Here's what's grounded in confirmed GLM-5.3 facts, what's a reasonable inference, and what simply isn't known yet.
Sep 21
Mozilla: Open-Weight Models Now 4 Months Behind FrontierA Sept. 20, 2026 AI news digest surfaced a bare headline: Mozilla says the gap between the best open-weight model and the best closed frontier model has narrowed to about 4 months. There is no linked report to verify it against yet — so this post explains what a "gap in months" claim actually needs to measure to be credible, using explainx.ai's own benchmark coverage as the grounding.
Sep 21
Pirate Face Turns Open-Weight AI Models Into Checksum-Verified TorrentsPirate Face syncs Apache-2.0 and MIT-licensed models off Hugging Face and republishes each one as a checksum-verified BitTorrent magnet link, seeded by a P2P swarm instead of one company's servers. It hit #1 on Hacker News as "Pirate Face Rescues LLM Models from Deletion" with 436 points and 133 comments — and the thread surfaced a real implementation gap, not just praise.
Sep 20
A Developer Put All 3.1M arXiv Papers Into One 16TB Hugging Face DatasetSep 20
Qwen-Image-2.1: 7B Params, Native Transparency, and a License DowngradeQwen-Image-2.1 shrinks Qwen-Image's visual generator from 20B to 7B params, unifies text-to-image and editing with native alpha-channel support, and topped Hacker News at 483 points — but it drops Apache 2.0 for a restrictive new research license Alibaba requires a separate deal to use commercially.
Sep 20
StepFun Step 5 Preview: A New Pareto Frontier for Agentic AIStepFun launched Step 5 Preview on September 20, 2026 — a 600B-parameter (27B active) Mixture-of-Experts model with 1M-token context and vision, positioned as its new flagship for agentic software engineering and finance-heavy knowledge work. The launch leans on two Artificial Analysis charts claiming a new Pareto frontier at roughly 65% lower cost than the previous efficient-tier ceiling, with open weights promised for October 15, 2026.
Sep 20
Open Models Now 78.4% of Vercel AI Gateway Token VolumeVercel's AI Gateway reportedly shows open-weight models — think DeepSeek, Kimi, GLM, Qwen — now accounting for 78.4% of the token volume routed through its platform, overtaking OpenAI specifically. That's a real signal about cost-sensitive, high-volume workloads, not a claim that open models have overtaken the market.
Sep 18
Qwen3.8-Omni-Flash: Omnimodal Agents That Edit Video, Not Just Watch ItAlibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026 — an omnimodal model that moves past describing audio and video toward acting on them: editing footage, translating dubbed dialogue while preserving voice, and building deep-research reports from a video's content. Audio input pricing drops more than 98%, and Agentic Understanding cuts token consumption by roughly 46% versus processing a whole video statically.
Sep 17
Cohere and Aleph Alpha Merge Into a Transatlantic AI DeveloperCohere and Aleph Alpha announced a merger forming what the companies describe as the first transatlantic AI model developer, combining Cohere's enterprise-AI positioning with Aleph Alpha's German sovereign-AI focus into a single roughly 1,000-employee organization.
Sep 17
GLM-5.3 Built Its Own Inference Stack. The Real Lesson Is Dense FeedbackSep 16
Alibaba Open-Sourced Its Internal Code Review AI: Open Code ReviewSep 16
DeepSeek Engineer: "I Have to Bury My Talent in Yesterday"Sep 15
Bolt Forge Gives Open-Weight Models Up to 50x More UsageSep 15
DeepSeek 4.1 Flash vs. GPT-6 Astra: What a Viral Planet-Simulation Demo Actually ShowsSep 15
Sakana Fugu Max and Ultra v2: Beating Opus 5 Without Calling ItSep 15
Sakana AI's PC-ALM: Training Deep Nets Without BackpropagationSep 14
Xi Proposes a BRICS AI Open-Source Zone on His First India Trip in 7 YearsSep 14
Yann LeCun Revives the 2019 GPT-2 Mockery — and the Staged-Release NuanceSep 12
The Moonshot AI Detention Rumor: What We Could and Could Not VerifySep 11
Cohere North Small Translate: WMT26 Leader at 83.6 (Open Weights)Sep 11
DeepSeek V4.1 Flash Cuts KV Cache HBM by ~75% — What ChangedSep 11
Hugging Face Open Alignment Team: What Builders Can Use TodaySep 11
NASA-IBM Lunar Foundation Model: Open-Source AI for Moon ScienceSep 11
Venice Adds DeepSeek V4.1 Flash — Private Access Without the Direct APISep 10
Hugging Face's security.txt Has a Note for AI Agents — And It's Not a JokeSep 10
A 4B Open-Source VLM Reportedly Beats Qwen 122B on GeoGuessr-Style BenchmarksSep 9
DeepSeek V4.1 Flash: A Two-Day Beta With a New Multimodal ArchitectureSep 9
OUI-1: The First Open-Weights Model for Generative UISep 8
Greg Diamos Built a 10K Tok/s CPU Model, and Found 3 SurprisesSep 7
China's AI Token Economy: Credit Cards, Bank Loans, and a Reported 500 Trillion Tokens a DaySep 4
Abu Dhabi's IFM Releases 6 AI Models — Weights, Data, and CodeSep 4
Cerebras Serves Qwen 3.8 27B at 1,500 Tokens/Second — Here's the CatchSep 4
GLM-5.3 Flash's 18x Price Cut, and the Fable 5.1 SimpleBench ClaimSep 2
DeepSeek Opens Its 305B V4 Flash Vision Model — Free Weights, Opus 4.8 NumbersSep 2
Qwen3.8-Flash Goes Live on QwenCloud — Same Qwen4 Preview, Now HostedSep 1
Abliteration.ai Hosts an Uncensored GLM-5.3 for Offensive Cyber WorkSep 1
GLM-5.3's "50% Coding Boost" Explained — What Z.ai Actually Measured
August 2026
Aug 31
Sarvam Champions: Creator and City Lead Program for Indian AI BuildersAug 30
GLM-5.3 Takes 3rd on Terminal-Bench 4.0 — Open Weights Beat GPT-5.6Aug 30
OrcaRouter Bakes Refusal Removal Into GLM-5.3-Flash's Native FP8 WeightsAug 30
Sarvam AI Powered a Live Hindi Dub of the Ather Konarc LaunchAug 28
Tencent Hy4 preview: 770B open weights, 1M context, Apache 2.0Aug 27
Cohere Parse 5: Near-Frontier Document Parsing at $1.50/1k PagesAug 27
GLM-5.3 Open Weights Delayed — Z.ai Misses Its Own Aug 28 TargetAug 27
Unsloth Ships 3-Bit GLM-5.3-Flash GGUFs — Runs on 128GB RAMAug 26
Alibaba Wan 3.0: 30-Second Document-to-Video APIAug 26
GLM-5.3-Flash: Ox Alpha Unmasked — 320B MIT Model on Chinese Chips (Aug 2026)Aug 26
IBM Granite 4.2: Open Reasoning Models With Agentic RLAug 25
Chinese Hackers Scale Attacks With DeepSeek and Low GuardrailsAug 25
Qwen3.8-Flash-Next: The 125B MoE Alibaba Teased on a Leaked ModelScope PageAug 21
DeepSeek V4-Flash-Vision-Exp: A Multimodal Model That Nears Opus-4.8Aug 21
GLM-5.3 Max "2nd Among Open Code Models": What the Numbers Actually ShowAug 21
Pliny's "OBLITERATED" Qwen3.8-27B: 0% Refusals, and Why That Went ViralAug 20
Ornith-1.5: Is This Open Model Really Self-Improving?Aug 20
Unsloth Ships Dynamic v3.0 GGUFs for Qwen3.8-27B — What Quant Should You Run?Aug 20
GLM-5.3 Ties Kimi K3 on the AA Intelligence Index — Without a New Base ModelAug 19
Build with Sarvam #4: 3 Voice Agents Built on India-First AIAug 19
RadixArk Miles v0.1: Production RL Stack for Frontier LLM Post-TrainingAug 18
J-Space Cognition Suite: A Community Harness Claims to Unlock DeepSeek V4 ProAug 18
OrcaRouter Ships an Uncensored Qwen3.8-27B MLX Build for MacAug 17
DeepSeek V4 Prices Just Went Up — Does It Really Match GPT-5.6?Aug 17
Qwen Hits 3 Billion Downloads — What That Actually MeasuresAug 16
GLM-5.3's 84.5% CyberGym Score Isn't Verified Yet — What "Opening to Researchers" Really MeansAug 15
Qwen3.8-27B Is Live — The Local Model Hacker News Put at #1Aug 14
GLM-5.3 Is Live: "Built to Code. Ready for Cyber Defense." — Full BenchmarksAug 14
Sakana Chat Adds Fugu, New Namazu, and In-Chat Code ExecutionAug 13
DeepSeek Harness v0.1: Run the Plugin-First Agent StackAug 13
DeepSeek V4 Pro Launch: Codex, Responses API, and New PricingAug 13
Transformers.js Crosses 10M Monthly Downloads: What It Means for BuildersAug 13
Kimi Slides: Research-to-PPTX That Stays EditableAug 13
Qwen3.8-Max Open Weights Are Live — Stripped, Relicensed, and Half-DeliveredAug 13
Sarvam Voice Agents Are Public: What Builders Can Actually UseAug 12
TwIL-LM3: A 3B Model That Beats GPT-OSS-120B — With Three Big AsterisksAug 11
antirez Ported MiniMax H3 to Apple Silicon — in C and MetalAug 11
Light Society: How a Chinese Lab "Simulated" One Billion LLM AgentsAug 11
Qwen-MM-Plugins: Make Claude Code, Codex and OpenClaw MultimodalAug 10
Mistral Patented "Code Implemented Tool Calls" — But Is It New?Aug 8
DeepSeek V4 Flash 0731 Scores 89% on ARC-AGI at $0.02/TaskAug 8
DOE Launches Genesis Open Models: A US Government Open-Weight AI ProgramAug 6
Castform + Neon: A 4B Open Model Matches GPT-5.6 Sol at 1/100th the CostAug 6
DeepSeek Warns of a "Significant" API Price Increase — No Numbers YetAug 5
Mind Lab Macaron-V1: Continual Learning via LoRA, Not Fine-TuningAug 5
Shieldstral: Mistral's 3B Moderation Model That Takes Your Policy as a PromptAug 3
Calacanis vs Musk: Is the Open–Frontier Gap Already Negligible?Aug 3
DeepSeek Flash Hit 8T Tokens in a Day — What OpenCode MeasuredAug 3
Inkling-Small: Thinking Machines’ 12B-Active MoE Matches Inkling at 1/4 SizeAug 3
Qwen3.8-Max: Coding and Cowork Pitch — Open Weights Still MissingAug 1
Crypto T-Shirts Then, Open Weights Now: “Because We Can”Aug 1
Tailscale on HF Breach: No Vuln, Still Should Have Stopped It
July 2026
Jul 31
DeepSeek-V4-Flash-0731: Codex Support and $0.14/$0.28 PricingJul 31
Kimi K3 1-Bit GGUF: 1.56TB Shrunk to 594GB, ~79% Accuracy KeptJul 31
What Are LLM Parameters? Top 10 Model Sizes (July 2026)Jul 30
Bulbul V4: Sarvam’s More Expressive Indian Voice ModelJul 30
Sarvam Code: The Coding Agent, GLM-5.2 Results and Open QuestionsJul 30
Sarvam Epoch 2026: Every Confirmed Launch and What Comes NextJul 29
Deltafin: Run Kimi K3 (2.8T MoE) on One Apple Silicon MacJul 29
Hugging Face Agent Intrusion Timeline: HDF5 Leak, Jinja RCE, Mesh PivotJul 29
Kimi K3 Architecture Explained: LatentMoE, NoPE, KDA, Attention ResidualsJul 27
Neutrino-1 8B: Ternary Weights on One ArtifactJul 27
Kimi K3 Open Weights Are Live — 2.8T Parameters, Day-0 on Together and ModalJul 27
Using an Open Model Feels Surprisingly GoodJul 27
Sakana Fugu in Claude Code: Setup, Pricing, and LimitsJul 26
Open-Weight AI’s Kubernetes Moment — Tobi KnaupJul 26
Top 10 Open-Weight Models You Can Actually Run on a LaptopJul 26
Why explainx.ai Supports Open-Source AIJul 24
The Stack v3: 5 Trillion Open Code Tokens from Hugging FaceJul 24
Open Weights Letter: OpenAI + Google Join — 70+ SignersJul 23
White House Accuses Moonshot AI of Distilling Claude Fable 5 Into Kimi K3Jul 22
Kimi K3 + Fable 5 Routing Beats Either Model Alone, Fireworks Study FindsJul 21
"American AI Is Losing" — The Open-Weights Op-Ed That Split Hacker NewsJul 21
Hugging Face Was Breached by OpenAI's Own Models During a Cyber EvalJul 21
Kimi K3 vs Claude Fable 5: Vibe Engineering on a 3D Football StadiumJul 21
Kimi Work: Moonshot Desktop Agent — WebBridge, Swarm, and Cron GuideJul 21
Microsoft Reportedly Tests Kimi K3 for Copilot and AzureJul 21
Qwen-Image-3.0: Dense Layouts, 10px Text, and a Meta-Keyword MessJul 21
Sakana AI's 'Diffusing Blame': Training Neural Networks Like Real NeuronsJul 21
Fugu-Cyber: Sakana AI Orchestrates Frontier Models for Cyber DefenseJul 20
Kimi K3 Subscription Pause: Moonshot Hits GPU Limits After Demand Spike (July 2026)Jul 19
Qwen 3.8-Max Preview: 2.4T Params, Token Plan Pricing, and Open Weights SoonJul 17
How to Run Kimi K3 on Mobile — App, API, and Agent Swarm GuideJul 17
Kimi K3 #1 on Next.js Evals and Frontend Code Arena — What the Numbers MeanJul 17
How to Run Kimi K3 Locally — Confirmed Hardware Tiers and vLLM Setup (2026)Jul 16
David Siegel on Open Source AI: The Fortune Fight Is Bigger Than SoftwareJul 16
Inkling: Thinking Machines Lab Open-Weights MoE for Customization (July 2026)Jul 16
Kimi K3: Moonshot's 2.8T Frontier Model — API, Pricing, and 1M Context GuideJul 15
PrismML Bonsai 27B — 1-Bit 3.9GB Qwen3.6 Model That Runs on iPhoneJul 13
geohot: I Love LLMs, I Hate Hype — Why Frontier Labs May Not Capture the ValueJul 10
Colibrì: Run GLM-5.2 on 25 GB RAM by Streaming MoE Experts From DiskJul 8
China May Restrict Overseas Access to Top AI Models — What Reuters ReportedJul 8
Liquid AI Antidoom: Final Token Preference Optimization Cuts Doom Loops 90%Jul 7
GLM-5.2 Goes Fully Open Under MIT: Code Arena #2, George Hotz's Daily Driver, and the Multi-Model StackJul 7
Tencent Hy3: 295B Open-Source MoE Model for Agentic Coding — Apache 2.0, Free API, 256K ContextJul 4
Fastest GLM-5.2 on AMD MI355X: Wafer AI Achieves 213 Tokens/Second
June 2026
Jun 30
Agents-A1: InternScience 35B MoE Agent Model — Long-Horizon Search, GAIA 96, and vLLM SetupJun 29
Apodex 1.0-mini: 35B Open Model Tops FutureX — Beats Sonnet 4.6 and GPT-5.5Jun 29
China's AI Playbook: Free Models, Cheap Compute, and What Happens If Intelligence Gets CommoditizedJun 29
Cline $9.99/mo GLM-5.2 Plan: Bundled Open-Weights Access ExplainedJun 29
DeepSeek V4 Official Release Mid-July 2026: Peak-Hour Pricing ExplainedJun 29
GLM-5.3: Zhipu AI Asks the Community — Vision Leads the WishlistJun 29
Qwen 3.6 27B Local Dev Guide: llama.cpp, OpenCode, and Why Dense Beats MoEJun 29
Top Chinese AI Companies and Startups in 2026: The Complete Landscape GuideJun 29
US vs Chinese AI Startups in 2026: Funding, Strategy, and Who Wins WhatJun 28
Chinese Firms Predict the AI Bubble Is Close to Bursting — What the Warning Signals MeanJun 28
Zhipu AI's New Model Matches Claude Mythos on Security Bug Detection — What It Means for the AI RaceJun 27
DeepSeek DSpark: speculative decoding for V4 Flash and Pro (51–400% faster inference guide 2026)Jun 27
Meta Llama 4: The Complete Open-Source AI Model Guide 2026Jun 26
CoffeeBench: Sakana AI Benchmarks 90-Day LLM Supply Chain ManagementJun 20
Sarvam AI: Full Capabilities Guide — Models, API, Speech, Vision & How to Run (2026)Jun 19
GLM-5.2 vs Claude Fable 5: Kilo Code's Planning Benchmark Shows a Near-Tie at 1/10th the PriceJun 19
How to Run GLM 5.2 in Claude Code, Pi, OpenCode & Every Harness (2026)Jun 16
VibeThinker 3B: A 3-Billion Parameter Model That Matches Opus 4.5 PerformanceJun 15
GLM-5.2 Beats Fable 5 on Reasoning — 24 Hours After the U.S. Export BanJun 10
Cohere North Mini Code: Open-Source Agentic Coding Model (Apache 2.0)
May 2026
May 24
DeepSeek V4 Pro Shakes the AI Industry: 34x Cheaper Than GPT-5.5 and What It Means for 2026May 23
DeepSeek V4-Pro locks in 75% permanent API discount: $0.435/M tokens, 20x cheaper than GPT-5.5May 22
Cohere Command A+: the first fully Apache 2.0 enterprise AI model that runs on 2 H100s (May 2026)May 22
Qwen 3.7-Max: The Agent Frontier and Long-Horizon AutonomyMay 6
DeepSeek-TUI: terminal coding agent for DeepSeek V4 (Rust, MCP, skills)May 4
DeepSeek V4-Pro: agent coding benchmarks, 1M context, and API economics
April 2026
Apr 27
DeepSeek V4 preview: V4-Pro, V4-Flash, 1M context API (2026)Apr 22
What are parameters in a large language model? Billions, MoE, and what 2026 model cards really sayApr 17
GLM-5.1 on Hugging Face & how to run it (Z.ai API, Ollama, vLLM) — 2026 guideApr 17
Netflix VOID on Hugging Face: video object removal that respects physics (model card recap)