Hacker News put MicroLLM Lab on the front page on 28 September 2026 (about 135 points when we pulled the thread): seven tiny language models you can try in a tab, no API key. The live experiment is stateofutopia.com/experiments/microllmlab; the repo is robss2020/microllm-lab. HN user logicallee posted it.
This is inference, not training. explainx.ai already covered in-browser fine-tuning over WebGPU as a harder claim, and WebGPU itself as the GPU API. MicroLLM Lab is the complementary object: a multi-model Q4 playground so you can feel what a 100-million-parameter network actually does — including the jokes, the hallucinations, and the Firefox crashes.
TL;DR — what people asked on HN
| Question | Direct answer |
|---|---|
| What is it? | Browser lab: Q4 small language models (roughly 26M–360M), WebGPU first, IndexedDB cache, no server after load |
| How many models? | HN title says 7. The GitHub catalog table lists 8 Q4 packs, including first-generation SmolLM 135M Instruct plus SmolLM2 |
| Useful for chat? | No. Useful as an edge layer: classify, route, extract, demo privacy. Not Opus/Fable |
| Zip size? | Download button on the live page says 589 MB; README and footer say current.zip (591 MB) |
| After load, is there a server? | Weights fetch from the origin once; chat runs locally. First visit still hits the host for the zip/weights |
| Tok/s on your GPU? | Do not generalize. README reports an Apple M4 / Safari / Metal Q4 run: PetitGPT 115 tok/s and SmolLM2 135M Instruct 66 tok/s on a short “Say hello…” prompt. We did not re-bench |
| Firefox? | Linux + some GPUs: GPUShaderStage is not defined. Windows Firefox 156.0.1 reportedly worked for the author |
| vs transformers.js? | Lab = catalog + custom kernels + UI. Transformers.js = production ONNX pipelines |
| Fine-tuning in the tab? | Not this project. That is the training-side explainer |
What shipped (from the live page and README)
The page title is “tiny LLMs, Q4, in your browser.” Tags on the hero: WebGPU · Zero-Server · Q4 Quantized. Tabs: Chat, Benchmarks, Compare & Certificate, Custom eval. A runtime HUD shows backend, tok/s, TTFT, JS heap, GPU buffers, IndexedDB. Prefer WebGPU is on by default and is documented as falling back to WASM, then JS — that fallback did not rescue the Firefox GPUShaderStage reports below, which fail at parse/load time.
Weights live in IndexedDB after you click Load. Opening index.html as file:// is blocked (ES modules, shaders, IndexedDB). Serve it:
python3 -m http.server
# open http://127.0.0.1:8000/
Lab code is Apache-2.0. Checkpoints keep upstream licenses (Apache-2.0 except GPT-2 MIT). README: greedy decode only, context capped at 2048 so the KV cache fits a laptop GPU, and the models will hallucinate.
Custom eval is explicit about risk: the studio eval()s JavaScript in this origin, then runs check functions on decoded text. Treat that as a local toy, not a multi-tenant eval SaaS.
The README names Metal and fused-NVIDIA WebGPU decoders for Llama-style GQA/SwiGLU and GPT-2/nanoGPT LayerNorm+GELU. Conversion helpers: tools/convert_hf_to_pgw.py and tools/convert_gpt2_to_pgw.py, then a row in models/catalog.json.
Which checkpoints are in the box?
Parameter counts and Q4 file sizes below are from the project README (not from our own download of every shard).
| Model | Params | Q4 size | License | Notes (README) |
|---|---|---|---|---|
| PetitGPT research-v1 | 124.6M | 74 MB | Apache-2.0 | yangqi0; Llama-style; lab’s usual speed baseline |
| SmolLM2 135M Instruct | 134.5M | 80 MB | Apache-2.0 | Hugging Face Smol Models Research, 2025; ChatML |
| SmolLM 135M Instruct | 134.5M | 80 MB | Apache-2.0 | 2024 first-gen SmolLM; same size class |
| L20-Edu 135M | 134.5M | 80 MB | Apache-2.0 | AliceYin; ~13B tokens on one L20; weaker on purpose |
| SmolLM2 360M Instruct | 362M | 216 MB | Apache-2.0 | Larger sibling; slower load/decode |
| MiniMind2 104M | 104M | 62 MB | Apache-2.0 | jingyaogong; stronger Chinese than English |
| MiniMind2 Small 26M | 26M | 15 MB | Apache-2.0 | Teaching checkpoint; shaky English |
| GPT-2 124M | 124M | 77 MB | MIT | Feb 2019; completion, not chat-tuned |
PetitGPT is a WebGPU port of yangqi0/petitgpt, not a replacement of that research repo. GPT-2 is openai-community/gpt2 (Radford et al.), the architecture Karpathy’s nanoGPT trains. README: greedy GPT-2 often loops; that is the checkpoint, not a decoder bug.
Q4 here is described as group-32 packing. For the general “why 4-bit exists” story, see explainx.ai’s quantization guide. The live page’s primer claims ~50–84 MB for 100M-class Q4 in browser memory; the README table is the more precise per-file list.
What are 100M models actually useful for?
This was the first real question on the thread. Someone pasted a bromine-tub / spa chemistry prompt into a model and got garbage. The reply that stuck: that is not what SLMs are for. They have little parametric knowledge and weak reasoning. Useful jobs named in-thread: sentiment, classification, entity extraction — not “replace Opus.”
That matches the lab’s own copy: compact models as an edge layer — triage, spam/intent filters, decide whether a cloud LLM is even needed. It is the same split explainx.ai already uses for Transformers.js in agent pipelines (classify/embed locally, call an MCP server only when the step needs a frontier model) and for Calvin French-Owen’s “small models have arrived” economics — except those essays are about cheap hosted small models, not 135M in a tab.
If your task is pick among N labels, you may not want a chat decoder at all. Julia 1 is a 144M encoder for typed decisions on CPU or WebGPU. MicroLLM Lab is the opposite demo: generative tiny decoders, which is why HN immediately used them as chatbots and then dunked on the answers.

Honest failures: 2+2, California, soup
The README is unusually blunt: a 135M net will fail arithmetic and invent facts, and the lab measures that instead of hiding it behind chat skin.
HN receipts (paraphrased, not a scientific sample):
- PetitGPT on “What is 2+2?” produced a rambling “add 2 to both sides / 2+2 = 4+2” answer (commenter tolugenius). Another user (inventor7777) got “2+2 is 2.” The author posted a screenshot of a correct PetitGPT 2+2 in their own run — variance is the point at this size, not a leaderboard.
- GPT-2 124M on the same prompt wandered into “3+3? 4+4?” completion noise (tecleandor). That is expected for a 2019 base model.
- SmolLM2 360M Instruct on “population of California” invented 153 million, then 325 million, then mixed in GDP (not2b). Fast, still wrong.
- A soup recipe degenerated into Saffa-Cake-Sweet-… repetition (NicuCalcea).
README’s own M4 Safari objective suite (greedy): PetitGPT 10/14 on an older short set; SmolLM2 135M Instruct 14/20 on an expanded set that includes 1+1=2, 2+2=4, Berlin, copy-from-context. Suite wall ~1.5 s vs ~6 s. Those numbers are author hardware, not yours.
Do not ship these weights as a medical, legal, or tub-chemistry assistant. Do use the suite to teach “small ≠ ChatGPT.”
GPT-2 2019 vs 2026 instruct in the same HUD
logicallee’s useful framing: GPT-2 is February 2019, years before the November 2022 ChatGPT preview. In 2019, output often looked like pasted web fragments. Putting that checkpoint next to SmolLM2 Instruct (2025) in one IndexedDB-backed UI is the pedagogy: same browser, same tok/s widget, different training recipe and year.
Commenters who expected ChatML-style answers from GPT-2 were running a completion model. README: it is not chat-tuned. MiniMind2 is the other mismatch: Chinese-first teaching weights in an English HN crowd.
Vibe-coded dense UI vs demo value
The second-longest subthread was not about transformers. It was about type size, above-the-fold demo, and AI-generated explainer walls.
demibabs: too small, too dense, a page of mostly useless intro before the interface; footer that links to itself and says “serve over HTTP.” janalsncm: humans are not transformers; the first thing you see should be the demo; the intro still called GPT-4 a frontier model. TZubiri: vibe-coded; give the model screenshots / a browser, not only source. askjdfksdbfhk: people upvoted in spite of the UI.
logicallee’s reply: the text is fine print so you can skip it; they counted ~600 words on the page; the footer is for people who unzip current.zip; the submission stayed on the front page for hours so they would not change it. Later they said the original UI was click a model, chat, and they asked a model to add vocabulary and three steps because a blank lab did not explain itself.
explainx.ai’s read: both can be true. The demo value (seven architectures, IndexedDB, local suite, GPT-2 vs instruct) is why it ranked. The Claude-dashboard density is a real 2026 failure mode — extra copy is cheap, attention is not. If you fork the lab, put Load + composer above the fold and keep the primer behind a details element. That is a product lesson, not a reason to ignore the kernels.
Firefox: GPUShaderStage and Radeon
This is the bug report to remember if you copy the architecture.
- anonimous-emacs: Firefox,
Uncaught ReferenceError: GPUShaderStage is not defined. - Author: works on Firefox 156.0.1 Windows; asked for OS/GPU; suggested unchecking Prefer WebGPU.
- ad_fontes: Firefox 156.0, Ubuntu GNOME, AMD Radeon 860M; unchecking WebGPU and reloading did nothing.
- cwnyth: same error, Intel Iris, Firefox on Fedora; Prefer WebGPU did not help.
- Author: hard to fix without a Radeon; would look the next day.
GPUShaderStage is a WebGPU global (VERTEX / FRAGMENT / COMPUTE). If the engine evaluates a module that names that binding before a feature-detect, Firefox-without-WebGPU (or with a stub) throws before the WASM path. That is why a checkbox that only changes runtime preference cannot help a load-time ReferenceError.
Chrome / Edge / Safari were the author’s happy path (Windows, Mac, iPhone). WebGPU support is still browser × OS × GPU. Our WebGPU guide is the API map; this thread is the interop footnote.
vs transformers.js, in-browser fine-tuning, Web Models API
Three adjacent objects people searched after the headline:
Transformers.js. Hugging Face’s JS port, ONNX Runtime, WebGPU or WASM, tens of millions of monthly npm downloads across current + legacy packages — covered in Transformers.js at 10M downloads. It is how you productize classification and embeddings in a tab. MicroLLM Lab is a research-style harness: custom Q4 format (pgw), fused kernels, Chart.js certificates, not the Hugging Face pipeline API. If you need embeddings in 7MB of WASM instead of a 80MB decoder, see Ternlight and the WebAssembly guide.
In-browser fine-tuning. That post is about backprop in a sandbox (likely LoRA, small data). MicroLLM Lab never claims training. Confusing the two is how “I ran a model in Chrome” becomes “I trained GPT in Chrome.”
Web Models API. Commenter kenzic pointed at a browser-standard sketch: webmodels.dev — an API for on-device open-weight models rather than every site shipping its own WebGPU decoder. MicroLLM Lab is existence proof that sites will ship decoders until the platform grows a shared cache and permission model. We are not endorsing the proposal’s current draft; we are noting it is the standards question this demo raises.
How to try it without fooling yourself
- Use Chrome, Edge, or Safari first if you just want the catalog. Firefox on Linux GPUs is currently the landmine.
- Load one 135M instruct model (SmolLM2 135M), not GPT-2, if you want to see the best of this size class.
- Run the Benchmarks tab on your machine. Do not quote README tok/s as yours.
- Prompt for labels (“spam or not”, “which of these three intents”) before you prompt for world knowledge.
- If you download the zip, HTTP server, not
file://.
If you are building an agent skill that only needs a router, a 135M decoder in-tab is still optional — a tiny classifier or Julia 1-style decision head is often the more honest architecture.
The bottom line
MicroLLM Lab is a front-page WebGPU inference zoo, not a new assistant. The catalog (PetitGPT, SmolLM/SmolLM2, L20-Edu, MiniMind2, GPT-2) plus IndexedDB plus an honesty-first eval suite is the contribution. The HN thread already did the product review: classification not chat, UI too dense, Firefox GPUShaderStage, 2019 GPT-2 vs 2025 instruct, California population fanfic. Use it to calibrate your expectations of 100M-class Q4 — then ship real work on transformers.js or a purpose-built encoder, not a chat box glued to MiniMind2-26M.
Related reading
- In-browser LLM fine-tuning on WebGPU — training vs this lab’s inference-only demo
- WebGPU complete guide (2026) — compute shaders, browser GPU access
- Transformers.js crosses 10M monthly downloads — the library you actually embed
- What is AI model quantization? — why Q4 exists
- Ternlight: 7MB browser embeddings in WASM — even smaller on-device text ML
- WebAssembly complete guide — CPU fallback when WebGPU is missing
- Julia 1: 144M decision model on CPU / WebGPU — classification without a chat decoder
- Small models have arrived (Luna economics) — cheap small models in the API era
External: Live lab · GitHub robss2020/microllm-lab · HN discussion · PetitGPT upstream · Web Models API sketch
HN score (~135), comment quotes, README model table, zip sizes (589 MB vs 591 MB), and M4 Safari tok/s figures are from sources retrieved on 29 September 2026. Token rates are the project’s own Apple M4 / Safari measurements, not explainx.ai benchmarks. Firefox GPUShaderStage reports are user comments, not a complete browser matrix. Check the live page and repo for catalog changes after publication.
