explainx / blog / topics
Local AI
Local AI means running models on your own laptop, workstation, or phone instead of a cloud API. It trades some capability for privacy, cost control, and offline use.
This page collects our coverage of local runtimes, on-device models, and the hardware that makes them practical.
53 stories · latest Oct 9, 2026
Start here
Surface Laptop Ultra: Which Configuration to Buy for Local AI
Microsoft opened preorders for the Surface Laptop Ultra on October 7 with prices from $2,599.99 and shipping from October 16. The base 24GB model and the 128GB model are very different machines for local AI. This guide does the memory math, ties it to the new MAI-Code-1.1 Flash, and says who should buy which tier or wait.
How AI Toothbrushes Work: Sensors, Models, and Dyson's Camera
"AI toothbrush" covers two genuinely different engineering approaches: a motion-sensor classifier that infers mouth position from an accelerometer and gyroscope, and — as of Dyson's September 2026 CameraJet — an onboard camera running a vision model in real time. Here's how each actually works, what data trains them, and what the mechanism means for anyone building their own edge-AI product.
What Is Ollama? $88M Funding, 9M Builders, and the Open-Models Bet (July 2026)
Docker Desktop founders Jeff and Michael raised $88M for Ollama on July 9, 2026 — 9M+ active builders per @ollama, hybrid cloud scaling, day-one open model support. explainx.ai explains what Ollama is and why it matters now.
SOTA LLMs Locally: Jamesob’s $46k RTX PRO 6000 Hardware Guide
Bitcoin developer jamesob released a comprehensive guide to running SOTA LLMs locally on a custom $46,000 GPU rig. We break down the hardware, BIOS configurations, kernel tweaks, and the ongoing local AI debate.
What Is llama.cpp? Install, Run GGUF Models, and Serve OpenAI-Compatible APIs
If you run open weights on your own hardware in 2026, you are almost certainly touching llama.cpp — directly or through Ollama and LM Studio. This guide explains what it is, how GGUF fits in, copy-paste install and run commands, and how to expose a local API for coding agents.
MacBook vs dedicated GPU for local LLMs: how much RAM you really get, and when each wins in 2026
MacBooks behave like a slow GPU with enormous shared RAM; dedicated cards are fast but VRAM-capped. The right buy depends on whether you wanted a laptop anyway, need privacy at 64k context, or need frontier-speed coding throughput.
Timeline
October 2026
Oct 9
Cactus Whistle: A 16.9 MB Speech-to-Text Model That Runs on a CPUOn October 2, 2026, Cactus Compute released Whistle, a 16.9 MB Apache 2.0 speech recognition model that runs on one CPU engine beside the Needle tool-calling model. It beats Whisper base on several benchmarks and loses on others. Here is the full picture.
Oct 9
Samsung LittleBit: A 13B LLM Under 1 GB, and What the Paper Really ShowsSamsung Research released LittleBit code that compresses LLMs to 0.1 bits per weight by factorizing and binarizing each weight matrix. The headline says a 13B model fits under 1 GB. The paper says so too, but the method needs training, and the license blocks commercial use.
Oct 8
Surface Laptop Ultra: Which Configuration to Buy for Local AIOct 7
Oki Home: A $1,799 "Memory Computer" With a Local Qwen 3.8 27B and a Swappable MemchipYC-backed Oki launched Oki Home on October 6, 2026: a 5-liter PC with an RTX 5060 Ti, 32 GB RAM and a credit-card-sized 2 TB Memchip that holds your photos, messages and files on one searchable timeline, with a local Qwen 3.8 27B model claimed at 106 tokens per second. Reservations are $59 and shipping is planned for mid-December. Here is what is claimed, what is not shown, and how it compares with doing it yourself.
Oct 5
Ghost Core: The $3,499 Personal AI Computer That Runs Models On DeviceGhost launched Core, a $3,499 personal AI computer that runs open-weight models entirely on device, ingests your apps, files and wearables, and acts without being asked. Batch 1 is already sold out. Here is what Ghost has said, what it has not, and how it compares with building your own local AI box.
Oct 3
DwarfStar 4: Can You Run ds4 on Your Mac This Week?Salvatore Sanfilippo's ds4 is back on the Hacker News front page. The useful question is whether your Mac, Spark, or Strix Halo can hold a project quant this week, and whether Claude Code or Codex can talk to it.
Oct 3
llama.cpp Adds Decision Model Support for Five Open ModelsOn October 2, 2026 the ggml team announced decision-model support in llama-server. You POST state plus typed questions to /v1/systemone and get option probabilities in one forward pass. Five open models ship as GGUF today — and the local runner map just got a third path next to Ollaya and Python/MLX stacks.
Oct 3
DGX Spark 64GB: What $4,999 Actually Buys on Oct 23Feed headlines said NVIDIA "halved" DGX Spark memory for a $4,999 Oct 23 launch. The official story is a new 64GB OEM SKU beside the 128GB model — same GB10, up to 100B params solo, 128GB when you pair two. Here is what that means for local LLM buyers.
September 2026
Sep 18
PrismML Bonsai 2 27B: 9x Smaller, 98.2% of Full-Precision QualityPrismML released Ternary Bonsai 2 27B on September 17, 2026: a 1.76-bit compression of Qwen3.8 27B that fits in 5.9GB while retaining 98.2% of the full-precision model's aggregate benchmark score — up from 95% in the first Bonsai release two months earlier. It runs coding agents locally at 143 tokens/second on an RTX 5090, but Hacker News users report inconsistent real-world throughput and at least one clear reasoning failure in testing.
Sep 17
BITCOS: Breaking the 1.58-Bit Barrier for Ternary LLMsA September 14, 2026 paper from Intel Labs researchers challenges an assumption baked into every ternary LLM deployed today — that the three weight values are roughly equally likely. They measured 29 real ternary models and found zeros can account for over half of all weights, then built BITCOS, a packing format that exploits this to beat the standard five-trit format on 26 of 29 models.
Sep 10
Edge0 Open-Sources a 35B Model That Runs in 2.5GB of MemoryEdge0, founded by Samuel Zeng, open-sourced a framework that runs a 35-billion-parameter language model using only 1-2.5GB of peak memory — small enough for a phone, in principle. The technique is "SSD expert offload": stream only the parameters a given step needs from storage instead of loading the whole model into RAM. explainx.ai covers how it works, what's actually shipped today versus the demo, and what it means for on-device AI.
Sep 5
How AI Toothbrushes Work: Sensors, Models, and Dyson's CameraSep 3
Tether Released Offline Translation Models for 19 African LanguagesTether — the company behind the USDT stablecoin — released TranslatePsy-AfriSLM on September 2, 2026, a family of open-source, on-device translation models covering 19 Sub-Saharan African languages, claiming the 800-million-parameter version beats much larger models like Qwen3.5-122B on translation benchmarks. Here's what's actually in it.
Sep 2
Gemma 4 26B A4B Runs 2x Faster on Mac — Community MLX OptimizationGoogle's official Gemma account credited a public MLX leaderboard for making Gemma 4 26B A4B run dramatically faster on Apple Silicon. The precise figure is 130.3%, not a clean 2x—here's what actually changed, who did the work, and what it means for running MoE models locally on a Mac.
August 2026
Aug 25
Apple M6 Mac mini: On-Device AI, Dual Neural Engine, and What 32GB Actually RunsOn August 25, 2026, Apple launched the M6 Mac mini — its first 2 nm chip with a Dual Neural Engine and Neural Accelerators in every GPU core. explainx.ai breaks down what that hardware actually runs locally, how it compares to the M5 Pro option and M5 Ultra Studio, and whether the $899 entry price still makes sense for AI builders after Namespace spent a week putting MacBooks in racks.
Aug 25
Mac Studio M5 Max and M5 Ultra: On-Device AI for Builders Who Outgrew the MiniOn August 25, 2026, Apple launched Mac Studio with M5 Max (128GB, from $2,499) and M5 Ultra (512GB, 4.3× peak AI vs M3 Ultra). explainx.ai breaks down what that silicon runs locally — MLX, LM Studio, Thunderbolt 5 clustering — and whether the fully specced $18,299 box beats a dedicated GPU for your workload.
Aug 25
Can You Use MacBooks as Servers? Namespace’s Rack Video ExplainedOn August 24, 2026, @namespacelabs posted a video of MacBooks going into server racks (~2.8M views). The internet asked one question: why laptops? explainx.ai answers whether MacBooks can be servers, what Namespace actually builds, and when Mac minis, Studios, or Nvidia boxes win instead.
Aug 25
Xiaomi Xring O3: 44MB Cache and the Wide-Core Trend Daniel Lemire SpottedOn August 24, 2026, Daniel Lemire unpacked Xiaomi's Xring O3 — ~3,945 Geekbench single-core, ~15,221 multi-core, 44MB cache, ARM C1-Ultra cores on TSMC N3P. explainx.ai maps what is real, what is lab-only, and why Qualcomm should worry.
Aug 24
FreeToken Runs a 753B MoE Locally — But the GPU Is Only Half the StoryThe viral version says FreeToken runs a 753B model on a single GPU. The paper does — at 14.9 tok/s — but on a workstation with a 96GB RTX PRO 6000, 512GB of system RAM, and a 433GB checkpoint. This guide explains the whole machine, why its bandwidth-adaptive MoE runtime matters, and which smaller tier is realistic for your hardware.
Aug 24
Pipette: Liquid AI’s Open On-Device Benchmark SuiteOn August 24, 2026, Liquid AI and Artificial Analysis released Pipette — an open platform measuring on-device AI as a full system (model × quant × runtime × device), not model cards alone. 1,000+ configs, iOS/Android clients, public data.
Aug 24
OCR It: Pin a Region, Hotkey Through a Book, Feed Your LLM OfflineHalf the documents you want in an LLM cannot be selected: scanned books, slide decks, locked PDF viewers. OCR It pins a rectangle once, captures each page with a hotkey or auto-run, and runs bundled Tesseract entirely offline. explainx.ai walks through install, auto-pagination, and the MV3 tricks.
Aug 20
S1-mini: Superwhisper's 0.6B On-Device Transcript CleanerSuperwhisper open-weighted S1-mini, a 0.6B Qwen3 fine-tune that sits after speech-to-text and rewrites messy ASR into clean English. It is not a replacement for Parakeet TDT v3 or Cohere Transcribe. explainx.ai covers the pipeline, the license catch, and how to run the GGUF locally.
Aug 19
DFlash-MLX Brings Lossless Speculative Decoding to Apple Silicon — Up to ~189 tok/s on M5 Maxbstnxbt/dflash-mlx ports DFlash block-diffusion speculative decoding to MLX on Apple Silicon with lossless greedy verification. On an M5 Max 64GB, Qwen3.5-4B jumps from 54 to 189 tok/s at 2048 tokens, while Qwen3.5-27B-4bit lands at 70 tok/s — the headline "~70 tok/s" figure — at roughly 2.1x over baseline. Speedup varies sharply by model size, architecture, and context length.
Aug 19
RAM Prices Are Up 500% — What That Means for Local AI Builds128GB of DDR5 now costs $3,399, roughly 10x the lowest price ever tracked, and average DDR5 kit prices are up 350-485% year-over-year — driven by AI datacenter demand locking up global memory production. For anyone building a local-inference rig, this changes the math significantly, and not for a year or two.
Aug 17
Matic Cues: How a Home Robot Runs Voice, Vision, and Mapping On-DeviceMatic Robots launched Cues, a voice-and-gesture control layer for its $115M-funded home robot, built entirely on an Nvidia Jetson Orin Nano. It's a real case study in edge AI system design — wake-word detection, 3D spill localization, and house-scale navigation, all running locally with no data leaving the device.
Aug 16
A 1.5B Model That Writes Your Shell Commands — On a Laptop CPU, No GPUA developer tired of googling tar flags fine-tuned Qwen2.5-Coder-1.5B on 125k natural-language/command pairs, quantized it to Q4_K_M, and got a 941MB model that matches an untuned 7B on InterCode-ALFA while running on four CPU threads. It still trails GPT-4o by 0.11. Here is the full recipe, the honest benchmark read, and the safety design worth copying.
Aug 12
Unsloth Desktop: One Local App That Both Trains and Runs AI ModelsEvery local AI app so far has done inference. Unsloth Desktop does inference and training in the same window — LoRA and full fine-tuning at 2x speed and 70% less VRAM, plus GGUF, MLX, diffusion and audio models, and a model-swap bridge into Claude Code and Codex. explainx.ai covers what it actually does and where the catches are.
Aug 5
LFM2.5-2.6B: Liquid AI's Biggest On-Device Agent Model YetLFM2.5-2.6B is Liquid AI's flagship on-device agent model — 2.6B parameters, 34 trillion training tokens, and benchmark scores that beat Gemma-4-E4B and match Qwen3.5-9B on tool use, while running under 2.5GB of memory on a phone. explainx.ai covers the numbers and where it fits next to Liquid's smaller LFM2.5-230M.
Aug 3
BitNet on a 1975 6502: A Language Model Inside 25KBWeights and inference fit in ~25KB of BBC Micro userspace. Ternary BitNet kills the need for multiply; Mamba kills the KV cache. The model is tiny and janky — and a perfect case study in mechanical sympathy.
July 2026
Jul 30
TurboFieldfare: Gemma 4 26B in ~2 GB RAM on Apple SiliconAndrey Mikhaylov’s TurboFieldfare keeps Gemma 4’s shared core and KV in RAM and preads routed experts from disk — ~2 GB process RSS for a ~14 GB model. explainx.ai covers how it works, scores, macOS 26 limits, and when to use MLX instead.
Jul 28
Framework Laptop 13 Pro: Modular Hardware for Local AI (2026)MKBHD’s Framework 13 Pro walkthrough lands as Framework ships a ground-up redesign: CNC aluminum, haptic trackpad, LPCAMM2 RAM, and claimed 20-hour battery — without abandoning repairability. Here’s what that means if you run models on a laptop.
Jul 26
28.9M Params on an $8 ESP32 — How It FitsA 28.9M-parameter model on a ~$8 ESP32-S3 writes stories to a tiny OLED with no Wi-Fi. explainx.ai unpacks Per-Layer Embeddings, the SRAM/PSRAM/flash split, and why this is architecture news — not ChatGPT on a chip.
Jul 23
Petals Resurfaces: Why BitTorrent-Style LLM Inference Still StrugglesA 2022 Hugging Face/BigScience project called Petals — run large language models at home, BitTorrent-style — hit the Hacker News front page again in 2026, reigniting a debate about whether peer-to-peer LLM inference is finally viable now that models are smaller, quantization is better, and newer projects like Mesh LLM and AI Horde have taken different approaches to the same problem.
Jul 19
Moonshine Micro: Voice AI on 80-Cent MCUs — RP2350 VAD, STT, and Neural TTS (2026)Pete Warden's Moonshine Micro brings voice activity detection, command recognition, and neural text-to-speech to microcontrollers — reference demo on the Raspberry Pi RP2350 (~80 cents) in as little as 470 KB RAM. MIT-licensed, TensorFlow Lite Micro, and a full Wi-Fi provisioning walkthrough.
Jul 17
LM Studio Bionic: Open-Model Agent for Code and Work ProjectsLM Studio shipped Bionic on July 16, 2026 — a dedicated agent app (not LM Studio itself) for code repos and work projects over local models, LM Link, or Secure Cloud with zero data retention. This guide covers what works, HN rough edges, closed-source trade-offs, and how it compares to OpenCode and Unsloth Studio.
Jul 14
Tencent Hy3 GGUF — 1-Bit and 4-Bit Quants for Single-GPU llama.cppEight days after Hy3's launch, Tencent dropped 1-bit and 4-bit GGUF quants claiming single-GPU serving via llama.cpp + MTP. That means 128GB unified memory — not a 16GB RTX 3060. explainx.ai breaks down hardware tiers, p-min flags, and honest tok/s from the X thread.
Jul 12
Mesh LLM v1.0: Split 235B Models Across Your LAN with iroh P2Pn0's Mesh LLM 1.0 exposes distributed inference as localhost OpenAI API. Run locally, route to peers, or Skippy-split giants across your mesh. explainx.ai architecture, benchmarks, and HN perf debate.
Jul 9
What Is Ollama? $88M Funding, 9M Builders, and the Open-Models Bet (July 2026)Jul 4
Local LLMs Keep Looping? Fix It With Samplers, Not More VRAMA Hacker News thread on jamesob's local-LLM rig turned into a practical guide on its own: why 4-bit models get stuck in loops on long tasks, which llama.cpp samplers actually fix it, which harness to run them in, and how to sandbox an agent that has full filesystem access.
Jul 4
SOTA LLMs Locally: Jamesob’s $46k RTX PRO 6000 Hardware GuideJul 3
Can You Self-Host Your Photos? Immich 3.0, Privacy, Costs, and When to Blur FacesJul 2
What Is llama.cpp? Install, Run GGUF Models, and Serve OpenAI-Compatible APIs
June 2026
Jun 29
MacBook vs dedicated GPU for local LLMs: how much RAM you really get, and when each wins in 2026Jun 29
Ollama 0.31: Gemma 4 Is ~90% Faster on Apple Silicon With Multi-Token Prediction (No Output Change)Jun 26
LFM2.5-230M: Liquid AI's 230M Model Built to Run Agents on Phones and RobotsJun 26
TREK: Self-Hosted Travel Planner with Real-Time Maps, Budgets, and AIJun 17
NVIDIA DGX Spark: The Best Setup for Running Local LLMs in 2026Jun 16
What Is AI Model Quantization? Running Frontier AI LocallyJun 15
Build Your Own Personal AI System: The Complete 2026 Guide to Local Models, Frameworks, and Workflows