explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: what changed and what to expect
  • Why was speculative decoding slow on Metal in the first place?
  • What the new kernels do
  • PR 30065: the remaining quantization types
  • Does this make speculative decoding 3.4x faster?
  • How to test it on your Mac
  • Who should care
  • Caveats
  • Related reading
← Back to blog

explainx / blog

llama.cpp Metal Gets Few-Row MMA Kernels: Up to 4x Faster Verification on M3 Ultra

llama.cpp, Local AI, Apple Silicon, Speculative Decoding, Open Source

Part of Local AI

llama.cpp b11404 adds Metal few-row matmul kernels that speed up speculative decoding on Apple Silicon. Real PR numbers, what they mean, how to test.

Oct 10, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
llama.cpp Metal Gets Few-Row MMA Kernels: Up to 4x Faster Verification on M3 Ultra

llama.cpp just fixed one of the quiet reasons speculative decoding disappoints on a Mac. Build b11404, published October 5, 2026, adds Metal "few-row" matrix kernels, and a follow-up merged on October 7 extends them to almost every quantization format. On an M3 Ultra, the author's own kernel benchmarks show the new paths running up to 4.4x faster on the small batches that speculative decoding creates.

A headline circulating this week puts it as "3.4x faster speculative decoding on M3 Ultra." That is not a number we could find in the primary sources, and it is easy to misread. The PRs report matmul kernel speedups, not end-to-end tokens per second. This post explains what changed, what the numbers say, what they do not say, and how to measure the benefit on your own hardware. It sits next to our earlier local-inference coverage, including the Gemma 4 MLX Mac speedup and the Perplexity Lily Apple Silicon engine.

TL;DR: what changed and what to expect

table · 2 cols
QuestionAnswer
What shipped?Metal few-row MMA mat-mul kernels for 2 to 16 rows (PR 29869, in build b11404), extended to the remaining types by PR 30065 (merged October 7).
Who benefits?Mac users on Apple7 and newer GPUs without the tensor API running speculative decoding, or other small-batch decoding.
Do I need a flag?No. Dispatch is automatic once the row count passes a per-type threshold.
Best published number?4.41x on Q1_0 at 8 rows, 4.30x on Q2_K at 8 rows, 4.03x on TQ2_0 at 16 rows (M3 Ultra, matmul kernel only).
Does plain chat get faster?Not much. One-row generation is a matrix-vector job and the PR says 512-row results are within 1 percent of before.
Is the 3.4x end-to-end figure confirmed?No. It is not in the PRs or release notes. Treat it as unverified.
Is the build stable?b11404 is marked a pre-release on GitHub. Like all llama.cpp builds it moves daily.
Weekly digest3.6k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Why was speculative decoding slow on Metal in the first place?

Speculative decoding is a trick for generating text faster without changing the output. A small, quick "draft" proposes several next tokens. The big model then checks all of them in a single forward pass and keeps the ones it agrees with. The idea comes from the 2022 paper on fast inference from transformers via speculative decoding, and it is now standard in local runtimes. We covered a production version in our DeepSeek DSpark guide.

A small fast arrow running ahead of a longer checking arrow with a tick, showing the draft-then-verify loop in speculative decodingA small fast arrow running ahead of a longer checking arrow with a tick, showing the draft-then-verify loop in speculative decoding

The catch is the shape of the verification math. Normal generation multiplies the model weights by one vector (one row). Verification multiplies them by a small matrix of two to sixteen rows, one per drafted token. On a GPU, that is the awkward middle: too many rows for a pure matrix-vector kernel to stay efficient, too few for a big-batch matrix kernel to pay off.

According to the b11404 release notes, Metal handled these cases with matrix-vector kernels whose cost grows with every extra row. The notes state that on an M3 Ultra, DFlash2 decoding was slower than serial decoding. In other words, the verification step ate the gains the draft was supposed to provide.

What the new kernels do

The release notes for PR 29869 describe the fix in three parts.

  • New kernels for 2 to 16 rows. They use 8x8 simdgroup matrices. Each weight is dequantized once and reused for all rows, and simdgroups in a threadgroup split the K dimension to keep the GPU busy.
  • Per-type thresholds. The kernels only kick in where they beat the old path on an M3 Ultra: at 6 rows for F32, 3 rows for F16, Q4_K, Q5_0 and Q5_1, and 2 rows for the other types. Q4_0, Q8_0 and Q5_K get dedicated kernels, while F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path.
  • Fusion. A matmul followed by an add can fold a same-shape residual into the matrix store, and long CONCAT rows are split across threadgroups when there are few rows.

They apply only on Apple7 and newer GPUs without the tensor API, so the benefit is tied to GPU family rather than to the M3 Ultra alone, although that machine is where the author tuned the thresholds.

Laptop outline with a small green chip representing Apple Silicon running local model inferenceLaptop outline with a small green chip representing Apple Silicon running local model inference

PR 30065: the remaining quantization types

The first PR covered the common formats. PR 30065, "metal : few-row MMA mat-mul for the remaining src0 types," by pratiknarola-t, was opened October 6, approved October 7 and merged the same day by Georgi Gerganov. It extends the kernel to every type that has a 16-weight dequantizer: BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0 and the IQ family (IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL, IQ4_XS).

The author's M3 Ultra measurements, for a 4096 by 14336 matmul using test-backend-ops perf -o MUL_MAT, relative to the prior master:

table · 5 cols
Type2 rows4 rows8 rows16 rows
BF16n/a1.34x2.83x3.63x
Q1_01.27x2.33x4.41x3.38x
Q2_Kn/a2.13x4.30x3.64x
TQ2_0n/an/a1.71x4.03x
IQ2_XS1.07x1.91x3.48x3.67x

Across the 9 to 16 row range the PR reports 3.06x to 4.14x for the newly covered types. Types that already had these kernels stayed within about 3 percent of master, and everything was within 1 percent at 512 rows, so large-batch prompt processing is unaffected. The author notes some run-to-run variance on the machine and reports test-backend-ops passing 3,314 of 3,314 MUL_MAT and MUL_MAT_ID cases on the M3 Ultra, plus the new cases passing on CUDA and Vulkan hardware. One cost: the mul_mv_mma library takes longer to compile, from 0.33 s to 0.67 s on an M5.

Does this make speculative decoding 3.4x faster?

No, and this is the part to keep straight. Three layers separate a kernel speedup from the tokens per second you see.

  1. Kernel versus model. A matmul is only part of a decode step. Attention, sampling, the draft model's own passes and memory traffic are not sped up by this change.
  2. Acceptance rate. Speculative decoding only wins if the big model accepts enough drafted tokens. A poor draft pair can lose even with a perfect kernel.
  3. Quantization. The biggest ratios in the table are on aggressive formats such as Q1_0, Q2_K and TQ2_0. If you run Q4_K_M or Q8_0, the gain is the smaller, dedicated-kernel kind.

Independent measurements already showed how uneven Metal results can be. A recent academic study, "Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware", reports a best-case 1.61x wall-clock speedup with three of five configurations slowing down, and attributes part of that to the quantized Metal backend running "parallel" verification serially. That is exactly the behavior these kernels target, so retesting after b11404 is worthwhile. A separate GitHub issue against an older build, b9330, reported multi-token-prediction drafting being up to 28 percent slower than baseline on Metal. That predates the new kernels, so it is a good candidate to rerun rather than a current verdict.

How to test it on your Mac

Pull a build at or after b11404 and compare with and without drafting on the same prompt.

bash
brew install llama.cpp   # or build from source at tag b11404 or later
llama-server -m big-model-Q4_K_M.gguf \
  -md small-draft-Q8_0.gguf \
  --draft-max 8 --draft-min 2 -ngl 99

Flag names change often in llama.cpp, so confirm with llama-server --help on your build. Then measure:

  • Run the same 500-token generation with the draft model disabled and enabled.
  • Record generation tokens per second and the draft acceptance rate the server logs.
  • Try draft lengths of 4, 8 and 16, since the new kernels cover 2 to 16 rows and the best value depends on acceptance.
  • If you maintain a regression suite, run test-backend-ops for MUL_MAT to confirm correctness on your device.

If you see a slowdown, check acceptance before blaming the kernel. A draft that is wrong half the time can cost more than it saves.

Who should care

A small house-shaped box with a glowing green core, representing a local LLM running on a home machineA small house-shaped box with a glowing green core, representing a local LLM running on a home machine

Mac Studio and MacBook Pro owners running big local models are the clearest winners: the slower the base model, the more a successful draft saves. That includes the very large mixture-of-experts models that people now run on high-memory Macs, as in our write-ups of antirez's DwarfStar ds4 local inference engine and the local AI buying guide for the Surface Laptop Ultra.

People who run heavily quantized models (1 to 3 bit) get the largest relative boosts because those were previously the least optimized on Metal.

People who only chat with a model at one request at a time should expect little change unless they enable drafting.

Tool builders that wrap llama.cpp, such as desktop apps, should watch for bundled-version bumps, because the gain arrives automatically when they update. Our llama.cpp decision model support post shows how fast the project's feature surface moves, and the Transformers GGUF parity update covers how its file format is spreading.

Caveats

  • Kernel benchmarks come from the PR author, on one machine, with acknowledged variance. We have not independently reproduced them.
  • b11404 is flagged a pre-release. Pin a version in anything you ship.
  • The kernels apply to Apple7+ GPUs without the tensor API. Newer chips with tensor support may take a different path.
  • Compile times for the Metal library roughly doubled in the author's test.

Details above are accurate as of October 10, 2026. Check the llama.cpp releases page for current builds.

Related reading

  • Gemma 4 26B A4B MLX Mac speedup
  • Perplexity Lily: Apple Silicon inference engine
  • DeepSeek DSpark: speculative decoding for V4
  • DwarfStar ds4 local inference
  • llama.cpp decision model support
  • Surface Laptop Ultra local AI configs
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 19, 2026

DFlash-MLX Brings Lossless Speculative Decoding to Apple Silicon — Up to ~189 tok/s on M5 Max

bstnxbt/dflash-mlx ports DFlash block-diffusion speculative decoding to MLX on Apple Silicon with lossless greedy verification. On an M5 Max 64GB, Qwen3.5-4B jumps from 54 to 189 tok/s at 2048 tokens, while Qwen3.5-27B-4bit lands at 70 tok/s — the headline "~70 tok/s" figure — at roughly 2.1x over baseline. Speedup varies sharply by model size, architecture, and context length.

Jul 30, 2026

TurboFieldfare: Gemma 4 26B in ~2 GB RAM on Apple Silicon

Andrey Mikhaylov’s TurboFieldfare keeps Gemma 4’s shared core and KV in RAM and preads routed experts from disk — ~2 GB process RSS for a ~14 GB model. explainx.ai covers how it works, scores, macOS 26 limits, and when to use MLX instead.

Oct 10, 2026

Underdog Saluki 27B: Qwen3.8-27B Squeezed Into 7.89 GB With Tool Calling Intact

ConwayResearch has published Underdog Saluki 27B 1.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB instead of 54 GB. Its model card says it keeps 96 percent of the full model across nine benchmarks and beats it on tool calling. Here is how to run it, what the numbers do and do not show, and where it clearly loses.