Most speech-to-text models are too big for the devices that need them most: watches, earbuds, robots, car dashboards and cheap boards. On October 2, 2026, Cactus Compute released Whistle, a speech recognition model that fits in 16.9 MB and runs on a CPU.
The release is not only a small model. Whistle shares an engine with Cactus's Needle model, so one binary can turn a voice clip into tool calls. This post covers how Whistle works, the benchmark results including where it loses, how to run it, and how it compares with other transcription options on explainx.ai. We read the Hugging Face model card, the Cactus announcement and the Needle 3 card. We did not run Whistle, so the numbers are Cactus's own.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A 16.9 MB open speech-to-text model for seven languages |
| License? | Apache 2.0 |
| Does it need a GPU? | No. CPU only, no dependencies |
| Speed? | 11.1 ms to first token and 1,319 tokens per second on an Apple M4 Pro CPU, per Cactus |
| Better than Whisper base? | Wins on some benchmarks, loses on TED-LIUM, AMI and MLS |
| Languages? | English, German, French, Spanish, Italian, Dutch, Polish |
| Longest clip? | 30 seconds per pass |
| Can it act on speech? | Yes, with Needle 3 in one call |
What Whistle is
The model card describes "a speech-to-text model for mobiles, wearables, robots, smart home, automotive and microcontrollers." The whole model is "a single 16.9 MB file, and it runs on the same CPU engine as Needle, from the same container and the same quantisation, with no dependencies and no GPU."
It does three jobs on the device:
- Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in seven languages. The language is detected unless you name it. Silence returns an empty transcript.
- Word timestamps. Every word with a start, an end and a probability, aligned from the decoder's own attention. An app can highlight, seek or cut on a word.
- Speech embedding. The encoder output, one row per 80 ms frame, for matching and retrieval without decoding a transcript.
It also supports keyword biasing. You pass names, places and product words, and the search favors them so the ones that matter survive.

Figure: Whistle at a glance. Source: the Whistle model card by Cactus Compute (Apache 2.0), used with credit.
How it works
The Cactus announcement, written by Jakub Mroz and Henry Ndubuaku, gives the architecture in detail.
| Stage | Detail |
|---|---|
| Front end | 16 kHz mono audio, 25 ms window, 10 ms hop, 80 log-mel bins, band-limited to 250 to 3500 Hz |
| Stem | A convolutional stem of 128 channels and kernel 9 halves the frame count three times, leaving 375 frames for 30 seconds, one per 80 ms |
| Encoder | Eight Simple Attention blocks with four mHC residual lanes and a Monarch Hadamard MLP. Attention is not causal |
| Decoder | Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, engram lookups at layers 3 and 7 over 18,432 slots |
| Cross attention | A gated cross attention at each decoder layer reads the encoder. K and V are computed once per clip |
| Decoding | Five beams scored by length-normalized log probability. Cap of 320 tokens. Vocabulary of 8,192 text pieces plus seven language tokens |
| Quantization | 2 to 4 bit with Cactus Quants |
Three design choices matter for developers.
One engine for speech and text. Whistle reuses Needle's .cact container, Cactus Quants, SIMD kernels and KV cache. A device that already runs Needle runs Whistle in the same binary, with no second runtime.
A decoder ladder. Every decoder depth from 2 layers up was trained as a model of its own. You pick one at load time with --audio-depth. The encoder is never sliced: all eight blocks run at every depth. That gives you a knob to trade accuracy for speed on a weak chip.
Cheap beams. Cross-attention keys and values are computed once when the clip arrives and then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.
Benchmarks, including where Whistle loses
Cactus compares Whistle with Whisper base and Moonshine tiny v2. Word error rate (WER) is scored with the Whisper normalizers. Whistle's WER is measured over 86,174 utterances. The Whisper and Moonshine figures are the ones their authors published, from the multilingual checkpoints.

Figure: Whistle against Whisper base and Moonshine tiny v2. Source: Cactus Compute, from the Whistle model card, used with credit.
Whistle's own WER values in the chart:
| Benchmark | Whistle WER (%) |
|---|---|
| LibriSpeech test-clean | 4.31 |
| LibriSpeech test-other | 10.49 |
| SPGISpeech | 7.65 |
| Earnings-22 | 19.01 |
| AMI | 26.07 |
| AMI cleaned | 22.87 |
| TED-LIUM | 7.61 |
| FLEURS (average of seven languages) | 21.4 |
| MLS (average of six languages, no English) | 24.9 |
In its own words, Cactus says: "Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9."
That is a fair reading of the chart. Whistle is not a clean win. It loses on meeting speech (AMI), on TED talks and on the multilingual MLS set. The Whisper figure for AMI is AMI-IHM, a different subset from the AMI the other two report, so that row is not directly comparable.
Speed and size
On 10 seconds of audio on an Apple M4 Pro CPU, with each model on its official runtime at its defaults:
| Metric | Whistle | Whisper base | Moonshine tiny v2 |
|---|---|---|---|
| Size | 16.9 MB | 145.3 MB | 41.9 MB |
| Time to first token | 11.1 ms | 73.2 ms | 22.8 ms |
| Decode speed | 1,319 tokens/s | 266 tokens/s | 262 tokens/s |
Whistle ran at 5 beams, Whisper through openai-whisper, and Moonshine through moonshine-voice in non-streaming mode over whole audio. Cactus notes that Whisper pads every input to 30 seconds, so its time to first token stays flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds.
Read the speed numbers with care. They are for one Apple laptop CPU. A phone, a watch or a microcontroller will run slower, and the card gives no figures for those. Precision also differs: Whistle runs at 2 to 4 bit, Whisper base in fp32 in memory, and Moonshine in int8.
The model card states that no test audio appears in Whistle's training or validation data. Cactus says it checked this by comparing audio checksums and speaker IDs across every reported test set.
How to try it
Install the Python package:
pip install cactus-needle
Then transcribe a clip:
import needle
print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lights
Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second. Useful options:
| Option | Effect |
|---|---|
word_timestamps=True | Adds each word with start, end and probability |
keywords=[...] | Favors names and product terms during search |
language="de" | Forces the language instead of detecting it |
needle.Whistle() | The model as an object for embed(audio) or a tuned .cact |
The [mic] extra adds microphone capture and other sample rates. A 16 kHz WAV or raw samples need nothing beyond the base install. The command needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs one clip through Whistle, Whisper and Moonshine side by side with timings.
The command-line engine
needle download macos-arm64
needle download whistle
./macos-arm64/needle --model whistle.cact --audio clip.wav --audio-word-timestamps
Per the announcement, the engine ships prebuilt for seventeen targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. The speech C API is three functions: needle_load, needle_transcribe and needle_embed. The engine reads no environment variables, so every run of the same file scores the same.
The announcement page also hosts a browser sandbox. The first press downloads the 16.9 MB model, and Cactus says audio never leaves your device.
Voice to tool calls in one binary
The part that sets Whistle apart is the pairing with Needle 3. The Needle card calls it "a foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers," in a single 8 to 29 MB file. It picks functions from your tool list and fills arguments.
Load both models:
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
The engine transcribes the clip, answers the transcript against your tools and returns one JSON object. A response looks like this:
{ "function_calls": [{ "name": "set_lights", "arguments": { "room": "kitchen", "on": false } }],
"confidence": 0.94,
"audio_text": "turn off the kitchen lights",
"audio_language": "en" }
The caller never handles the transcript. For a smart-home panel or a robot, that is one fewer hop and one fewer place for latency and bugs. The confidence field lets you route on act, confirm or refuse.

How it compares
| Option | Where it runs | Notes |
|---|---|---|
| Whistle | CPU, on device, 16.9 MB | Seven languages, 30 s clips, pairs with Needle for tool calls |
| Whisper base | CPU or GPU, 145.3 MB | Broader language coverage, ahead on TED-LIUM, AMI and MLS in Cactus's chart |
| Moonshine tiny v2 | CPU, 41.9 MB | English only |
| Superwhisper S1-mini | On device | A 0.6B cleaner for transcripts, not an ASR model |
| MAI-Transcribe-2-Streaming | Cloud API | Real-time streaming, ranked on Artificial Analysis WER |
| Gemini 3.5 Transcribe | Cloud API | Hosted, large-model accuracy |
| Meta Muse Voice Transcribe | Cloud | Real-time ASR with diarization and endpointing |
| Interfaze-1-Lite | GPU class | Open weights for OCR, speech and extraction together |
The choice is about where audio may go. A cloud model gives higher accuracy on hard audio and many languages. Whistle gives privacy, no per-minute fee, offline use and a tiny footprint. For a voice assistant on a Raspberry Pi or a wearable, those traits decide the choice. For call-center transcription of noisy multi-speaker audio, they do not.
What this means for what you build or pay
- Per-minute fees can drop to zero for short commands. Voice control, dictation of short notes and wake-then-command flows fit in 30-second clips. Local inference removes the API bill and the network round trip.
- Privacy is easier to promise. Audio stays on the device. The Cactus sandbox states this for its own demo, and the same holds for any app that runs the model locally.
- Plan for the weak spots. If your users speak in meetings, on TED-style stages or in languages beyond the seven, test hard before shipping. Whistle trails Whisper base on those sets in the vendor's chart.
- One engine simplifies a voice agent. Speech and tool calling share a runtime, a file format and a quantization, so a device carries one binary instead of two stacks.
For a deeper map of building voice agents with open tools, see Hugging Face's speech-to-speech guide.
Limits and open questions
- Clip length. 30 seconds per pass. Longer audio needs chunking, and the card does not describe a streaming mode.
- Languages. Seven only.
- Vendor-measured WER. Whistle's results are Cactus's own, while the baselines come from their papers, which use different test conditions in places.
- Device coverage. The speed figures are for an Apple M4 Pro CPU. We found no numbers for phones or wearables.
- Download counts. The model page showed about 2,600 downloads and 196 likes on October 9, so community testing is early.
- Hacker News. The Cactus post appeared on Hacker News on October 2 with no comments when we checked, so we have no developer reports to quote.
- Repository access. The
Cactus-Compute/needle3repository, which holds the engine, returned a 404 on GitHub when we fetched it. The Hugging Face repository of the same name loads fine, so we relied on the model cards and the announcement.
Summary
Whistle is a credible tiny ASR model with a clear niche: private, offline, low-latency transcription of short clips on weak hardware, and a one-binary route from voice to tool calls. It is not a Whisper base replacement everywhere. Test it on your own audio, especially meeting speech and non-English languages, before you rely on it.
Related reading
- MAI-Transcribe-2-Streaming and the Artificial Analysis WER ranking
- Superwhisper S1-mini: on-device transcript cleaning
- Gemini 3.5 Transcribe: what builders get
- Meta Muse Voice Transcribe
- Hugging Face speech-to-speech voice agent guide
- Interfaze-1-Lite open weights for OCR and speech
Primary sources: the Whistle model card, the Cactus announcement and the Needle 3 model card.
Specs, benchmark figures and download counts are accurate as of October 9, 2026. Re-check the model card before you build on it.
