Underdog Saluki 27B 1.0 is a 2-bit GGUF build of Qwen3.8-27B that fits in a single 7.89 GB file, down from roughly 54 GB for the full-size model. Its author, ConwayResearch, published it on Hugging Face under Apache 2.0 and says it was tuned to keep tool calling intact. On the card's own 120-task tool-calling test, it passes 88 tasks against 84 for the uncompressed original.
That is a surprising result for a model squeezed to about a seventh of its size, and it is worth reading with some care. This post covers what the model is, how to run it with stock llama.cpp, what the benchmark table does and does not prove, and who should reach for it.
TL;DR: questions people will ask first
| Question | Answer |
|---|---|
| What is it? | Qwen3.8-27B compressed to a 2-bit "IQ2-mix" GGUF, 7.89 GB |
| Who made it? | ConwayResearch (the card credits "Underdog"), built on the Qwen team's model and ISTA-DASLab's quantized base |
| License? | Apache 2.0, same as both bases |
| Does it need special software? | No. Stock llama.cpp and apps built on it |
| Does it do images? | Yes, with an optional 629 MB or 928 MB vision add-on file |
| Is it faster? | Unknown. The card publishes no tokens-per-second figures |
| Does it beat the full model? | On tool calling, by the card's test. On math and long reasoning, no |
| Downloads so far? | 15,274 in the last month and 184 likes at the time of writing |
What exactly was released
The repository holds one main file, Underdog-Saluki-27B-1.0-IQ2-mix.gguf, plus two optional vision projector files. The model card describes the base as Qwen3.8-27B and names ISTA-DASLab's Qwen3.8-27B-GSQ-RCO-GGUF as the quantized derivative it builds on. The architecture tag is qwen35 and the model is 27 billion parameters.
The card is vague on method. "IQ2-mix" signals a mixed 2-bit scheme in the llama.cpp family of quantization formats, and the repository tags include imatrix, which usually means an importance matrix was used to decide which weights deserve more precision. But the card does not describe the mix, the calibration data, or what the "tuned to keep tool calling intact" step actually involved. That is the main gap for anyone who wants to reproduce the result.
For background on the base model family, see our coverage of Qwen3.8 Flash and the Qwen 4 architecture preview and the Qwen3.8 Flash Next 125B MoE release. The 27B size is the one hardware vendors keep showing off: Cerebras ran Qwen3.8 27B at about 1,500 tokens per second, and the Oki Home memory computer ships with a local copy.
How to run it
The model card gives two commands. First, download the file:
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .
Then serve it with llama.cpp:
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
The --jinja flag matters. It turns on the Qwen3.8 chat template, which is what handles tool calls and the thinking mode. Without it you lose both. The flags -ngl 99 offload all layers to the GPU, -fa on enables flash attention, and -c 32768 sets a 32K context window. The server speaks the OpenAI chat API on port 8080, so most agent frameworks can point at it with a base URL change.
To add vision, download one of the mmproj files (928 MB in F16 or 629 MB in Q8_0) and pass it with --mmproj. The main file is text only.
The card recommends two settings profiles:
| Use | Settings |
|---|---|
| General use, reasoning, instructions | Thinking on (default), temperature 0.6, top_p 0.95, top_k 20 |
| Fast, direct tool calls | Thinking off via chat_template_kwargs with enable_thinking false, temperature 0 |
That second profile is how the tool-calling numbers below were produced, so use it if you want to match them.
The benchmark table, read carefully
A gripper choosing the right tool from a tray, illustrating tool calling in a small local model
The headline result is on "Underdog Bench," which the card describes as 120 tasks taken from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before any model was tested, run with thinking off at temperature 0.
| Model | Size | Passed (of 120) |
|---|---|---|
| Underdog Saluki 27B 1.0 | 7.89 GB | 88 |
| Qwen3.8-27B, full size | 54 GB | 84 |
| Bonsai 2 | 5.95 GB | 70 |
On 100 BFCL v4 parallel-call tasks scored with the official checker, Saluki gets 42 and the full model 35. These two results are the basis for the claim that tool calling survived compression.
Two cautions. The card itself says 120 tasks is a modest test and that a few tasks of difference is within run-to-run variation, so 88 against 84 is better read as "roughly on par" than "better." And these are the authors' own numbers on a set they selected; nobody independent has reproduced them yet. Still, parity on tool calling at one seventh of the size is the claim that matters for agent builders, and it is easy to test yourself.
Where it loses
The wider table is more sobering, and the card deserves credit for publishing it.
| Benchmark | Saluki | Full size |
|---|---|---|
| SWE-bench Verified (50 issues fixed) | 30 | 33 |
| IFEval (prompt-loose) | 93.5 | 91.5 (public) |
| IFBench (prompt-loose) | 72.7 | 71.0 (public) |
| MBPP+ | 78.0 | 83.9 (public) |
| MuSR | 67.5 | 79.6 (public) |
| AIME 2025 (avg@4) | 79.2 | 96.7 (public) |
| AIME 2026 (avg@4) | 80.0 | 94.6 (public) |
A test bench with a probe touching a block, illustrating benchmark evaluation of a quantized model
Instruction following holds up, coding drops a little, and competition math falls by around 15 points. The card puts the retention at about 82 to 85 percent of the full model's math score and says average retention across nine benchmarks is 96 percent. It also notes that the "public" scores for the full model come from a different harness than the one used for Saluki, so the comparison is not apples to apples; where the authors ran the full model themselves, they used the same harness and settings.
The card lists further limitations: it is weakest at letter-level puzzles such as palindromes and alphabetical ordering, about one in five parallel-call replies carry small formatting slips, and with thinking on it often reasons at length before answering. For a 2-bit model, none of this is shocking. Letter-level tasks are a known weak spot for tokenized models, and aggressive quantization tends to cost the most on long multi-step reasoning, which is exactly what AIME measures.
Who should use it
The sensible use case is a local agent that spends most of its time choosing and calling tools rather than proving theorems: a coding helper wired to a few commands, a home-automation controller, a document pipeline, or a private assistant that must stay on the machine. Our guide to agent skills and the explainer on MCP show the kinds of tool surfaces such a model would drive.
It is a poor pick if you need strong math, long chains of reasoning, or exact string manipulation. In that case a larger quantization of the same base, or the full model, is still the safer choice.
Memory is the other consideration. A file under 8 GB brings a 27B-class model within reach of a 12 GB graphics card or a laptop with 16 GB of unified memory, though the KV cache for a 32K context adds to that, and the card gives no minimum hardware spec. If you are weighing machines for this kind of workload, our Surface Laptop Ultra configuration guide for local AI walks through how much memory different model sizes need, and llama.cpp's decision model support covers recent runtime changes worth keeping current with.
How it compares with other extreme-compression efforts
Saluki sits in a crowded field of "small enough to run anywhere" work. Samsung's LittleBit pushes a 13B model under 1 GB with sub-1-bit weights, but it is a research paper result rather than a file you can load in llama.cpp today. Saluki's pitch is the opposite: a modest 2-bit level, a standard GGUF, and a benchmark focused on the one capability agents depend on. If you have been following the on-device privacy push, Underdog's private personal AI launch is a separate product post of ours; the model card does not say the two are connected, so we are not claiming they are.
What to check before you rely on it
- Run your own tool schema. BFCL tasks are not your tools. Test with your actual function definitions, including nested arguments and parallel calls.
- Watch for formatting slips. The card admits roughly 20 percent of parallel replies have small formatting errors. Add schema validation and a retry.
- Pick thinking mode deliberately. Off and temperature 0 for fast tool calls; on for harder reasoning, accepting the longer outputs.
- Measure speed. With no published tokens per second, benchmark on your hardware before building a latency budget around it.
- Check the NOTICE file. The card points to a NOTICE file for attribution to the Qwen team and ISTA-DASLab, both Apache 2.0.
Bottom line
Underdog Saluki 27B is a practical, openly licensed shrink of a strong base model, with an honest limitations section and a claim that is cheap to verify. Treat the tool-calling win as promising rather than proven, treat the math losses as real, and download it if you want a 27B-class agent brain in under 8 GB. The unanswered questions are how the "tuned to keep tool calling intact" step works and whether independent testers reproduce the 88 and 42 scores.
Model details, file sizes and benchmark figures are from the Hugging Face model card as of October 10, 2026 and may change as the repository is updated.
Related reading
- Qwen3.8 Flash and the Qwen 4 architecture preview
- Qwen3.8 Flash Next 125B MoE release
- Cerebras runs Qwen3.8 27B at 1,500 tokens per second
- Oki Home: a local Qwen 3.8 27B memory computer
- Samsung LittleBit: a 13B LLM under 1 GB
- llama.cpp decision model support
- What is MCP
- llama.cpp releases and the base Qwen3.8-27B model
