Jeff hit the Hacker News front page on 28 September 2026 as "Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms" (item 49883844, 289 points and 117 comments when we pulled the thread on 29 September 2026). The linked repo is firelex/jeff — 363 stars and 9 forks on that same morning, created 28 September 2026. This is not another write-up of Kev, OpenJev, or the two-day clone wave. It is the home-trained Qwen3.5 / Gemma checkpoint family that speaks TypeSafe Jev's request format without being TypeSafe.
The author's own launch comment is the right briefing: open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification; calibrated probabilities in one forward pass; no text generation; Apache 2.0 weights; Jev-compatible API; not affiliated with TypeSafe. Training ran on one RTX PRO 6000, synthetic data from Qwen3.8-Flash-Next on two DGX Sparks, tests on a MacBook. explainx.ai checked the GitHub README and Hugging Face cards against that HN paste. Where they agree, we cite the repo. Where HN rounded, we use the README.
TL;DR: What people are asking
| Question | Direct answer |
|---|---|
| Is this just structured outputs / constrained decoding / logprobs? | No for JSON-mode and constrained decoding. Not "just" logprobs either: trained readout + full-weight fine-tune + fitted temperature. |
| vs BERT / FLAN-T5 / ModernBERT / embeddings+logistic? | Those are usually fixed-label classifiers (or a separate encoder+head). Jeff's options are described at request time (up to 255). Embeddings+logistic can still win on labeled, high-class-count tasks — HN said so. |
| vs hosted Jev vs Kev vs Laya vs OpenJev? | Jev = cloud API. Jeff = local Jev-shaped HTTP. Kev = Palmer LoRA family. Laya = encoder typed-decision + MLX port. OpenJev = in-browser readout vs generation. |
| Why does 0.8B beat 2B? | Author: 2B more risk-averse; training amplified it. Panel accuracy still favors 2B. |
| Zero-shot vs fine-tune? | Voice-nav held-out: 31.7% → 95.8% in under ~30 minutes on one GPU (autojev-train --initial-checkpoint). |
| Price/speed vs DeepSeek Flash? | Jeff has no token invoice once weights are local (~22 ms RTX PRO 6000, 28 ms M4 Max MLX for 0.8B). Flash is a generative API you pay per token. Do not compare them as the same product. |
| Honest 70% vs 94% / job ads? | Real HN reports. Treat panel 83.1% as not your job-ad accuracy. |
What Jeff actually is (repo, not digest)
Project jeff version 0.2.0 (pyproject). GitHub description: "Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification." Hugging Face cards (publisher mstrasser):
| Checkpoint | Base | Panel accuracy (Jeff) | ECE | ~16-bit size |
|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | Qwen/Qwen3.5-0.8B | 79.1% | 0.049 | 1.7 GB |
| Jeff-Qwen3.5-2B | Qwen/Qwen3.5-2B | 83.1% | 0.028 | 4.2 GB |
| Jeff-Gemma4-E2B | google/gemma-4-E2B-it | 81.6% | 0.031 | 9.3 GB stored (2B effective) |
Jev's published five-benchmark overall is 83.0%. The README is explicit: Jev and AutoJev figures were measured on a different sample of the same benchmarks. Do not treat 83.1 vs 83.0 as a controlled win.
Training recipe (README): full-weight supervised fine-tuning, one epoch, batch 256, cross-entropy over option letters, then one fitted temperature for calibration. Checkpoints are chosen on a development set, never on the published panel. That is a different stack from Kev's LoRA adapter and pointer head. Code is MIT (copyright Mathias Strasser / Jeff, plus Denis Yarats / AutoJev). Model weights: Apache 2.0. Training data is not released; sources and licences live in the repo's docs/data-sources.md.
A System One Model here means choice / noul / score over enumerated options, not a chat LLM.

Is this just structured outputs / constrained decoding / logprobs?
Structured outputs / JSON mode: no. Jeff's README: "No generated text, no parsing." The server exposes POST /v1/systemone, not a chat completions grammar.
Constrained decoding: no. There is no token-by-token sampler constrained to a JSON schema. The model scores option codes in one forward pass.
Logprobs on a frozen chat model: closer, still incomplete. The AutoJev/Jeff design uses a trained answer readout over single-token option codes (initialized from output embeddings in the generic decoder path). You are reading a distribution over a small label set. You are not wrapping GPT-style generation and scraping logprobs. Weights are fine-tuned; a temperature is fitted for calibration. Calling that "just logprobs" is like calling a fine-tuned BERT "just softmax." Related explainx.ai frame: Jev vs XGBoost and BERT.
vs BERT, FLAN-T5, ModernBERT, embeddings + logistic
HN commenter nico (replying to the 70%-vs-94% thread) reported embeddings plus a logistic head matching or beating Jev and Laya on AG News, Emotion, MASSIVE Intent, and Banking77, especially on taxonomies with more than 50 classes, at under 1 ms, and failing on reasoning-ish sets like XNLI where Jev/Laya look stronger.
That is the honest split:
- Labeled, stable taxonomy (job industry, Banking77): a dedicated classifier or encoder often wins on accuracy and latency. Jeff's 0.8B failing job-ad fields is consistent with that.
- Runtime-described options (this ticket's three queues, this turn's legal Doom moves): you need something that does not freeze a softmax width at train time. That is the System One pitch Jeff, Jev, and Kev share.
- ModernBERT / Laya: encoder, PPO-trained Laya is a different base than Qwen3.5 decoder fine-tunes. See the Laya-MLX write-up for Apple-silicon encoder numbers, not Jeff's Qwen MLX path.
- FLAN-T5: still typically generates (even if short). Jeff is trained not to decode a continuation.
Jeff vs hosted Jev vs Kev vs Laya vs OpenJev vs AutoJev
| Project | What it is | What Jeff is not |
|---|---|---|
| Hosted Jev | TypeSafe cloud System One API, undisclosed size, RLCD branding | Jeff is local weights + compatible shape |
| Jeff | Home-trained 0.8B/2B/Gemma-E2B, AutoJev-derived SFT | Not TypeSafe; not 27B AutoJev |
| AutoJev | Open recipe (Denis Yarats) fine-tuning Qwen3.8-27B; Jeff README cites published AutoJev-27B panel 84.9% | Jeff kept the recipe and shrunk the student |
| Kev | Palmer 0.8B/4B/9B Qwen3.5, LoRA + pointer head, kev.train --init_from | Different training; different sizes |
| Laya / laya-mlx | ModernBERT-scale typed decisions; MLX port claims sub-14 ms on M3 Max | Different architecture; Jeff's 0.8B is 28 ms on M4 Max in Jeff's table |
| OpenJev | Browser WebGPU, compare readout vs generation | Not a trained Jeff checkpoint |
Speed table from the README (median over 200 ~200-token questions, one at a time, text to probabilities):
| Model | NVIDIA RTX PRO 6000 | Apple M4 Max (MLX) |
|---|---|---|
| Jeff-Qwen3.5-0.8B | 22 ms | 28 ms |
| Jeff-Qwen3.5-2B | 24 ms | 60 ms |
| Jeff-Gemma4-E2B | 29 ms | — (MLX is Qwen-only in this repo) |
| Jev (published Doom) | 114–212 ms per API call including network |
Those Jev times are not the same hardware as the Mac numbers. Jev's 200x/400x marketing is a separate fact-check; Jeff does not inherit those multipliers.
The benchmark table you should actually read
4,599 questions across five public sets, plus JevBench hard (105 items, scored separately). Numbers are the project's.
| Benchmark | Jeff 0.8B | Jeff 2B | Jeff Gemma E2B | Jev published | AutoJev-27B published |
|---|---|---|---|---|---|
| Overall (5) | 79.1 | 83.1 | 81.6 | 83.0 | 84.9 |
| BBH | 64.0 | 68.0 | 66.4 | 94.3 | 82.8 |
| Financial PhraseBank | 96.4 | 96.3 | 96.1 | 77.0 | 84.2 |
| JudgeBench | 62.6 | 64.6 | 60.6 | 78.6 | 78.9 |
| RAGTruth | 86.1 | 88.9 | 87.4 | 77.3 | 88.9 |
| WinoGrande | 68.6 | 79.0 | 77.4 | 90.7 | 83.3 |
| JevBench hard | 47.6 | 53.3 | 48.6 | 73.3 | 70.3 |
HN's "BBH 64–68% vs Jev 94%" and "JevBench hard ~50% vs 73%" match this table (0.8B hard is 47.6%, 2B 53.3%). Overall closeness is classification and grounding; reasoning is where the small students lose, which is what you should expect from 0.8B–2B versus a larger hosted model. For how to read self-tables, see how to read AI benchmarks.
Why 0.8B beats 2B, wording, zero-shot vs fine-tune
Games (20 episodes, seed 1234), README:
| Model | Doom kills | Frogger crossings | Pac-Man pellets / 98 |
|---|---|---|---|
| Hand-coded rule bot | 6.55 | 10.25 | 94.1 |
| Jeff-Qwen3.5-0.8B | 6.55 | 10.3 | 57.0 |
| Jeff-Qwen3.5-2B | −0.9 | 6.0 | 41.2 |
| Untrained Qwen3.5-0.8B | 5.0 | 1.0 | 25.8 |
| Jev published Doom | 6.55 with aiming rule; −0.60 without | — | — |
The 2B is worse than the 0.8B on Doom and Pac-Man after training. The author: bigger is not better for "fast option picking"; the 2B is "more cautious." Untrained Gemma 4 E2B beats untrained Qwens on the panel (62.5%) and plays games worst.
Wording: giving Frogger's goal step the same language as other forward options moved one episode from 15 to 23 crossings. Jev's Doom prompt (bearing number + aiming rule) does not work for Jeff; consequence-in-words options do.
Fine-tune: voice navigation, ~11k app-specific examples, ~half an hour, one GPU, held-out 31.7% → 95.8%, ~40 ms on M4 Max. Command shape in the README: autojev-train --initial-checkpoint pointing at a Jeff checkpoint, --epochs 1. Zero-shot is the demo; domain data is the production path — same lesson as Kev's --init_from, different trainer.
Price and speed versus DeepSeek Flash (and hosted Jev)
Jeff: electricity + disk. Optional JEFF_API_KEY on the local server. TypeSafe's list price in explainx.ai's Jev coverage is $42 per billion input tokens, output free — a usage meter. DeepSeek Flash is a full generative API (explainx.ai has covered its token-volume and coding-benchmark stories separately). You do not "replace Flash" with Jeff unless your workload is typed decisions, not code/chat reasoning.
Rough practitioner math: a 200-token Jeff decision at 28 ms locally has zero Flash/Jev invoice. A Flash call that writes a paragraph is a different SKU. If you were using Flash as a classifier, Jeff (or BERT, or embeddings+logistic) is the category to A/B, not "is Flash faster than 28 ms" without matching the task.
HN asked for a "price comparison for the masses." The accurate answer is: Jeff's marginal cost is your GPU/CPU; Jev's is TypeSafe's meter; Flash's is DeepSeek's meter and generation length.
Honest HN: 70% vs 94%, job ads, "they all suck"
Do not sand this down.
- AgentMasterRace: compared Jeff to Jev on current use cases — "70% vs 94%. for classification, it's unacceptable."
- Oras: job ads (industry, remote/hybrid/onsite, full-time/part-time). Side-by-side Gemini 2.5 Flash Lite, Jev, Jeff. "0.8B model, completely useless." Jeff-Qwen3.5-2B better, still missed job type.
- zergrush: "i've tried all the open source me too Jevs they all suck."
- tbeseda: the point is an open-weight MVP you can fine-tune, not displacing Jev on day one.
- senko: ModernBERT ~0.4B already fine-tunes; Jev's value is zero-shot without fine-tuning.
- clhodapp: needing fine-tuning changes the product category.
- ijustlovemath / neuronexmachina: "isn't Jev just a less nuanced classifier?" — "zero-shot is what makes these different."
The README already said small models "won't match Jev" on reasoning and that you should fine-tune if zero-shot is not enough. The thread is that warning with production labels attached. If you wire Jev into agent routing, do not drop in Jeff-0.8B without the same eval set.
How to call the Jev-compatible API locally
From the official README (NVIDIA/CPU). Apple silicon: uv sync --extra mac and JEFF_BACKEND=mlx.
uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"model": "jeff-latest",
"state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
Question types: choice (up to 255 options), noul (yes/no as a probability), score (a scale you describe). Several independent questions in one request. README ops advice: reason in code, decide with Jeff; short option keys; describe consequences, not forecasts ("a car arrives in 2 turns" is random-level).
What this means for builders
If you need local, open-weight, Jev-shaped HTTP and you can measure your own labels, Jeff is a real checkpoint family with an unusually honest README (games vs panel, 2B vs 0.8B, wording). If you need hosted calibration and zero-shot on hard reasoning, Jev is still the product HN defenders called "king." If your taxonomy is fixed and labeled, try the classifier you already know before a 0.8B decoder.
Jeff does not replace the six clones from launch week; it is a later, AutoJev-student, train-at-home data point in the same category.
Related reading
- Kev: open-source Jev clone, Qwen3.5 0.8B/4B/9B
- TypeSafe AI launches Jev
- OpenJev: browser Jev clone
- Laya-MLX on-device alternative
- Six Jev clones in two days
- What is a System One Model?
- Jev vs XGBoost and BERT
- Jev speed/cost claims fact-check
- Official: firelex/jeff · Hugging Face
mstrasser/Jeff-Qwen3.5-0.8B· HN item 49883844
Figures, licences, and API snippets are from firelex/jeff and the mstrasser Hugging Face cards as of 29 September 2026 (repo created 28 September 2026). Hacker News score 289 / 117 comments and GitHub 363 stars were live API reads that morning — they will move. explainx.ai did not retrain Jeff or re-run the 4,599-question panel.
