Supersonic Labs announced Julia 1 on September 27, 2026: a 144.3 million parameter decision model you can download and run on a laptop CPU this week. The lab's post on X had about 67,600 views. The weights are Apache 2.0, about 550.5 MiB, and the same graph is packaged for a WebGPU browser. The primary writeup is the lab's Julia 1 page. The files live in two Hugging Face repos: the Python CPU and CUDA runtime and the ONNX WebGPU build.
Readers who have been following Jev, Laya, and the same-day GLM-5.3-Flash one-token method now have a third artifact with a different shape. Julia 1 is the checkpoint. You install it and call it. The GLM piece leaves a generative model as shipped and scores a closed list from one token. Jev remains the hosted System One Model from the September 15 launch. This post does not re-derive RLCD. That training story is already in how Jev works.
Figures below are Supersonic Labs' measurements, as of September 27, 2026. explainx.ai did not re-run them.
TL;DR
| Question | Answer |
|---|---|
| What is it? | A decision model. Context, a question, and 2 to 20 answers in. A choice and scores out, in the order the options were supplied. Classification, routing, ordered scales, and yes/no. |
| Can I run it locally? | Yes. Python 3.11+ on CPU, or the ONNX graph in a WebGPU browser. CUDA is available when PyTorch sees a BF16-capable GPU. |
| Does it beat Jev? | On one 2,000-item suite, by 0.45 points (73.15% vs a 72.70% reference). Banking77 goes the other way: 64% vs 87% on 100 examples. |
| How many options? | 2 to 20 on a native call. A Router can narrow a longer list, and that narrowing can delete the correct label. |
| License? | Apache 2.0 for the released artifacts. The training pipeline is not in the repository. |
| What does a call cost? | Planned API price: $0.025 per million input tokens and $0.00 per million output tokens. Local weights have no per-token meter. |
| Is the API live? | No. The lab says access will open later. This week the path that exists is the download. |
What you actually download
Julia 1 starts from mmBERT-small, a multilingual ModernBERT encoder and tokenizer from Johns Hopkins CLSP. Supersonic Labs says training a multilingual foundation from scratch was beyond the budget of this first release, so they kept that encoder and tokenizer, added a decision head, and trained the checkpoint to score answers supplied with a question. The released model is 144.3 million parameters, decision components included. Supersonic Labs also says Julia 1 is not a fine-tuned Qwen. Their cloud GPU spend on training and experiments was about R$540 (US$104.08), their figure. Julia 2 is planned on an architecture of their own and is still in development.
The interface is small on purpose. You pass context, a question, and between 2 and 20 answers. The model picks one and returns scores in the order you supplied the options. The same call covers a classification label, a routing destination, an ordered scale such as low / medium / high, or a yes/no. On the model card those three shapes are choice, score, and noul. A choice maps your ids to descriptions and returns the winning id. A score takes an ordered rubric and returns an expected index. A noul is the probability of true, with false first and true second when you use the older list form.

That is the whole product. There is no chat reply to parse, and there is no new output head for each workflow. A new label set is a new list of options. The category explainer covers why a small, enumerated output space changes latency compared with decoding text. Julia 1 is one open checkpoint in that category. It is also a first public test of Supersonic Labs' training setup, which is why the lab publishes the miss next to the wins.
How Julia 1 compares with Jev, Laya, and one GLM token
The useful comparison is the artifact, not a single accuracy cell. Four ways to get a typed decision are in circulation this month. They do not share weights, hardware, or a referee.
| Weights you can download | Where it runs | Images | Option-list behavior | Who measured accuracy | |
|---|---|---|---|---|---|
| Julia 1 | Yes. Apache 2.0, 144.3M, about 550.5 MiB | Python on CPU or CUDA. Browser WebGPU from the ONNX repo | Text in the published evals | 2 to 20 per native call. Longer lists go through a Router that can drop the right label | Supersonic Labs. explainx.ai did not re-run them |
| Jev | Hosted. explainx.ai's coverage treats it as an API | TypeSafe's service | No image input in the launch coverage | Choice, score, or yes/no on answers you supply. The Jev column in Julia's table is a protocol reference | The protocol's reference values, not a new Jev run on the same machine |
| Laya and Ollaya | Laya checkpoints ship with the local runtimes. Ollaya serves other people's models | Laya on Apple MLX. Ollaya on macOS, Windows, and Linux | Text in the comparisons explainx.ai has covered | Laya option names share a token budget. The GLM writeup reports 192 tokens, enough for Banking77 and short of CLINC150 | Project latency tables, and a different accuracy suite from Julia 1 |
| GLM-5.3-Flash, one token | The generative checkpoint, used as shipped | Wherever you already serve that model | Yes. The same chat request can carry images | One token per option index the tokenizer spells as a single token. Their stated ceiling is 191, with a 128-logprob serving cap | Privatemode, September 24, 2026. explainx.ai did not re-run it |
If the job is "put a classifier on a laptop tonight," Julia 1 and the Laya/Ollaya path are the ones with files. Ollaya is a local server for several open decision models, Laya among them. Julia 1 is a single 144.3M encoder with its own Python package and a browser build. If the job is "I already serve GLM-5.3-Flash and I need images," the one-token method is the one that accepts a scan. Jev is the hosted meter, with output tokens priced at zero in TypeSafe's own launch pricing, covered in the launch post. Julia 1's planned meter is lower on input ($0.025 per million tokens versus the $0.042 per million figure in that launch coverage) and also prices output at zero, and it was not switched on with this release.
Does Julia 1 beat Jev?
On the headline cell, barely, and only against a reference column. Supersonic Labs' September 24, 2026 evaluation used H200 BF16 inference and strict encoding. The Jev numbers are the comparison protocol's reference values. The lab says they are not a new Jev run, and they are not a same-hardware bake-off. Abstentions count as errors.
| Task | Julia 1 | Jev reference |
|---|---|---|
| Typed decisions, 2,000 items | 1,463 / 2,000 = 73.15% | 72.70% (+0.45 points) |
| AG News, 4 labels, 100 examples | 94 / 100 = 94% | 91% |
| DAIR Emotion, 6 labels, 100 examples | 86 / 100 = 86% | 48% (+38 points) |
| Banking77, 72-label shortlist, 100 examples | 64 / 100 = 64% | 87% |
The typed suite is 400 cases with five decisions each. On that H200 run the split was Choice 428/600, Noul 484/600, and Score 551/800. The +0.45 point gap is the number to remember, because it is the broadest suite in the table and the margin is smaller than a rounding argument. A 100-example pilot can move several points from a handful of items. The emotion gap of 38 points is that kind of pilot: 86 versus 48 on 100 rows, six labels. It is a large difference on a small sample. It is also Supersonic Labs' measurement of their checkpoint against a stored reference, not a result explainx.ai reproduced.
Banking77 is the result that should decide whether you try Julia 1 on a long menu of similar classes. Seventy-two banking intents do not fit in a native call. The page says a Router narrows the list first, and that narrowing can remove the correct label before the model ever scores the finalists. The model card is more specific: the 72-label pilot went through a ranking and a top-16 shortlist, and a grouped result is not a global probability over all 72 labels. Julia 1 got 64 of 100. The Jev reference is 87 of 100. Twenty-three points, on the task where the option list is the hard part, is a larger fact than half a point on the mixed suite.
A CPU run dated September 25, 2026, on the same tasks, used PyTorch 2.14.0+cpu, strict encoding, a 1,024-token limit, and weights whose SHA-256 prefix is df853bf7fe42. Typed decisions moved to 1,451 / 2,000 = 72.55%. AG News stayed 94/100. DAIR Emotion stayed 86/100. Banking77 fell to 60/100 correct, with 3 abstentions, 97% coverage, and 60% accuracy, because abstentions count as errors. The model card's September 26 CPU reproduction (FP32, torch 2.14.0) reports the typed split as Choice 426/600, Score 542/800, and Noul 483/600, which adds up to the same 1,451 / 2,000. That split is the CPU package. The 428 / 484 / 551 split above is the H200 comparison. They are close, and they are not the same run.
MASSIVE, in this evaluation, is scenario classification: one of 18 scenarios. Julia 1 scored 110,573 / 154,648 = 71.50% across 52 locales, 2,974 examples each. The model card calls that all-locale figure macro accuracy, which matches an unweighted mean when every locale has the same count. Portuguese (pt-PT) was 2,565 / 2,974 = 86.25%. English (en-US) was 2,580 / 2,974 = 86.75%. Brazilian Portuguese, intent classification, and slot filling are outside this test. A strong European Portuguese scenario score is not a claim about Brazilian Portuguese intents.
How fast is a CPU decision?
Supersonic Labs published three machines. The workloads differ, so the table is a set of conditions, not a ranking of chips.
| Machine and test | What they ran | Median | p95 | Rate |
|---|---|---|---|---|
| Apple M4, one call at a time | 128 decisions, 100 context words, 4 options, 4 CPU threads, public mmBERT-small tokenizer, local Python | 33.15 ms | 44.23 ms | 28.01 decisions/s |
| Apple M4, batches of 16 | Same 128 decisions, time is per batch | 312.11 ms per batch | — | 51.20 decisions/s. 370.6 MiB RAM at the end of the batch run |
| Samsung SM-X510 | 40 decisions, ONNX Runtime 1.27.0, Android 16, Python 3.13.13, native Rust tokenization | 203 ms | 205 ms | 5.0 decisions/s, about 92 tokens per decision |
| Intel Core i5-1235U, typed | September 25 package, 2,000 decisions | 294.81 ms | 428.49 ms | — |
| Intel Core i5-1235U, AG News | 100 examples | 107.83 ms | — | — |
| Intel Core i5-1235U, DAIR Emotion | 100 examples | 89.83 ms | — | — |
| Intel Core i5-1235U, Banking77 | 100 examples, 72-way narrowing | 3,713.54 ms | 5,125.10 ms | — |
On the M4, a single short decision is a few tens of milliseconds, and batching 16 at a time raises throughput from 28.01 to 51.20 decisions per second. That batch figure is the time for the batch, 312.11 ms, divided into the 16 rows. It is the same laptop, a different call shape. The tablet is a different stack. ONNX Runtime listed NNAPI and then fell back to CPU per operator. XNNPACK was left out because it miscompiles a Reshape on this graph. Peak RSS was 393.1 MB, with the 550.1 MB weight file memory-mapped. Five decisions a second on that fallback path is a real measurement of a tablet running the graph on CPU. It is not a WebGPU number, and it is not comparable to the M4's four-option Python loop.
The i5 row is dominated by Banking77. Typed decisions sit at a 294.81 ms median. AG News and emotion, four and six labels, sit near 108 ms and 90 ms. The 72-label path, with narrowing, sits at 3.7 seconds median and 5.1 seconds at p95. If your product is a four-way router, the i5 typed and AG News rows are the closer cousins. If your product is "which of 70 similar intents," budget for the Router, and budget for it to be wrong in the way the accuracy table already shows.
What context length was actually scored?
The evaluated configuration on the lab's page accepts up to 1,024 tokens for context, question, and options together. Python 3.11 or newer. That is the limit on the September 24 and September 25 accuracy runs.
explainx.ai opened the Hugging Face model card before writing this. The card does state a longer runtime limit, so it belongs here with the card's own caveat attached. The card says the runtime supports up to 8,192 combined tokens, that this is the default, and that historical benchmarks used 1,024. It also says an 8,192-token CPU smoke test passed, and that task accuracy at 8,192 tokens is not established. A secondary blog that quotes 8,192 as if the 73.15% figure were measured there is ahead of the published scores. Use 1,024 if you are trying to match the tables. Treat 8,192 as a runtime ceiling the lab has smoke-tested, not as a scored context length.
Strict encoding rejects overflow. The card's example also sets a 512-token budget for the question and options together, and a 48-token limit on each option. Those knobs are in the load call below. They are the card's example, not a second benchmark.
What a Python call looks like
The model card documents the install and the named-question API. This sketch follows that card. It is the interface, not a result explainx.ai executed.
python -m pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('SupersonicLabs/Julia-1', local_dir='Julia-1')"
python -m pip install -e ./Julia-1
from julia import load_model
engine = load_model(
"Julia-1",
device="cpu",
strict_encoding=True,
max_length=8192,
head_length=512,
)
result = engine.predict(
state="I was charged twice for the same order.",
questions={
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Billing and payment disputes",
"shipping": "Shipping and delivery",
"access": "Account access and login",
},
},
},
)
print(result["answers"]["team"]["choice"])
print(result["answers"]["team"]["probabilities"])
max_length=8192 is the card's example and the runtime default the card describes. The accuracy tables used 1,024. Set max_length=1024 if the point of the run is to match those tables. Caller ids such as billing come back unchanged. Probabilities on this named interface are a full softmax, keyed by those ids for a choice, by zero-based index strings for a score, and by false/true for a yes/no. The card still supports an older list form, predict on a list of state, question, options, and type, which returns an index into the list you passed and display-formatted probabilities. engine.logits returns raw scores. Keep the engine loaded between requests. JULIA_CPU_THREADS defaults to 4 if you set it before Python starts. The package is not a Transformers text-classification pipeline, and AutoModelForMaskedLM will not turn these weights into Julia 1.
Can it run in the browser?
The launch page and the ONNX repo both say yes, with WebGPU. The ONNX repository exports the same weights and the same decision graph. You call load() once against a base URL, then predict(rows) for indexes and probabilities, or logits(rows) for raw scores. Strict encoding is the default. A bundled Rust WebAssembly tokenizer encodes the request. ONNX Runtime compiles the graph to WGSL. model.onnx and model.onnx.data have to stay side by side. If you serve those files yourself, the lab says the computation stays in the browser. A browser without WebGPU cannot run this package.
The ONNX card also reports a parity check, separate from the accuracy suite: 100 validation requests, batches of four, in Brave on Linux, with 100 of 100 predictions matching the original runtime and a maximum absolute logit gap of 0.00225. Median time for those 100 decisions was 7.55 seconds, about 75.47 ms each, after warmup. The original accuracy suite was not rerun on WebGPU. A near-tie can still flip when the logits are close. That parity file is evidence about numeric agreement on one request set. It is not a new Banking77 score, and it is not a license to line the browser up against the M4 table above.
What people are asking
Is there a multimodal Julia 1? Someone asked for a multimodal version on the September 27 launch thread. The lab page, the Python card, and the ONNX card do not answer it. As of this post, that question is unanswered. The published model is a text encoder over context, a question, and written options. Image input shows up in the GLM one-token setup, because that path is still a chat model. It does not show up in Julia 1's evals.
Is the Jev comparison fair? Partly, and the lab labels the part that is not. The Jev column is a protocol reference, not a stopwatch on the same H200 with the same batch. The three classification numbers are 100-example pilots. Emotion's 38-point gap and AG News' three-point gap can both move when the sample grows. Banking77 is the cleanest warning in the set: Julia 1 never scored all 72 labels in one native distribution. A Router, described on the card as a top-16 shortlist after ranking, threw labels away first. If the correct intent was in the discarded group, the final softmax cannot recover it. A reference that handles the long list differently will look better for reasons that include the shortlist, not only the encoder. Read 73.15% versus 72.70% as "close on this suite, against a stored reference." Read 64 versus 87 as "long, similar label sets are where this checkpoint lost."
Does a typed answer mean a correct answer? The output is constrained to the options you passed, so a malformed JSON object is not the failure mode. A wrong option with a confident score still is. That is the same class of miss explainx.ai already tracked for hosted decision models in where Jev actually fails: the schema is valid, and the label is wrong. Abstentions in the CPU Banking77 run are scored as errors, which is the honest accounting. A production router still has to decide what to do when the top score is 0.4 on a lookalike pair.
Can I match the CPU package? The September 25 record names PyTorch 2.14.0+cpu, strict encoding, a 1,024-token limit, and the weight hash prefix df853bf7fe42. The card's reproduce script checks the published weight hash and a pinned typed-decisions Parquet, then writes a CPU FP32 run. That harness does not regenerate the classification pilots or MASSIVE. If your torch build or thread count differs, expect the latency table to move even when the argmax stays put.
Honest limitations
- The +0.45 point typed-decision edge is Supersonic Labs' H200 number against a protocol reference. explainx.ai did not re-run Jev or Julia 1.
- Banking77 at 64% (CPU rerun 60%, with 3 abstentions) is the load-bearing failure. Similar labels plus a narrowing Router can delete the right class. A grouped score is not a probability over the original list.
- AG News and DAIR Emotion are 100-example pilots. The emotion gap is large and the sample is small.
- MASSIVE here is 18-way scenario classification across 52 locales. It does not include Brazilian Portuguese, intent classification, or slot filling.
- External knowledge and multi-step math are out of scope. The model compares answers you provide. The card says it should not be counted on to supply a missing fact, solve an algebraic equation, or carry a chain of calculations because the right option happens to be in the list.
- Consequential decisions need a person. Evaluate the questions and options from your own domain before you route refunds, access, or anything else a wrong label would harm.
- The hosted API at $0.025 per million input tokens was a plan on September 27, 2026, not an endpoint you could call.
- Accuracy at 8,192 tokens is not established. The scored runs used 1,024.
- WebGPU accuracy was not rerun. The ONNX repo's 100-request parity check is a numeric comparison, and floating point can still flip a near tie.
- The tablet path fell back from NNAPI to CPU, and XNNPACK was excluded because of a bad Reshape compile. "Runs on almost anything" still depends on the execution provider you actually get.
What to run this week
Download the Apache-2.0 checkpoint if you want a 144.3M encoder on a CPU, with a browser build beside it, and you can keep each native call inside 2 to 20 distinct options. Use the September 24 table as Supersonic Labs' claim: a 0.45 point edge on 2,000 mixed typed decisions, solid four-label and six-label pilots on 100 rows each, and a real loss on a 72-way banking shortlist. Time your own hardware. The M4, the i5, and the tablet were measured on different inputs, and only your prompts will tell you which row you resemble.
If you need images, or you already pay for a GLM-5.3-Flash server, stay with the one-token method. If you want a local catalogue of several decision models behind one HTTP shape, Ollaya is that runtime, and Laya-MLX is the Apple Silicon port of Laya. If you want the hosted original of this product category, that is still Jev. Julia 1 is the new file in the set: small enough to map into a few hundred megabytes of RAM on the machines the lab timed, open enough to inspect, and explicitly weak when the label list gets long and the labels look alike.
Related reading
- How GLM-5.3-Flash scores a closed choice in one token
- How Jev works, without re-deriving it here
- TypeSafe's Jev launch
- What a System One Model is
- Ollaya, a local runtime for decision models
- Laya-MLX on Apple Silicon
- Where a typed decision is still the wrong label
Primary sources: Supersonic Labs, Julia 1, SupersonicLabs/Julia-1, SupersonicLabs/Julia-1-ONNX.
Accuracy, latency, pricing, and context-length figures are Supersonic Labs' measurements as of September 27, 2026. explainx.ai did not re-run them. The hosted API was not open at publication, and task accuracy at 8,192 tokens was not established on the model card.
