explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • Quick answers
  • Where Clef-Omni fits: the Clef family
  • How Clef-Omni works
  • Try it: the card's example
  • Benchmarks: where it wins and where it does not
  • Clef-Omni versus the rest: which should you pick?
  • What people will ask next
  • A practical evaluation plan
  • Limits and open questions
  • What this means for builders
  • Related reading
← Back to blog

explainx / blog

Cloudflare Clef-Omni: An Open Multimodal Decision Model for Audio, Video and Images

Decision Models, Cloudflare, Open Source, Multimodal, Models

Part of Decision Models

Cloudflare Clef-Omni is an Apache-2.0, 30B-A3B decision model that scores typed questions over text, images, audio and video. Benchmarks, VRAM, how to run it.

Oct 9, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Cloudflare Clef-Omni: An Open Multimodal Decision Model for Audio, Video and Images

Cloudflare has put a new open model on Hugging Face: Clef-Omni, a 30B-A3B mixture-of-experts that reads text, images, audio and video and answers structured questions with probabilities instead of prose. It is the third member of Cloudflare's Clef family of decision models, and it is the one built on an omni backbone with an audio encoder, so it can listen to an audio clip as well as read images and video before it decides. The model card lists an Apache-2.0 license and a Jev-compatible API, which means it can drop into existing decision-model pipelines.

If you are new to the category, start with our guide to what decision models are. Short version: instead of generating an answer and parsing it, you ask typed questions and read back a probability for each option.

Quick answers

table · 2 cols
QuestionAnswer
What is it?A multimodal decision model: state plus typed questions in, per-option probabilities out.
Size and architecture30B-A3B MoE, listed as 35B parameters in bfloat16.
Base modelQwen3-Omni-30B-A3B-Instruct, post-trained by Cloudflare.
LicenseApache-2.0, following the base model.
InputsText, JSON, images, audio, video (up to 64,000 tokens by default).
GPU needsAbout 64 GB in bfloat16; tested on one H200.
APIJev and SystemOne compatible.
ServingSGLang support is listed as coming soon.

Where Clef-Omni fits: the Clef family

Cloudflare announced the first Clef models on October 1, 2026, about two weeks after TypeSafe launched Jev, according to The Register. That coverage describes Clef as post-trained on Qwen3.8-27B and the smaller Clef-flash on Qwen3.5-9B, both able to answer yes/no, multiple-choice and ranking questions. The Register reported pricing of $0.24 per million tokens for Clef on Workers AI, against $0.042 for Jev, and noted that Cloudflare's scores were self-reported. Clef-Omni follows as the audio-and-video variant.

If Jev is new to you, our write-ups of the Jev and System One launch and the Jev speed and cost fact-check give the baseline that Clef is measured against. Cloudflare is also not alone: OpenAI has its Decisions API, Perplexity shipped pplx-decider, and Strands released an open 2B decider. Our ranked comparison lives in the top 10 decision models.

A signpost with three arrows, one highlighted, picturing how Cloudflare Clef-Omni picks one option per typed questionA signpost with three arrows, one highlighted, picturing how Cloudflare Clef-Omni picks one option per typed question

How Clef-Omni works

The model card describes four moving parts.

  1. Record. You pass a JSON record with a state (any JSON value or string), optional images, audio and videos, and a questions map.
  2. Backbone. The Qwen3-Omni thinker reads the state, including media, and produces final hidden states. Video files are sampled at 2 frames per second and, when every video has a soundtrack, the audio is heard alongside the frames.
  3. Joint schema head. A small transformer head routes evidence from the state to each question and scores all options of all questions together.
  4. Output. One logit per allowed option per question. A softmax per question yields probabilities.

Each question is one of three types: noul (true or false), choice (named options with descriptions) and score (ordered options, returning an expected score). Because there is no generation step, there is nothing to parse and the output shape is guaranteed. That is the whole pitch of this category, and the reason latency stays predictable: one forward pass answers every question. Our explainer on System One models walks through why this is different from a chat model.

Try it: the card's example

This is the usage snippet from the model card, shortened. It checks an invoice and asks two questions.

python
import sys, torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef-omni")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")
record = {
  "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "status": "overdue"}},
  "questions": {
    "status": {"type": "choice", "instructions": "What is the invoice status?",
               "criteria": {"paid": "Paid.", "overdue": "Past due.", "draft": "Not sent."}},
    "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
  },
}

Then call encode_record, collate_records and run the model under torch.inference_mode(). The card tested torch 2.11 and transformers 5.10.2. For audio or video, add audio and videos lists to the record. The card's example asks whether glass is breaking in a call recording and whether a dashcam clip shows a collision.

A Jev-style call is also supported through the systemone helper, which takes a /v1/systemone request body and returns model, answers and usage. If you already run Jev in production, the card's claim of full API compatibility suggests swapping the model name is the migration path; test it first. For local runtimes, see our note on llama.cpp support for decision models.

Benchmarks: where it wins and where it does not

The numbers below come from Cloudflare's internal run of Decision Index 0.2.1 as published on the model card. Treat them as vendor-reported.

table · 5 cols
BenchmarkClef-OmniClefClef-flashJev
MMLU (accuracy)92.790.391.891.7
CRUXEval (accuracy)88.886.786.173.0
CLadder (accuracy)97.294.097.772.6
BANKING77 (macro-F1)94.894.290.979.7
GPQA Diamond (accuracy)47.448.051.078.3
RAGTruth (hallucination F1)42.079.435.676.5
ForecastBench (Brier, lower is better)11.713.910.617.4

Some takeaways:

  • Strong on classification-style tasks. Intent and domain classification (BANKING77, CLINC150+OOS at 97.7) and logic tasks like CLadder are strong.
  • Weak on hard reasoning and hallucination detection. GPQA Diamond and RAGTruth are well behind Jev and the dense Clef. If you are building a hallucination judge, the dense Clef is the better pick from this family.
  • Mixed on workflows. On Typesafe Evals, Clef-Omni scores 60.2 exact actions on invoice processing versus 61.8 for Jev, and 71.6 on customer service versus 76.0 for Jev. It is not a clear upgrade over what you may already run.
  • Audio and video are the point. Most Decision Index rows are text. The card gives no audio or video benchmark, so the multimodal strength is unquantified in the published results.

Two shapes on matching plinths joined by a dotted line, a visual for comparing Cloudflare Clef-Omni with Clef, Clef-flash and JevTwo shapes on matching plinths joined by a dotted line, a visual for comparing Cloudflare Clef-Omni with Clef, Clef-flash and Jev

Clef-Omni versus the rest: which should you pick?

table · 3 cols
If you needPickWhy
Decisions over audioClef-OmniThe only Clef variant with an audio encoder (The Register reports Clef also handles images and video)
Fast, cheap text classificationClef-flashSmaller dense model
Best accuracy on text and imageClefLeads the family on several rows
Hosted with no GPUCheck Workers AI listingsThe card documents self-hosting; hosted options are on Cloudflare's platform
Hallucination or faithfulness checksClef or JevClef-Omni scores 42.0 on RAGTruth

What people will ask next

Is it really "open source"? The weights and loading code are Apache-2.0 on Hugging Face, which is what most developers need to self-host, fine-tune and ship commercially. The training datasets are not published, so it is open weights rather than fully open data. That mirrors its base model, Qwen3-Omni, whose license the card says it follows.

Why a decision model instead of a chat model with JSON mode? A chat model generates tokens one at a time, can drift from the schema, and gives you a confidence you have to guess at. A decision model scores the allowed options directly, so the answer is always one of the options you defined and comes with a probability you can threshold. Our structured output and JSON mode guide covers the generation-side alternative and where it still wins, such as free-text explanations.

What does the joint head buy you? Scoring all questions together lets one pass answer many questions about the same state. If you ask five questions about a support call, the audio is processed once. That is a large saving over five separate calls, and it keeps the answers consistent because they share the same evidence.

How should I set thresholds? Use the probabilities as routing signals. For example, auto-approve above 0.95, send 0.6 to 0.95 to a cheaper review step, and escalate below 0.6 to a human. Calibrate these numbers on a labeled sample of your own traffic, because published benchmark accuracy does not tell you how well-calibrated the probabilities are on your data.

What about latency? Cloudflare has not published latency for Clef-Omni on the card. Expect it to be slower than the 9B-class Clef-flash because a 30B-A3B model activates about 3B parameters per token but still has to load all 35B parameters into memory, and audio and video inputs add many tokens. Measure on your hardware before you commit.

A practical evaluation plan

If you want to decide whether Clef-Omni deserves a slot in your stack, a one-day test is enough.

  1. Collect 200 labeled examples from real traffic, including a spread of easy and borderline cases and at least 40 with audio or video.
  2. Define your questions as typed schemas. Start with noul questions because they are the easiest to score and calibrate.
  3. Run Clef-Omni, Clef-flash and your current model on the same set and compare exact-match accuracy, not just average confidence.
  4. Check calibration. Bucket answers by predicted probability and compare to the actual hit rate in each bucket.
  5. Measure cost per decision. Include GPU hours for self-hosting against per-token pricing for hosted models.
  6. Review the misses by hand. Look for patterns such as failures on noisy audio or long videos, and decide whether a fallback is needed.

This mirrors the approach in our guide to classifier feature engineering and calibration, which explains why calibration matters more than headline accuracy when you act on probabilities.

Limits and open questions

  • Self-reported scores. The Register noted Cloudflare's earlier numbers had not been reproduced on the official Decision Index. The same caution applies here.
  • Hardware cost. 64 GB of GPU memory means an H200 or a pair of smaller cards, far above the 9B-class decision models.
  • No generation. You cannot ask it to explain a decision. It returns probabilities only.
  • Weights note. The card says the base model's speech-output weights are included unchanged but unused, so the download is larger than the decision path needs.
  • Serving. SGLang support is "coming soon" per the card; for now you run the provided Python code.
  • Training data. The card does not describe the post-training datasets. The Register reported a Cloudflare product manager saying the datasets for the Clef release are not public, so "open" here means weights, not data.

What this means for builders

Audio and video decisions have until now meant a pipeline: transcribe or caption, then classify. A single model that hears the call and watches the clip removes a stage and a failure mode. Good early uses are call-center triage, claims review with photos and recordings, content moderation of short clips, and safety checks on agent screen recordings. As always with this category, calibrate the confidence thresholds on your own data before you trust a probability, and keep a human in the loop for high-stakes calls. Our post on cheap verification checkpoints in agent pipelines shows where a decision model slots in. If your agents take real-world actions, AgentBeam, the agent security platform from the explainx.ai team, is built to stop agents before they take dangerous actions.

Related reading

  • What are decision models?
  • What is a System One model?
  • Top 10 decision models ranked
  • OpenAI Decisions API vs Jev
  • Strands 2B open decision model
  • llama.cpp decision model support

Specs and scores are accurate as of October 9, 2026 and come from the Hugging Face model card; check it for changes.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 10, 2026

Talorys: A Personal AI Agent You Deploy to Your Own Cloudflare Account in One Command

Talorys, a Show HN project, packages a personal AI assistant (chat, memory, tasks, notes, reminders) that runs entirely in your own Cloudflare account on the free plan. We cover setup, the architecture, the real limits and the self-hosted debate.

Oct 8, 2026

Top 10 Decision Models in October 2026: Jev, Perplexity, OpenAI and Open Source

In three weeks the decision-model category went from one hosted product to a crowded field of hosted APIs and open weights. This is a ranked list of ten you can use today, starting with Jev, with prices, latency, licenses and the caveat that matters most for each.

Oct 7, 2026

Cloudflare Open-Sources Its Security Audit Skill: How the Six-Phase Agent Workflow Works

Cloudflare published the single-repo skill that seeded its fleet-wide vulnerability discovery harness. It turns a coding agent into a security auditor across six phases with independent verification. Here is how it works, how to install it, and what developers say about cost and noise.