Model News

Cloudflare releases Clef-omni with audio and video support

Cloudflare launches Clef-omni, a decision model processing audio, video, images, and text in one API call, with price cuts and speed gains for existing models.

Headline card: Cloudflare releases Clef-omni with audio and video support
On this page
  1. What changed
  2. Why it matters
  3. What to test
  4. The conclusion

What changed

Cloudflare announced the release of Clef-omni, a new decision model that handles audio, video, images, and text input natively in a single pipeline. The company says this expands decision model capabilities beyond the text-only approach that has dominated the category since TypeSafe’s Jev debut.

According to Cloudflare’s blog, the company built Clef-omni on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation and removed the text-to-speech output components to optimize for decision making rather than generation. The model discards transcription and captioning overhead, mapping media elements directly into a unified sequence where video and audio sync with visual frames for joint processing.

Alongside the Clef-omni launch, Cloudflare cut pricing for Clef-flash to $0.038 per million input tokens, down from the original $0.09. The company says this pricing now undercuts Jev. Clef pricing remains at $0.24 per million input tokens. Clef-omni launched at $0.15 per million input tokens.

Cloudflare also made the Clef model faster through serving infrastructure optimizations, not model weight changes. The company moved to SGLang for serving and reports median speedups of 1.7x to 2.0x across different input sizes, with text-only decisions returning in about 130 milliseconds at the median and full 21-second video clips with sound scoring in about 1.5 seconds.

One trade-off: Clef-flash’s hosted context window dropped from 64k to 24k tokens. Cloudflare says usage data shows only 0.24 percent of requests exceed 24k input tokens, so the company made this change to enable lower pricing. The open-weight model on HuggingFace retains the 256k context window for self-hosted deployments.

Why it matters

The release addresses a practical gap in the decision model space. Until now, decision models focused on text-only classification or required users to build cascading pipelines to handle multiple input types separately, transcribing audio and extracting video frames before feeding data to models.

Clef-omni eliminates this workflow complexity by accepting audio (WAV or MP3), video (MP4 or WebM), images, and text in a single API call. For teams building AI agents that need to classify or make decisions based on multimodal inputs, this reduces infrastructure friction and removes the need to orchestrate separate models for each modality.

The pricing changes make decision models more accessible to cost-conscious teams. Clef-flash now undercuts Jev on price, which matters for high-volume workflows where model cost scales with request volume. Teams that previously chose Jev on price grounds now have a Cloudflare alternative to consider.

The speed gains for Clef, driven by moving to SGLang for serving, reduce latency for existing users. A 1.7x to 2.0x median speedup matters for latency-sensitive applications, particularly those handling large context windows. The move from 262ms to 152ms median latency at 800 tokens represents a meaningful improvement for real-time decision workflows.

What to test

Before adopting Clef-omni or the updated Clef and Clef-flash models, teams should verify several claims and trade-offs:

Performance verification

Cloudflare published benchmark results comparing Clef-omni against Clef, Clef-flash, and Jev across multiple datasets. These benchmarks show mixed results. Clef-omni underperforms the earlier Clef model on several benchmarks: BFCL drops from 98.47 to 98.2, ToolRet drops from 69.19 to 66.6, and on Home appliances the gap is more significant at 69.3 versus 82.95. On TypeSafe workflow evals, Clef-omni trails Clef on most metrics. Teams should test whether these performance differences matter for their specific use case rather than assuming multimodal capabilities come without cost.

Context window trade-offs

Test whether the 24k context window reduction for Clef-flash affects your workflows. Cloudflare’s data suggests 99.76 percent of requests fit within 24k tokens, but your traffic pattern may differ. If you exceed this limit, you will need to either switch to Clef (64k) or handle batching yourself. The hosted version enforces the limit while self-hosted deployments on HuggingFace support the full 256k window.

Latency in production

Cloudflare reports median latencies of about 130ms for text, 150ms for images, and a few hundred milliseconds for audio. A 21-second video takes about 1.5 seconds. Run these through your actual latency requirements. What counts as acceptable depends on your application’s tolerance for decision-making delay. The speed improvements to Clef may matter more than Clef-omni’s absolute latency for some use cases.

Multimodal input handling

Test Clef-omni’s actual behavior with your media types. The model accepts audio (WAV or MP3) and video (MP4 or WebM) natively, but verify that compression formats, audio codecs, and video frame rates your system produces work as expected. The API documentation provides examples, but real-world media often includes edge cases.

Open-weight deployment

Cloudflare released model weights on HuggingFace. If self-hosting matters to your team, test the deployed model with SGLang 0.5.22 or later. Verify that the self-hosted version actually supports the larger context windows Cloudflare advertises and that the new SGLang integration works as documented.

The conclusion

Cloudflare has moved quickly to expand its decision model family from text-only to multimodal. Clef-omni represents a convenience gain for teams that need to classify decisions across audio, video, and images without building orchestration layers around separate models.

However, the performance trade-offs evident in the benchmarks suggest that combining all four modalities may come at a cost. Clef-omni underperforms Clef on multiple benchmarks, and on TypeSafe workflow evals it trails Clef across most metrics. Teams considering Clef-omni should benchmark against Clef for their specific tasks rather than assuming the new model is strictly better.

The pricing moves and speed gains feel more incremental but meaningful. Undercutting Jev on Clef-flash price and delivering 1.7x to 2.0x speedups on Clef make both models more compelling for cost and latency sensitive workflows. The context window reduction for Clef-flash is a real constraint for some workloads, but Cloudflare’s data suggests it affects fewer than one percent of requests.

Early feedback on multimodal decision accuracy in production will reveal whether Clef-omni actually solves the workflow problem Cloudflare claims, or whether users prefer specialized models for each modality.

AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure