jsonscraper

Cloudflare releases Clef-omni for structured audio and video analysis

The model processes text, images, audio, and video, then scores responses against a predefined schema. In the same update, Cloudflare cut the price of Clef-flash and reduced the context window of its hosted version.

Editorial Policy Report an error

Checking in a single request whether a label is legible in a photo, a fan is running in a video, and an unusual noise can be heard in an audio recording is one example Cloudflare gives for its new Clef-omni model. On October 9, 2026, the company announced its release: the model accepts text, images, audio, and video, and scores response options against specified questions.

Article: Cloudflare releases Clef-omni for structured audio and video analysis
Summary: On October 9, Cloudflare announced Clef-omni, an open-weight model for structured decisions based on text, images, audio, and video. The company also cut the hosted Clef-flash price from $0.09 to $0.038 per million input tokens and reduced its context window from 64K to 24K tokens.
Image role: cover — the central idea
Section: 
Visual subject: Industrial equipment inspection with
© jsonscraper · AI-generated illustration

Clef-omni continues the Clef family of decision-making models, which Cloudflare introduced a week earlier. The developer defines a schema of questions and acceptable answers; the model returns scores for those options, ready for structured processing. This is suitable for classifying requests and multimedia content when the result needs to be passed to an application or workflow.

One schema for different input types

In Cloudflare’s Clef-omni announcement, the model is described as built on Qwen3-Omni-30B-A3B-Instruct with a mixture-of-experts architecture. It supports text and JSON, images, WAV and MP3 audio, and MP4 and WebM video. For each multiple-choice question, the model scores the available options, allowing an application to process the result in a predefined format.

For example, when triaging fault reports, a team could ask whether a label is legible in a photo, whether a recording contains noise, and whether an indicator is flashing in a video. A single request combines multimodal evidence into a structured response for subsequent routing.

Cloudflare reports median latency of about 130 ms for text decisions and about 150 ms for images. For a 21-second video with audio, the company reports about 1.5 seconds. These are results from tests reported by the company; evaluating your own system also depends on its files, question schema, and deployment method.

Open weights and cloud deployment

Cloudflare published Clef-omni weights on Hugging Face under the Apache-2.0 license. The model card describes local deployment and testing on a single H200 accelerator; the bfloat16 configuration is listed as requiring about 64 GB of GPU memory. This gives teams the option to evaluate self-hosting in light of their hardware and configuration requirements.

Inside the Nebra Smart LoRa Gateway
Pi Supply · Unsplash License

For cloud-hosted Clef-omni, Cloudflare announced a price of $0.15 per million input tokens. The cost depends on the input mix: Cloudflare describes separately how images and audio are converted into billable tokens, while video also affects the volume of input data. Check the current terms in the Workers AI documentation before estimating a budget.

Clef-flash gets cheaper, while the hosted version’s context window shrinks

In the same update, Cloudflare cut the price of hosted Clef-flash from $0.09 to $0.038 per million input tokens. The hosted version’s context window was reduced at the same time, from 64K to 24K tokens. According to the company, only 0.24% of requests exceeded the new limit.

Before switching to the lower price, teams should check the length of their actual requests and documents. The published Clef-flash weights have not changed: Cloudflare says the self-hosted model supports a context of up to 256K tokens. That limit applies to the open-weight version, not the hosted version with its 24K window.

Clef-omni’s practical idea is to combine multimodal input with choosing from predefined options in a single request. Teams building products that classify content, compare audio with images, or fill fixed fields can evaluate it using their own examples. When choosing between cloud and self-hosted deployment, they should separately consider pricing, hardware requirements, and the context length they need.

People

No people listed for this article yet.

Keep readingGoogle Opens SynthID Detector to Check Images, Video, and Audio
Read the next article

Turn what you read into a working integration

Explore jsonscraper's social-data APIs, test requests and build your next workflow.

Explore APIs