Cloudflare introduces Clef-omni, an open-weight multimodal decision model capable of processing audio, video, image, and text inputs in a single API call. This architecture streamlines workflows by eliminating the need for cascading pipelines of single-modality models and leverages a mixture-of-experts (MoE) foundation for efficient, schema-constrained scoring. The article also highlights performance optimizations for existing Clef models, including faster inference and reduced pricing.
Read original on Cloudflare BlogTraditional AI pipelines often involve multiple specialized models for different modalities (e.g., speech-to-text, object detection, text processing), requiring orchestration and data transformation between stages. Clef-omni presents an architectural shift by consolidating these capabilities into a single model. This simplifies the client-side integration and reduces system complexity, as developers no longer need to manage cascading calls or build custom pipelines to integrate diverse data types. The goal is to interact with data more naturally, akin to human perception, by processing all modalities concurrently.
Key Architectural Benefit
Consolidating multimodal processing into a single model significantly reduces system complexity, latency, and operational overhead compared to chained single-modality models. It enables more holistic context comprehension for decision-making.
Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts (MoE) foundation. This architecture allows the model to natively handle various input types (text, imagery, audio, video) within a single pipeline. A crucial design decision for Clef models is to skip output token generation (unlike Large Language Models), focusing solely on decision-making by scoring valid parameter options. This eliminates the overhead of transcription or captioning, leading to faster inference.
Cloudflare continuously optimizes its Clef family of models for speed and cost-efficiency. For instance, Clef-flash received a significant price reduction by optimizing the model and making a trade-off on the context window for the hosted version (24k instead of 64k), based on usage data indicating that most requests do not exceed 24k tokens. The larger context window (256k) is still available for self-hosted versions. The Clef model also saw speed improvements, primarily due to optimizations at the serving infrastructure layer, including a transition to SGLang for model serving on Workers AI. These infrastructure-level changes are critical for delivering low-latency AI services at scale.