Cloning Jev in Two Weeks: Cloudflare Open-Sources 27B Decision Model Clef

Cloning Jev in Two Weeks: Cloudflare Open-Sources 27B Decision Model Clef

CloudflareDecision ModelsClefJevAI AgentsOpen Source Models

Sources:Cloudflare Blog + The Register + HN · HN

How short can the lifecycle of a new AI model category be? In mid-September, Typesafe AI launched Jev, demonstrating how a compact model can output probability scores for an agent’s branching actions. Just two weeks later, Cloudflare dropped two structurally identical models onto Hugging Face under the Apache 2.0 license, dubbed Clef.

This is not merely a follow-up demo—it upends the table. Jev is closed-source, proprietary in architecture, and API-only. Clef releases full open weights and claims the top spot on the very Jev Decision Index its rival championed. Yet amid this “two-week reproduction” race, the most revealing question surfaced in the Hacker News comments: What is actually new here?

What Are Decision Models: From Generating Answers to Scoring Options

To understand this category popularized by Jev, consider how conventional LLMs handle tasks like “should this ticket be escalated?” They rely on autoregressive generation, outputting tokens one by one to “write” an answer—a process that is sluggish and non-deterministic. Decision models eliminate token generation entirely: you pass in an input state alongside a typed question schema, and the model directly emits probability scores for each predefined option without generating a single word of intermediate text.

Cloudflare’s blog illustrates this with ticket triage: submit an incoming customer message, and the model returns parallel, typed scores such as “Urgent: Yes 87%” and “Assigned Team: Technical 91%”. Downstream application logic routes, escalates, or flags for human review based directly on these calibrated probabilities. With zero text generation, the design fits naturally into the critical hot path of agent decision loops.

Decision model input-output flow: ticket body plus schema for parallel scoring Figure: Decision model workflow—ticket text and schema inputs yield parallel probability distributions across all queries. Source: Cloudflare Blog

The real breakthrough lies at the product tier. Historically, this capability was simply called a “classifier.” Back in the BERT era, training a narrow domain classifier took an hour on a laptop and consumed under 1GB of VRAM during inference, outpacing API roundtrips. Jev’s contribution was packaging this capability into a generalized product: swap out the JSON schema to repurpose the model without retraining, all wrapped in an ergonomic developer API. Jev validated product-market fit; now the floodgates are open.

Clef’s Technical Blueprint: Qwen Backbone with Prefill-Only Scoring

The Clef family arrives in two tiers: Clef (27B) and Clef-flash (9B), built on top of Qwen3.8-27B and Qwen3.5-9B respectively. Training keeps the backbone weights frozen, tuning only a rank-256 low-rank adapter alongside a custom routing head.

At inference time, the Qwen backbone executes a single prefill-only pass, followed by parallel scoring across all valid schema candidates. Because the decision stage is non-autoregressive and bypasses iterative token decoding, it offers a structural latency advantage over general-purpose LLMs. Cloudflare describes this architecture as “two-phase attention routing”: candidate options pull prompt-relevant context, fields cross-attend before referencing the raw input, and schema-constrained scoring completes the pass. The training recipe combines label-smoothed cross-entropy with Brier loss for probability calibration, supplemented by RLCD (Reinforcement Learning from Categorical Distributions) to assign partial credit across adjacent ordinal outcomes.

Cloudflare reported the following comparative benchmarks:

MetricClefClef-flashJev
Decision Index61.257.157.9
Median Latency209ms38.8ms524ms
Context Window64k64k32k (state + single question)
Vision InputSupported (images + video)SupportedNot supported
WeightsApache 2.0 Open SourceApache 2.0 Open SourceClosed source
Workers AI Pricing$0.24 / M tokens$0.09 / M$0.042 / M

Two figures stand out. Clef-flash achieves a Decision Index of 57.1 in 38.8ms—nearly matching closed-source Jev (57.9 score, 524ms) at roughly one-thirteenth the latency. Meanwhile, the flagship Clef takes the crown on the official Decision Index benchmark, outscoring Jev by over three points while cutting latency in half. In internal testing paired with Browser Run for website categorization, Cloudflare reported crawling, rendering, and classifying a domain in 2.2 seconds; their general-purpose LLM gpt-oss-120b took 4.7 seconds while supporting only two classification labels.

Scatter plot of Decision Index vs Latency: Clef models along the Pareto frontier Figure: Jev Decision Index vs. latency comparison, placing the Clef suite on the Pareto efficiency frontier. Source: Cloudflare Blog (self-reported)

Scrutinizing the Scorecard

A dose of skepticism is warranted. Across the scatter plot, data points are flagged as “Cloudflare (self-reported).” As The Register verified, these figures have not yet gone through the official Hugging Face Decision Index board validation pipeline. The blog’s visual distinction between “Jev (closed)” and “Open models (board-validated)” underscores that self-reported claims and verified results are two different standards.

Pricing is another telling indicator. At $0.24 per million tokens, Clef costs nearly six times as much as Jev ($0.042/M), while Clef-flash ($0.09/M) runs more than double. As HN commenters noted, the Pareto frontier conveniently omits cost. While Clef appears dominant on accuracy and latency alone, factoring in the dollar sign redraws the frontier entirely: the performance lead is real, but you pay a premium for it.

Local deployment hurdles remain substantial. Cloudflare Product Manager Michelle Chen confirmed to The Register that Clef demands 85GB of VRAM, with Clef-flash requiring 41GB—assuming single concurrency and a 64k context window. While “open weights” make local hosting technically possible, consumer GPUs are out of the question. HN contributors pointed out that for well-defined narrow tasks, fine-tuning a BERT-style model on a laptop in an hour delivers lower latency than any hosted API. General-purpose decision models shine primarily when domain-specific labeled data is scarce—that is their true niche.

Finally, training datasets remain closed. While model weights are released under Apache 2.0, The Register confirmed the underlying training data is proprietary. The purity of the “open” label depends on whether you value weights or data.

The HN Debate: Novel Innovation or Repackaged Classifier?

The most heated debate on the 478-point Hacker News thread coalesced around the top comment: “How are so many people building decision models within days to weeks? Isn’t this an old concept?”

Top responses unpacked the technical reality: transformers natively generate probability distributions over vocabularies. Leveraging constrained decoding and structured output schemas for classification has been common practice for years—prompting a model for a single token and inspecting logprobs to rank candidates works well even on modest architectures. Several engineers went further: Jev’s innovation was largely in interface design and API packaging, making an established technique intuitive to mainstream developers. Because APIs are trivially imitated, Hugging Face saw an influx of alternatives (AutoJev, Jebadiah, Kev, and now Clef) within a fortnight of Jev’s debut.

Yet counterarguments highlighted a critical distinction: probability calibration is the genuine moat. Building a fast classifier is straightforward; ensuring output probabilities reliably correspond to real-world confidence is difficult, and Jev remains widely regarded for superior calibration. A data engineering perspective summarized it cleanly: model architecture is “the fun and easy part,” whereas curated datasets and evaluation rigor represent the real bottleneck. Without high-quality labeled data, one cannot even rigorously verify classifier accuracy.

Both viewpoints capture complementary truths. While the core mechanics are accessible, credible calibration demands massive data engineering. Cloudflare’s ability to ship an open alternative in weeks reflects fifteen years of proprietary network traffic and annotation pipelines—the asset that separates it from weekend clone projects.

The Bundled RL Platform: Cloudflare’s Strategic Play

Tucked away in the latter half of the announcement is a move far more consequential than Clef itself: Cloudflare introduced an RL fine-tuning service, allowing enterprises to adapt Clef to proprietary domains using their own traffic.

The workflow integrates existing platform primitives: AI Gateway captures production traffic for dataset curation, Workers AI generates rollouts, Containers run sandboxed reward functions, a new Trainer component updates weights, and BYO Model redeploys back to edge nodes. Cloudflare dogfoods this pipeline across Trust & Safety moderation, support ticket routing, and Bot Management classification—all narrow tasks backed by vast troves of internal labeled logs.

RL fine-tuning platform: from production traffic to model redeployment Figure: Architecture of the RL fine-tuning pipeline, capturing production events via AI Gateway and training models for edge deployment. Source: Cloudflare Blog

The business strategy is straightforward. Raw inference token sales yield modest margins. The real prize is locking the end-to-end loop—“foundation model → RL training loop → global edge deployment”—into Cloudflare’s ecosystem. Customer traffic captured by AI Gateway remains proprietary, but the operational pipeline stays anchored to Cloudflare. This directly aligns with the company’s “agent cloud” ambitions: decision models represent the highest-frequency calls on an agent’s critical path, and whoever controls that bottleneck commands the gateway to autonomous agent traffic.

Outlook

Replicating Jev in two weeks demonstrates that the barrier to building decision models is remarkably low: take an off-the-shelf Qwen base, bypass autoregressive text generation, and score schema options in parallel. Any engineering team with solid infrastructure can pull it off.

The real divergence lies ahead: whose probability calibration holds up under edge cases, whose data pipelines can sustain continuous RL fine-tuning, and whose edge network can deliver on sub-40ms latency across global geographies.

For developers, the immediate takeaway is practical: Clef’s API is drop-in compatible with Jev, and the weights are fully open. Testing your own evaluation datasets against both models takes under ten minutes—letting empirical numbers decide which best fits your operational pipeline.

Reference Links:

  • Cloudflare Blog: Introducing Clef — our open-source decision models, and new RL fine-tuning platform
  • The Register: Cloudflare tries to outplay Jev with open-weight Clef models
  • Hacker News Discussion: Clef — Open-weight decision models, and new RL fine-tuning platform
  • Hugging Face: Cloudflare/clef Model Card
  • Cloudflare Developer Docs: Workers AI Clef Documentation