# Clef and Strands Decider: Open Models Move the Fight to Workflow Control

**Plutonous** | October 3, 2026 | 7 min read

> Cloudflare and Strands launched decision models on October 1. The commercial contest is shifting toward deployment, calibration and who owns the feedback loop.

Tags: Cloudflare, Clef, Strands Decider, Decision Models, Open Models, AI Agents, Model Evaluation, AI Infrastructure

---

**TL;DR: Cloudflare and Strands announced decision models on October 1, opening another front in the market Jev helped popularize.<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-2">[2]</a></sup> Cloudflare lists Clef at $0.24 and Clef-flash at $0.09 per million input tokens; the bigger business question is who operates, evaluates and improves the decisions those tokens produce.<sup><a href="#source-5">[5]</a></sup><sup><a href="#source-6">[6]</a></sup>**

The real story isn't another model claiming it can pick the right label. It is the emerging competition to become the layer that chooses what an application does next. When classification moves into tool selection, escalation and routing, the supplier sits beside the application's operating logic.

Cloudflare's Clef family and Strands Decider arrived on the same day with different commercial implications. Clef has a hosted route into Workers AI. Strands puts local experimentation at the center of its announcement. This October 3 analysis examines that distribution contest; our [Jev evaluation article](/news/jev-benchmarks-decision-reliability) covers the reliability problem, and our [support-routing workflow](/news/jev-support-routing-workflow) shows a bounded application.

The AWS connection is specific. AWS introduced Strands Labs in February as an organization for experimental projects, with participation open to Amazon development teams and a release cycle separate from the production SDK.<sup><a href="#source-10">[10]</a></sup> Decider is a Strands Labs model release. Calling it a new managed Bedrock service would misdescribe the announcement.


### Why This Matters Now

An open model can make the inference supplier replaceable. A proprietary collection of labeled outcomes, escalation rules and production integrations can still make the surrounding platform expensive to leave.


*Cover: Generated editorial illustration of mechanical switches choosing paths. It represents workflow routing, not measured model behavior.*

## The Product: A Decision Layer Becomes a Buying Choice

Clef's model card describes a 27B model built from Qwen3.8-27B, with a vision encoder and Apache-2.0 licensing. Clef-flash uses Qwen3.5-9B and the same license. Both cards describe probability outputs over allowed answers and compatibility with the Jev/SystemOne API.<sup><a href="#source-3">[3]</a></sup><sup><a href="#source-4">[4]</a></sup>

Strands Decider's architecture documentation identifies a 1.9B backbone and a pointer head that scores supplied options. The question defines the available labels.<sup><a href="#source-8">[8]</a></sup> The launch releases v19 and emphasizes local execution and accessible training materials.<sup><a href="#source-2">[2]</a></sup>

| Documented option | Practical distinction | Purchasing implication |
| --- | --- | --- |
| Clef, 27B | Multimodal weights; Cloudflare hosts a 65,536-token context model.<sup><a href="#source-3">[3]</a></sup><sup><a href="#source-5">[5]</a></sup> | Evaluate managed inference alongside operating the weights yourself. |
| Clef-flash, 9B | Smaller multimodal variant; hosted context is also 65,536 tokens.<sup><a href="#source-4">[4]</a></sup><sup><a href="#source-6">[6]</a></sup> | Test whether its errors justify the lower input price. |
| Strands Decider 2B | Local experimentation with a 1.9B backbone.<sup><a href="#source-2">[2]</a></sup><sup><a href="#source-8">[8]</a></sup> | Budget for your serving infrastructure and evaluation work. |

*Specifications come from the respective vendors and project maintainers. This table compares deployment choices, not performance; LLM Rumors has not run these models.*

Compatibility lowers integration work, but a shared response shape does not make two models interchangeable. The same support ticket can receive different labels, and the same numerical confidence can describe a different error pattern. Buying a replacement therefore requires replaying real cases, not merely changing an endpoint.

## The Economics: A $150 Saving Can Disappear in One Workflow

At Cloudflare's published input rates, an illustrative workload of 1,000,000 requests with 1,000 billed input tokens each contains 1,000,000,000 input tokens. The input charge is $240 for Clef or $90 for Clef-flash, a $150 difference. These are calculations from the listed US-dollar rates, not measured customer bills.<sup><a href="#source-5">[5]</a></sup><sup><a href="#source-6">[6]</a></sup>

That example excludes retries, other services, human review and the consequences of mistakes. If one additional error costs $150 to resolve, it consumes the entire input saving in this hypothetical workload. A procurement team that compares only the token price is optimizing the smallest visible invoice rather than the operating result.

Self-hosting changes the invoice rather than abolishing it. Hardware utilization, maintenance and staff time remain costs even when the weights are downloadable. Low request volume may favor a managed endpoint; sustained traffic and an existing operations team may justify local infrastructure. The right comparison needs the organization's actual workload and staffing assumptions.

Here's the genius of distributing the model openly: a vendor can reduce resistance to adoption while competing to run everything around it. That is our reading of the strategy, not a disclosed revenue forecast.

## The Confidence: Coverage Is Part of the Product

The Strands maintainers report 95.2% accuracy in v19's confidence band at or above 0.9, but that band covers 23% of their unseen short-classification sample. They explicitly limit the interpretation to that task family and describe different calibration behavior on longer documents and answer-adequacy judgments.<sup><a href="#source-7">[7]</a></sup>

Those two percentages belong together. A classifier can perform well on the subset it accepts while leaving most work for another route. A buyer needs both the accepted subset's error rate and the proportion of traffic that reaches it. Neither figure alone tells the finance team how much labor disappears.

What's often overlooked is that confidence is also an interface decision. Strands documents different confidence calculations for Choice and Score; Noul returns a probability without a separate confidence field.<sup><a href="#source-8">[8]</a></sup> A generic rule that reads one field across every primitive can misinterpret the output before model quality even enters the discussion.

For a migration, preserve the original cases and downstream outcomes. Measure how many actions change, which previously correct cases regress, and which cases now require review. Recalibrate thresholds using representative traffic. An API-compatible response is the beginning of an evaluation, not its conclusion.

## The Openness: Failed Experiments Are Useful Product Evidence

The Strands research log records five versions promoted despite missing their preregistered bars. It also records v20 failing four predictions and leaving v19 as the reference recipe.<sup><a href="#source-9">[9]</a></sup> That record is more useful to an engineering buyer than a release sequence in which every experiment is presented as an uncomplicated win.

Open artifacts permit inspection; they do not supply an independent replication. A team still has to reproduce the behavior it intends to rely on. But a visible record of failed hypotheses gives that team somewhere concrete to start: which changes helped, which tradeoffs the maintainers accepted, and which result did not justify replacement.

Let's be clear: a bounded answer set is not an authorization system. A model can select a valid option for the wrong reason. Access controls, transaction limits and approval rules belong in application code, with the model supplying evidence to that policy. An open checkpoint does not alter that division of responsibility.

## The Strategy: Own the Feedback Before Choosing the Host

Cloudflare's announcement offers hands-on fine-tuning with its forward-deployed engineering team. The self-service platform is a later plan, not a capability readers should assume is already available.<sup><a href="#source-1">[1]</a></sup> The strategic direction is clear: move from hosting a generic decision model toward helping customers improve it for their own traffic.

Our recommendation is to retain the portable assets first: question definitions, allowed answers, labeled cases, observed outcomes and the policy that consumes each result. Keep those versioned independently of the inference provider. A provider switch becomes less painful when the organization owns the acceptance test and can run it again.

For an initial purchase, choose one reversible workflow. Compare providers on that same case set, include abstentions and failures, and price the complete path through review and recovery. Expand only when those results support the next use case. This is a buying discipline, not a claim that either release has already passed it.


### The Key Insight

Open weights create an exit route for inference. Portability also requires control over the evaluation data, decision policy and feedback loop surrounding the model.


The uncomfortable truth is that cheaper decisions make operational ownership more valuable. Cloudflare and Strands are expanding the supply of models. The durable advantage belongs to whoever can turn a decision into an outcome, measure the mistake and improve the next attempt.


## Sources

<a id="source-1"></a>
1. [Clef launch and fine-tuning roadmap](https://blog.cloudflare.com/clef-decision-models/)

<a id="source-2"></a>
2. [Introducing Strands Decider 2B](https://strandsagents.com/blog/introducing-strands-decider/)

<a id="source-3"></a>
3. [Clef model card](https://huggingface.co/Cloudflare/clef)

<a id="source-4"></a>
4. [Clef-flash model card](https://huggingface.co/Cloudflare/clef-flash)

<a id="source-5"></a>
5. [Clef hosted model documentation](https://developers.cloudflare.com/workers-ai/models/clef/)

<a id="source-6"></a>
6. [Clef-flash hosted model documentation](https://developers.cloudflare.com/workers-ai/models/clef-flash/)

<a id="source-7"></a>
7. [Strands Decider evaluation results](https://github.com/strands-labs/strands-decider/blob/ddd11994bc451ffc78fa30f43037c291fd4d44b6/evaluation/results.md)

<a id="source-8"></a>
8. [Strands Decider architecture](https://github.com/strands-labs/strands-decider/blob/ddd11994bc451ffc78fa30f43037c291fd4d44b6/docs/architecture.md)

<a id="source-9"></a>
9. [Strands Decider research record](https://github.com/strands-labs/strands-decider/blob/ddd11994bc451ffc78fa30f43037c291fd4d44b6/research/README.md)

<a id="source-10"></a>
10. [Introducing Strands Labs](https://aws.amazon.com/blogs/opensource/introducing-strands-labs-get-hands-on-today-with-state-of-the-art-experimental-approaches-to-agentic-development/)


*Last updated: October 3, 2026*

---

*Source: [LLM Rumors](https://www.llmrumors.com/news/open-decision-models-clef-strands)*
