Genie Generate a free company AI assistant Try it
← Back to Blog

Open decision models arrive: Strands Decider, Clef and llama.cpp support

Open decision models arrive: Strands Decider, Clef and llama.cpp support

Key Takeaways

  • Strands Decider and Clef launched October 1; llama.cpp support followed October 2.
  • Model licenses, modalities and calibration differ despite a shared interface.
  • Finite outputs simplify parsing but do not establish correctness or authorization.
BLOOMIE
POWERED BY NEROVA

Produced by Bloomie for Nerova AI using automated editorial checks. Sources used for factual claims are listed below.

Open decision models gained several concrete release paths at the start of October 2026: Strands Decider 2B and Cloudflare Clef on October 1, followed by llama.cpp server support on October 2. These models score predefined choices rather than generating an unrestricted answer. Together, the releases make bounded decisions easier to evaluate across hosted and local workflows.

Strands Decider targets small local decisions

The Strands announcement introduces a 2B model for experimentation and local development. Its official card lists Apache 2.0 and candid limitations: long multi-step documents are difficult, question phrasing can have limited effect, and confidence calibration needs measurement on the user’s own traffic.

That is useful guidance for choosing a first task. Short request routing is a better starting point than a complex legal judgment assembled from many documents. Keep the category definitions clear and include an option for cases the workflow cannot handle. Small models can reduce serving requirements, but a low resource cost does not make an incorrect decision cheap.

Clef adds a hosted and downloadable multimodal option

Cloudflare’s release introduces Clef and Clef-flash on Workers AI. The Clef card describes an Apache 2.0 multimodal checkpoint and a SystemOne-compatible interface. Cloudflare also offers assisted fine-tuning work; the announcement describes a future self-serve platform, not an already completed self-service product.

Image or video input creates additional evaluation requirements. A classification can depend on text visible only in part of a screenshot or on an event that occurs between selected video frames. Preserve the input representation when inspecting a wrong answer. Test whether the workflow still behaves acceptably when an image is resized, a field is absent or the supplied options use unfamiliar wording.

llama.cpp adds a shared local serving interface

The ggml-org announcement introduces /v1/systemone, which returns probabilities for allowed options. The implementation pull request merged October 2. Supported models have different sizes, modalities and licenses; the serving interface does not make every checkpoint commercially permissive.

A shared endpoint can reduce integration work when comparing models, but it does not guarantee identical interpretation of a question. Run the same labeled examples through each candidate and inspect disagreements. Pin the server build and checkpoint so that a conversion or runtime update does not silently change the comparison.

Keep policy and authorization outside the model

A finite output simplifies parsing and makes application code easier to reason about. It does not prove the selected choice is correct. Define what happens below an accepted confidence threshold and measure the errors that matter to the workflow, including uncommon cases.

Routing, evaluating an agent step and authorizing an action are separate responsibilities. A classifier can recommend a queue or indicate that a screenshot resembles success; the application should independently verify permissions and consequential state changes. These October releases are useful building blocks for hybrid agents that reserve larger-model reasoning for harder tasks, provided the smaller decision layer has its own evaluation and escalation path.

Nerova context

Custom AI agents for business operations

Nerova builds custom AI agents for business operations. Companies use Nerova when they need AI support for customer intake, support, sales follow-up, research, website audits, internal handoffs, and workflow automation.

Nerova can help turn websites, business context, and operational workflows into practical AI systems: website chatbots, single-purpose agents, AI teams, audits, and automation workflows built around a clear business outcome.

Ask Bloomie about this article