Open decision models gained several concrete release paths at the start of October 2026: Strands Decider 2B and Cloudflare Clef on October 1, followed by llama.cpp server support on October 2. These models score predefined choices rather than generating an unrestricted answer. Together, the releases make bounded decisions easier to evaluate across hosted and local workflows.
Strands Decider targets small local decisions
The Strands announcement introduces a 2B model for experimentation and local development. Its official card lists Apache 2.0 and candid limitations: long multi-step documents are difficult, question phrasing can have limited effect, and confidence calibration needs measurement on the user’s own traffic.
That is useful guidance for choosing a first task. Short request routing is a better starting point than a complex legal judgment assembled from many documents. Keep the category definitions clear and include an option for cases the workflow cannot handle. Small models can reduce serving requirements, but a low resource cost does not make an incorrect decision cheap.
Clef adds a hosted and downloadable multimodal option
Cloudflare’s release introduces Clef and Clef-flash on Workers AI. The Clef card describes an Apache 2.0 multimodal checkpoint and a SystemOne-compatible interface. Cloudflare also offers assisted fine-tuning work; the announcement describes a future self-serve platform, not an already completed self-service product.
Image or video input creates additional evaluation requirements. A classification can depend on text visible only in part of a screenshot or on an event that occurs between selected video frames. Preserve the input representation when inspecting a wrong answer. Test whether the workflow still behaves acceptably when an image is resized, a field is absent or the supplied options use unfamiliar wording.
llama.cpp adds a shared local serving interface
The ggml-org announcement introduces /v1/systemone, which returns probabilities for allowed options. The implementation pull request merged October 2. Supported models have different sizes, modalities and licenses; the serving interface does not make every checkpoint commercially permissive.
A shared endpoint can reduce integration work when comparing models, but it does not guarantee identical interpretation of a question. Run the same labeled examples through each candidate and inspect disagreements. Pin the server build and checkpoint so that a conversion or runtime update does not silently change the comparison.
Keep policy and authorization outside the model
A finite output simplifies parsing and makes application code easier to reason about. It does not prove the selected choice is correct. Define what happens below an accepted confidence threshold and measure the errors that matter to the workflow, including uncommon cases.
Routing, evaluating an agent step and authorizing an action are separate responsibilities. A classifier can recommend a queue or indicate that a screenshot resembles success; the application should independently verify permissions and consequential state changes. These October releases are useful building blocks for hybrid agents that reserve larger-model reasoning for harder tasks, provided the smaller decision layer has its own evaluation and escalation path.