The right production AI model is the one that completes your workflow accurately, within its latency and cost constraints, with failures your team can detect and handle. The September 2026 releases widen the shortlist: GPT-6.1 Sol, Claude Sonnet and Opus 5.5, Grok 4.7, and Gemini 3.8 Flash deserve evaluation. They do not establish one universal winner. This guide is an editorial selection framework based on vendor documentation, not a Nerova hands-on benchmark.
What the latest releases change in your shortlist
OpenAI positions GPT-6.1 Sol around agentic coding, computer use, and professional work. Anthropic's Sonnet 5.5 announcement emphasizes coding efficiency, while Opus 5.5 is a separate higher-capability choice. Grok 4.7 targets coding and knowledge work; Gemini 3.8 Flash adds another general-purpose option. Restricted cyber variants are separate access decisions.
These positioning statements tell you which tasks to test first. They do not prove that a provider's model will win inside your application. Tool definitions, available context, reasoning effort, and retrieval quality can change results as much as the model choice. A release comparison becomes useful when it narrows an evaluation, not when it substitutes a leaderboard for your acceptance criteria.
Start with the business output, then choose the evaluation
| Workflow | Measure first | Failure to investigate |
|---|---|---|
| Code changes | Correct changes that pass relevant checks and human review | A plausible patch that breaks an adjacent workflow |
| Document extraction | Field accuracy, source references, and missing-value handling | Confidently invented values in incomplete documents |
| Customer operations | Correct routing, policy adherence, and useful escalation | Unauthorized promises or actions |
| Research assistance | Evidence coverage and traceable reasoning | Unsupported conclusions dressed as sourced analysis |
| Computer-use tasks | Successful completion and safe recovery from interruptions | Duplicate actions after a partial failure |
For each workflow, assemble representative successful cases, difficult cases, and examples where the agent should stop. Use de-identified or synthetic inputs where possible. Decide in advance what counts as an acceptable answer or action. Otherwise the evaluation can reward whichever model writes the most convincing explanation instead of the one that produces the correct outcome.
Keep model comparisons fair enough to make a decision
Give candidates equivalent task information and tool access. Record the exact model version, reasoning settings, harness, retry policy, and evaluation date. If one provider is tested with a specialist coding agent and another through a bare text prompt, label that as a comparison of systems. It is not a clean model comparison.
Evaluate more than an average success rate. Inspect the errors that could cause customer harm, data corruption, or expensive human rework. A model that misses an optional formatting detail and one that writes to the wrong customer record have not failed in equivalent ways. Break out those categories before deciding whether a small aggregate-score advantage matters.
Vendor results remain useful as evidence of tested capability. They are a hypothesis about your workload. Any apparent lead should survive your own inputs, relevant tools, and review process before you route valuable production work to it.
Compare completed-work cost and latency
Token prices are only part of cost. Include input and output volume, eligible cached input, tool charges, repeated attempts, orchestration, and reviewer time. Track cost per accepted result. A lower token price can still produce a more expensive workflow if it needs longer outputs or frequent corrections.
For interactive applications, measure time to a useful response and the slow tail of the latency distribution. For background jobs, completion time and throughput may matter more. Do not pay for a premium speed tier by default if a nightly batch already finishes comfortably inside its deadline.
Make a narrow deployment decision
Select one workflow and one candidate that satisfies its quality, permission, latency, and budget requirements. Deploy to a limited slice with logs that reveal failure categories and with a clear way to return to the previous configuration. Keep the evaluation cases versioned so subsequent releases can be compared against the same baseline.
A more complex routing system is justified only when measured workload differences warrant it. If one model handles the task well, adding three providers, several fallbacks, and hidden retries creates more operational responsibility without automatically creating value. The practical choice is the smallest system that reliably delivers the required result.