OpenAI introduced GPT-6.1 Sol on September 29, 2026, as an upgrade for coding, computer use, and professional work. Standard API pricing is $2 per million input tokens and $10 per million output tokens; cached input is $0.10 per million. Its claimed proximity to GPT-6 Astra should be evaluated against the work your application actually performs.
Price advantage and capability claims are separate evidence
The announcement reports near-Astra performance on selected evaluations at one-fifth of Astra’s standard token prices. It makes the model available in the API, ChatGPT Work, and Codex for eligible plans, while noting it is not yet in Chat. Sol Ultrafast is described as forthcoming rather than part of the initial availability.
A token-price ratio does not establish a task-price ratio. A lower-cost model can take more steps, call more tools, or require more review. Compare the total cost of an accepted result and keep the premium model for task classes where it materially changes that result.
Safety evaluations do not replace application checks
The system-card addendum documents vendor safety evaluation context. An application still needs its own tests for permission boundaries, correct citations, and behavior when tools fail.
Include an unavailable source, a revoked connection, and a misleading instruction in retrieved content. Look for explicit reporting of missing evidence rather than an answer filled in from memory. For coding work, require tests and a reviewable diff; for professional document tasks, check numbers against their original tables.
Separate model comparisons from a recent vision fix
OpenAI’s September 25 changelog records a fix for image encoding that degraded GPT-6 Sol and Luna image understanding. Evaluations from before that correction are not a clean baseline for a new vision comparison.
Rerun the earlier models on the same screenshots or documents before concluding that all improvement comes from GPT-6.1 Sol. Record evaluation date, input preprocessing, effort, and model identifier with the result. A benchmark without this context can perpetuate a resolved defect as a permanent capability judgment.
A practical routing decision
Start with a bounded category such as routine code fixes or document extraction. Preserve the current prompts and tools, compare quality blind where practical, and route difficult cases deliberately. The launch makes a lower-cost tier worth testing; adoption should follow demonstrated reliability rather than a claim that one model can replace every tier.