SpaceXAI released Grok 4.6 on August 12, 2026, with a focus on long-running agents, coding, and visual or interactive work. For teams evaluating it, the central question is whether those changes reduce the cost and supervision needed to complete a real task in their existing workflow.
Launch availability and pricing
The announcement lists availability in Cursor, Grok Build, and the API, with additional distribution partners. It gives starting token prices of $2 per million input tokens and $6 per million output tokens, with a fast variant at twice that price. A one-week included-usage promotion was separate from those API rates.
Use those figures as launch context and confirm the applicable endpoint and current billing before deployment. A model served through a partner can have different limits, feature support, and commercial terms. The API documentation is the implementation reference; a product's model picker does not establish every API capability.
What the capability claims mean
SpaceXAI reports gains over Grok 4.5 on several coding and professional-work evaluations, and describes more sustained execution and self-checking. Its comparison table draws competitor figures from published material. That supports reporting the company's claim, but it is not a matched evaluation conducted by Nerova.
The relevant work is often more specific than a leaderboard category. A model that builds an attractive first version may still struggle with authorization, persistence, or regressions in an existing application. Separate visual fidelity from functional correctness, and evaluate both if the intended use combines them.
Test longer work under a controlled budget
Choose representative tasks that include the difficult transitions: locating the relevant code, making a minimal change, handling a failed test, and explaining the final diff. Hold the repository, instructions, tools, and acceptance criteria constant across candidate models. Otherwise a stronger scaffold can be mistaken for a stronger model.
Set bounds for runtime, output, and tool calls. Record how frequently the agent restarts, repeats an unhelpful action, or asks for human assistance. A long trajectory that eventually succeeds can still be operationally unsuitable when it monopolizes compute or creates an expensive review burden.
Adopt where the improvement survives review
Start with reviewed tasks in a disposable development environment. Keep production secrets and destructive actions outside the evaluation. Have the same checks validate both the candidate's result and the current model's result, including tests that the generated patch did not write itself.
A switch is justified when quality, elapsed time, and total cost improve for the targeted work without weakening control. If Grok 4.6 improves only one task family, route that family explicitly rather than replacing every model call. This is a release to evaluate with evidence, not a universal migration instruction.
Cloud and coding-tool distribution followed the launch
The official dated release index records Grok 4.6 arriving in GitHub Copilot on August 14, Amazon Bedrock on August 19, Gemini Enterprise Agent Platform on August 21 and Microsoft Foundry on August 26. These are distribution milestones for the same model, rather than four new model launches.
A team already using one of those platforms may be able to evaluate the model within its established billing and access workflow. Verify the exact model version, regional availability, tools and pricing on that route. A model name appearing in two providers does not establish that their service contracts or supported features are interchangeable.