OpenAI simplified its paid API usage tiers to Build, Launch and Grow on October 6, 2026. Organizations move between tiers as total credit purchases reach the relevant threshold. The changelog records the change. For operators, the useful distinction is between account qualification, traffic limits and the controls that stop spending.
Purchase thresholds determine tier qualification
The rate-limit guide lists total credit-purchase thresholds of $5 for Build, $100 for Launch and $500 for Grow. It also lists monthly usage limits separately. These thresholds are not subscription prices and do not mean a model request has a fixed cost.
An organization’s tier can change without an application’s request pattern changing. Conversely, a workload can grow while the organization remains in the same tier. Capacity planning should therefore use the actual model limits displayed for the account rather than assume that a label establishes enough throughput for a launch.
Request and token limits constrain different workloads
A service making many short calls can encounter a request limit before a token limit. A document-processing job can encounter the opposite pattern. Record both dimensions when estimating traffic, including background work that competes with user-facing requests.
Peak demand matters more than a daily average. A scheduled batch starting at the same time as a product notification can consume available capacity unexpectedly. Use bounded concurrency and a queue where the work can wait, and make rate-limit failures visible to the component responsible for retrying them. Retrying every call immediately can worsen the overload.
Spend alerts and hard limits are different controls
The documentation distinguishes a spend alert, which allows traffic to continue, from a hard spend limit that rejects affected requests at the configured cap. The account’s monthly usage limit is another separate concept. Confirm which control is active before relying on it to constrain a budget.
Test how the application behaves when a limit is reached. A failed model call should not be reported as a completed job, and a partial workflow should not repeat external actions blindly. The simpler tier structure makes qualification easier to understand, but production readiness still depends on the account’s verified limits and the application’s behavior under constrained capacity.