Genie Generate a free company AI assistant Try it
← Back to Blog

Terminal-Bench 4.0: Better Resource Controls Change How Coding Agents Should Be Compared

Terminal-Bench 4.0: Better Resource Controls Change How Coding Agents Should Be Compared

Key Takeaways

  • The official project index dates Terminal-Bench 4.0 to August 28.
  • The update calibrates resources, fixes tasks, and removes saturated tasks.
  • Earlier-version scores should not be silently compared as equivalent 4.0 results.
  • Model, harness, resource budget, and cost belong together in an evaluation.
BLOOMIE
POWERED BY NEROVA

Produced by Bloomie for Nerova AI using automated editorial checks. Sources used for factual claims are listed below.

Terminal-Bench 4.0 was released August 28, 2026, according to the project’s official benchmark index. The update calibrates task resources, fixes tasks, and removes saturated tasks to reduce measurement noise in coding-agent evaluation.

The release explanation describes a flat eight-hour agent timeout and fewer infrastructure-related failures. The official benchmark index establishes the release date. A score from an earlier version should not be treated as directly equivalent to a 4.0 result.

Resource limits can change a measured outcome

An agent can fail because it cannot solve a task or because the environment does not provide sufficient time or resources. The update’s calibration aims to reduce the latter influence. That helps make the score more representative of the evaluated agent rather than a hidden infrastructure limit.

Internal evaluations need the same discipline. Record time limits, CPU, memory, network access, and execution failures. A timeout should remain visible in the result, with enough information to understand whether it is an agent behavior or an environment problem.

The harness is part of the result

A model is tested through tools, prompts, and an execution harness. These determine how it gathers evidence and acts in a terminal. A leaderboard entry is therefore a model-and-system configuration, not a context-free measurement of intelligence.

When comparing releases, preserve the harness and settings or explain what changed. If a new model uses a different tool interface or larger budget, the comparison should expose that choice. Readers should not infer that a small score difference comes exclusively from the model.

Cost and success need the same denominator

A complete run can include many model turns and long execution time. Cost comparisons should account for unsuccessful attempts as well as solved tasks. A cheap request can still produce an expensive evaluation if the agent loops without making progress.

For an engineering team, accepted changes are a more useful internal denominator than generated responses. Include reviewer time and defects discovered after the agent declares success. Benchmark work and production engineering share a need for a verifiable final result, even when their environments differ.

How to use 4.0 in a model decision

Use the benchmark as one source of evidence for a shortlist. Then test scoped tasks from the actual repository, including cases that need domain knowledge or careful workflow preservation. Keep versioned inputs so later comparisons remain reproducible.

Nerova’s assessment is that the update makes benchmark methodology part of the news. It improves the basis for comparing agents, while reminding buyers to preserve versions, resource budgets, and harness details. A leaderboard can inform a decision without replacing application-specific verification.

Nerova context

Custom AI agents for business operations

Nerova builds custom AI agents for business operations. Companies use Nerova when they need AI support for customer intake, support, sales follow-up, research, website audits, internal handoffs, and workflow automation.

Nerova can help turn websites, business context, and operational workflows into practical AI systems: website chatbots, single-purpose agents, AI teams, audits, and automation workflows built around a clear business outcome.

Ask Bloomie about this article