Genie Generate a free company AI assistant Try it
← Back to Blog

SWE-2 and Devin Local Fusion: Evaluate the Coding Model and Harness Together

SWE-2 and Devin Local Fusion: Evaluate the Coding Model and Harness Together

Key Takeaways

  • SWE-2 launched September 10; Local Fusion reached Desktop and CLI September 11.
  • Fusion separates a planning/review lead from an execution sidekick.
  • Reported cost savings apply to tested model and harness combinations.
  • Compare accepted changes, rework, review time, and total task cost.
BLOOMIE
POWERED BY NEROVA

Produced by Bloomie for Nerova AI using automated editorial checks. Sources used for factual claims are listed below.

Cognition introduced SWE-2 on September 10, 2026, then brought its Fusion harness to Devin Desktop and CLI on September 11. The two releases should be evaluated together: SWE-2 supplies a coding model, while Fusion coordinates a stronger planning-and-review model with a less expensive execution model.

The SWE-2 announcement describes availability in Desktop and CLI, with Web and Fusion rollout. The Local Fusion announcement recommends Fable 5.1 as lead with SWE-2 as sidekick. Reported efficiency results apply to evaluated combinations, not every repository or every pair of models.

Why the model is only part of a coding-agent result

A coding model operates inside an environment that determines which files it can inspect, which tools it can use, how tests run, and when the task ends. Comparing two models through different harnesses can confound model capability with tool behavior or task setup.

Cognition’s approach makes the division of work explicit. Planning and review can use a model with stronger reasoning, while execution uses another model. That is a useful architecture to investigate when a team spends much of its agent budget on routine edits, but it creates coordination work that also has to be measured.

Read the benchmark table as a set of tradeoffs

Cognition reports SWE-2 results across several coding evaluations, including FrontierCode, DeepSWE, and Terminal-Bench. Its own table shows that relative performance varies by benchmark. A headline about approaching frontier performance does not establish equality on every difficult engineering task.

For an internal evaluation, hold task instructions, repository snapshots, tools, time limits, and review criteria constant. Measure accepted output and the effort needed to fix it. Record abandoned tasks as well as successes. If the agent solves a problem only after a reviewer supplies the central insight, that collaboration should remain visible in the result.

Local Fusion changes the cost conversation

The Fusion release reports savings for tested lead-and-sidekick combinations, sometimes alongside lower scores. That is exactly why cost per request is insufficient. A team needs to know the cost of an accepted change, including the lead’s planning, sidekick execution, retries, and final review.

A local development workflow also needs a clear workspace boundary. Separate concurrent tasks so one agent cannot accidentally consume another task’s unfinished edits. Maintain a reviewable diff and run the project’s actual checks. Permission to edit code should not automatically include permission to deploy it.

When a team should try the combination

Start with maintenance tasks whose outcomes can be checked: a scoped bug, a dependency migration, or a performance change with a reproducible baseline. Include a few tasks that require deeper judgment so the evaluation reveals whether the lead model adds useful supervision rather than overhead.

Keep the current workflow as the reference point. Fusion is valuable if it reduces complete cost or reviewer effort while preserving quality. A faster initial edit is not enough when it produces regressions or leaves the verification burden to a human.

Nerova’s assessment is that the releases make model-and-harness evaluation more important. SWE-2 is an option worth measuring; Fusion is a coordination design worth testing. Neither warrants a blanket claim that autonomous coding can replace engineering ownership.

Nerova context

Custom AI agents for business operations

Nerova builds custom AI agents for business operations. Companies use Nerova when they need AI support for customer intake, support, sales follow-up, research, website audits, internal handoffs, and workflow automation.

Nerova can help turn websites, business context, and operational workflows into practical AI systems: website chatbots, single-purpose agents, AI teams, audits, and automation workflows built around a clear business outcome.

Ask Bloomie about this article