Naïve raises $28.5M Series A to build autonomous company infrastructure

See more ↗
← Benchmarks

SWE-Bench Pro

Real repository issues resolved end-to-end, graded by hidden tests. Four harnesses ran the identical task set on three models (deepseek-v4-flash, claude-sonnet-5, claude-opus-5).

Last run: August 2026Task suite: SWE-Bench Pro ↗

Cheapest solve on every model · vs Claude Code, same model

$0.0498 / solve · 2.0× cheaper

Vetta posts the lowest cost per solved task on all three models. On deepseek-v4-flash it solves a task for $0.0498 — 2.0× cheaper than Claude Code on the same model ($0.1012) while resolving more of the suite (80.0% vs 66.7%).

Efficiency leaderboard?All harness × model cells — filter by model, click a column to re-sort.

Swapping the model changes the price of a token by two orders of magnitude; the harness decides how many of them a task needs. Only same-model rows are a like-for-like comparison.

?Show all harness × model cells, or only one model.

Sorted by $ / task

#HarnessModel
1Vettadeepseek ◆$0.039880.0%$0.049813 min21.9k
2Pideepseek$0.048280.0%$0.060314 min24.5k
3Hermesdeepseek ◆$0.064786.7%$0.074618 min25.9k
4Claude Codedeepseek$0.067566.7%$0.101211 min23.2k
5Vettasonnet$0.553866.7%$0.83035 min14.3k
6Pisonnet$0.586560.0%$0.97757 min13.4k
7Vettaopus$0.805286.7%$0.92876 min11.2k
8Claude Codesonnet ◆$1.0547100.0%$1.05478 min17.8k
9Piopus$1.2479100.0%$1.24799 min13.5k
10Claude Codeopus$1.7002100.0%$1.700210 min13.1k
11Hermessonnet$1.773666.7%$2.659115 min32.2k
12Hermesopus$3.6503100.0%$3.650322 min33.9k

◆ on the cost-efficiency frontier — no other configuration is both cheaper and higher-scoring. Dollars are rate-card token costs for the listed model.

Tokens and turns?What a harness consumes per task: tokens read, tokens generated, model calls made.

What each harness actually consumes to finish a task: how much context it reads, how much it generates, and how many model calls it takes to get there.

?Filters the panels below to one model — every harness ran all three.

Input tokens per task

Mean tokens read per attempt, cache reads included.

Vetta
1.70M
Pi
2.22M
Hermes
3.18M
Claude Code
3.62M

Output tokens per task

Mean tokens the model generates per attempt across the run.

Vetta
21.9k
Claude Code
23.2k
Pi
24.5k
Hermes
25.9k

Agent turns per task

Mean model calls per attempt — how many steps a task takes.

Vetta
51
Hermes
54
Claude Code
56
Pi
58

Why we run this benchmark

Issue resolution is the workload most teams actually buy an agent for. Running four harnesses over the identical tasks on three models — from a budget open-weights model to the priciest frontier one — shows how much of the bill is the model and how much is the machinery around it, with Vetta benchmarked as just another row against the strongest public alternatives.

Methodology

Tasks come from SWE-Bench Pro, pinned to a fixed dataset commit. Each attempt starts from a clean containerized checkout of the repository at the issue commit and runs unattended; grading is the suite's hidden tests — binary and automatic, no partial credit, no human judging.

Every harness × model cell ran the identical task set at identical sampling settings. Dollars are rate-card token costs for the listed model, so same-model rows compare exactly; attempts that failed for infrastructure reasons are excluded rather than estimated.

Want the winning configuration? Deploy Vetta or explore the CLI.