Naïve raises $28.5M Series A to build autonomous company infrastructure

See more ↗
← Benchmarks

Long Horizon Terminal Bench

Multi-hour terminal tasks run unattended in real sandboxes, graded by verification tests. Same model everywhere — glm-5.2 — only the harness is swapped, so every gap on this page is harness, not model.

Last run: August 2026Task suite: terminal-bench ↗

Top resolve rate · vs Claude Code, same window

81.3% · 2.7× cheaper

Vetta on the priority window posts the highest resolve rate in the field, and at the like-for-like default window finishes a task 2.7× cheaper than Claude Code ($0.2232 vs $0.5995) while completing more of them.

Efficiency leaderboard?All harness × window cells — filter by window, click a column to re-sort.

The completion window is a single field on the request — immediate answers now, priority soon, loose eventually. Only same-window rows are a like-for-like comparison.

?Show all harness × window cells, or only one completion window.

Sorted by $ / task

#HarnessWindow
1Vettaloose ◆$0.106862.5%$0.170945 min16.2k
2Hermesloose$0.161946.7%$0.346956 min10.7k
3Vettapriority ◆$0.189581.3%$0.233251 min20.5k
4Claude Codeloose$0.201568.8%$0.293158 min21.8k
5Vettaimmediate$0.223275.0%$0.297625 min18.0k
6Codexloose$0.263570.2%$0.375354 min48.2k
7Hermespriority$0.366268.8%$0.532756 min15.4k
8Claude Codepriority$0.429375.0%$0.572563 min29.9k
9Codexpriority$0.505551.4%$0.983557 min44.7k
10Claude Codeimmediate$0.599568.8%$0.872026 min21.7k
11Piimmediate$0.619756.3%$1.101630 min23.0k
12Hermesimmediate$0.684462.5%$1.095038 min11.9k
13Codeximmediate$0.764757.7%$1.325326 min34.8k
14Pipriority37.5%33 min18.6k
15Piloose37.5%34 min18.8k

◆ on the cost-efficiency frontier — no other configuration is both cheaper and higher-scoring. Dollars are what the vendor billed, not list price; cells without a verified invoice reading show no cost and are unranked on spend.

Tokens and turns?What a harness consumes per task: tokens read, tokens generated, model calls made.

What each harness actually consumes to finish a task: how much context it reads, how much it generates, and how many model calls it takes to get there.

?Filters the panels below to one completion window — immediate is the default, lowest-latency tier.

Input tokens per task

Mean tokens read per attempt, cache reads included.

Vetta
0.74M
Pi
1.01M
Codex
1.42M
Hermes
1.88M
Claude Code
2.06M

Output tokens per task

Mean tokens the model generates per attempt across the run.

Hermes
11.9k
Vetta
18.0k
Claude Code
21.7k
Pi
23.0k
Codex
34.8k

Agent turns per task

Mean model calls per attempt — how many steps a task takes.

Vetta
31
Pi
36
Hermes
42
Codex
42
Claude Code
47

Why we run this benchmark

Most benchmarks measure a model; this one measures the machinery around it. Over multi-hour horizons the harness differences compound — context budgeting, checkpointing, recovery from failed builds. Running the same model through every harness across all three completion windows isolates the one variable we sell, with Vetta benchmarked as just another row against the strongest public alternatives.

Methodology

Tasks come from the long-horizon split of terminal-bench, pinned to a fixed dataset commit. Each attempt starts from a clean containerized sandbox and runs unattended until it finishes or its window expires. Grading is binary and automatic — no partial credit, no human judging.

Every harness runs the same open-weights model — glm-5.2 — at identical sampling settings, with repeated attempts per task per window. Dollars are measured, not modelled: every figure is corrected against vendor invoices rather than list price, because rate-card accounting over-states spend by a different factor for every harness and would mis-rank them. Configurations whose invoice readings could not be verified carry no cost and are unranked on spend; attempts that failed for infrastructure reasons are excluded rather than estimated.

Want the winning configuration? Deploy Vetta or explore the CLI.