Profile
Back to NewsBack
Hacker News 3 min
Reader Mode
Show HN: Benzi – A Code Intillegence/Harness Beating Claude Code and CodeGraph

Show HN: Benzi – A Code Intillegence/Harness Beating Claude Code and CodeGraph

8 hours ago

Benzi · benchmarks

Three ways we've measured it.

Apples to apples comparison: Benzi harness vs mainstream harnesses, 24 GitHub issues, 10 languages.

Benzi harness on SWE-bench Verified. 78.2% of 500 real issues resolved at under 10¢ a fix.

Apples to Apples comparison: Code Graph's code intelligence vs Benzi's AI-native code intelligence

Every task, every attempt, verbatim — nothing held back.

Learn more about Benzi: benzi.fly.dev/about

Benzi's KPI (Key Performance Indicator) — source lines read

Every harness opens more source as bugs get harder. The question is the slope. Each point is one bug; the 24 are laid out easiest to hardest, left to right.

Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty. Hover any point for the bug and its count.

source lines read · 24 bugs · lowest of the four marked
bug Benzi
Sonnet
Benzi
DeepSeek
Claude Code
Sonnet
DeepSeek Harness
DeepSeek
mux64643831,372
commons-cli39235456771
addressable1802402001,603
jsoup6117060616
yaml-cpp120383194461
cJSON17187105854
dayjs64336307610
gson7815880516
CsvHelper42390335805
semver3534685631,857
hashie153172300458
money6815726611,892
rich2555147361,421
fmt6043152641,097
sqlglot227115800777
scrapy3861,5801,2592,127
marked9058072,1674,091
http-parser4491,2011,2702,635
zod8411,4761,9563,608
quartznet4881,7071,8974,197
sqlparser6201,0071,2252,334
nlohmann/json3799291,7645,201
ts-pattern1,1301,2321,1391,910
nats-server9892,1492,5832,385
all 249,12516,40720,70443,598

Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The green figure in each row is the lowest of the four.

The same 24 bugs in the same order, with wall clock in place of lines read.

Wall clock is raw here — unlike the tables above, Benzi's per-repo index build is not subtracted, so these seconds run slightly higher than the warm figures quoted elsewhere on this page. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit.

And the same again with dollars on the vertical axis.

Priced at the published per-token rates, same run selection as the chart above it. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline — the per-bug figures behind them are in the DeepSeek table further down. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here.

cost per fix · USD at list price · lowest of the four marked
bug Benzi
Sonnet
Benzi
DeepSeek
Claude Code
Sonnet
DeepSeek Harness
DeepSeek
mux$0.39$0.036$0.25$0.023
commons-cli$0.26$0.030$0.29$0.014
addressable$0.57$0.043$0.35$0.042
jsoup$0.23$0.025$0.30$0.051
yaml-cpp$0.25$0.068$0.37$0.024
cJSON$0.27$0.036$0.44$0.073
dayjs$0.21$0.038$0.55$0.040
gson$0.19$0.015$0.45$0.022
CsvHelper$0.25$0.038$0.58$0.042
semver$0.62$0.10$1.43$0.097
hashie$0.69$0.051$0.92$0.043
money$0.79$0.036$0.95$0.052
rich$0.33$0.11$1.30$0.053
fmt$0.54$0.14$1.12$0.071
sqlglot$0.62$0.033$1.23$0.059
scrapy$0.73$0.10$1.35$0.19
marked$1.08$0.20$3.59$0.13
http-parser$0.36$3.44$0.23
zod$1.53$0.13$2.47$0.33
quartznet$0.95$0.23$3.99$0.31
sqlparser$0.72$0.14$3.22$0.33
nlohmann/json$1.03$0.17$3.68$0.44
ts-pattern$3.74$0.31$4.33$0.053
nats-server$1.99$0.23$2.94
all 24$17.96$2.66$39.54$2.70

Priced at published per-token rates. The green figure in each row is the lowest of the four; the two DeepSeek columns are cheaper largely because that model costs roughly twenty times less per token. Blank cells are the two runs that never produced a fix.

Chat with me