Can agents design libraries other agents can use?

One agent designs a library. Other agents solve problems with it. The library is scored on how correct and how short their code is.

Model · click to see runsScore ± 95% CILibrary $
Opus 5.5mini-SWE
Opus 5.548.9, 95% CI 48.0 to 49.9
48.9 ± 0.9
$9.62
Fable 5.1mini-SWE
Fable 5.147.5, 95% CI 46.4 to 48.5
47.5 ± 1.1
$14.81
GPT-6 Astramini-SWE
GPT-6 Astra45.1, 95% CI 44.3 to 45.8
45.1 ± 0.7
$3.63
GPT-6 AstraCodex
GPT-6 Astra44.1, 95% CI 43.4 to 44.9
44.1 ± 0.8
$4.51
Kimi K3mini-SWE
Kimi K344.0, 95% CI 43.0 to 45.0
44.0 ± 1.0
$8.70
GLM 5.3mini-SWE
GLM 5.341.9, 95% CI 40.8 to 42.9
41.9 ± 1.1
$21.06
Fable 5.1Claude Code
Fable 5.139.9, 95% CI 37.4 to 42.4
39.9 ± 2.5
$14.39
Grok 4.6mini-SWE
Grok 4.639.7, 95% CI 38.6 to 40.7
39.7 ± 1.1
$2.63
GPT-5.6 SolCodex
GPT-5.6 Sol39.5, 95% CI 38.5 to 40.5
39.5 ± 1.0
$2.14
GPT-6 Solmini-SWE
GPT-6 Sol38.5, 95% CI 37.8 to 39.2
38.5 ± 0.7
$0.30
DeepSeek V4 Promini-SWE
DeepSeek V4 Pro31.2, 95% CI 30.4 to 32.1
31.2 ± 0.8
$0.31
Reference arms: the same implementers with no library, and with the task's production library
PProduction library
Production library46.6, 95% CI 46.1 to 47.2
46.6 ± 0.5
—
–No library
No library34.4, 95% CI 33.9 to 34.9
34.4 ± 0.5
—
Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. The error bars and ± show the 95% confidence interval.