The reference

One benchmark.
One global frontier.

DeepSWE summarizes performance across 113 software-engineering tasks. But the global view assumes every kind of repository matters equally.

Global Pareto

113 tasks · 91 repositories · 51 configurations

DeepSWE score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

DeepSWE tasks span repos with varied stars, sizes, and languages

91 repositories. Dot size is task count; color is primary language.

1001k10k1k10k100kGitHub starsFiles in default branch
TypeScriptGoPythonJavaScriptRust

Shade 01 · Popularity

Shades of SWE by GitHub stars

Repository popularity separates DeepSWE into three equally readable views of the cost–performance frontier.

≤2,000 stars

43 tasks · 34 repositories

≤2,000 stars task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

2,001–8,000

34 tasks · 30 repositories

2,001–8,000 task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

>8,000 stars

36 tasks · 27 repositories

>8,000 stars task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMclaude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

These slices are descriptive. Stars may also reflect repository age, language, domain, or project type.

Shade 02 · Scale

Shades of SWE by repository size

As codebases grow, agents must navigate and coordinate changes across more files. The frontier shifts with that scale.

≤200 files

35 tasks · 31 repositories

≤200 files task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

201–600

39 tasks · 28 repositories

201–600 task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

>600 files

39 tasks · 32 repositories

>600 files task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXkimi-k3MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Shade 03 · Ecosystem

Shades of SWE by language

Language and ecosystem change the mix of work in the benchmark—and the models that offer the best tradeoff.

TypeScript

35 tasks · 28 repositories

TypeScript task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Go

34 tasks · 29 repositories

Go task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Python

34 tasks · 28 repositories

Python task score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMclaude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Rust and JavaScript are omitted because each has only 5 tasks, which is too few for stable slice-level estimates.

Your repository

Pick your shade of SWE

Paste a public GitHub repository. We’ll inspect its language, size, and stars, then use similar DeepSWE repositories to draft a personalized Pareto frontier.

Experimental model

The result is a draft, not a measured benchmark run. We’ll show the repositories contributing to the estimate and how much support it has.

We only inspect public repository metadata. No sign-in required.

Language AutomaticFiles AutomaticStars Automatic