One DeepSWE
Many Shades of SWE

When I look at the DeepSWE leaderboard...

DeepSWE score0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

...I wonder how many of the 113 tasks/91 repos are related to MY work as a SWE?

1001k10k1k10k100kGitHub starsFiles in default branch
TypeScript · 27Go · 28Python · 27JavaScript · 4Rust · 5

Shade 01 · Popularity

Shades of SWE by GitHub stars

When I work on niche repos, which model should I use?

deepseek-v4-flash performs surprisingly well on niche repos: 18.3× cheaper than gpt-5.6-sol (medium) at 1.2× the score.

≤2,000 stars task score43 tasks / 34 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on more mainstream repos, which model should I use?

opus 5 works the best on more mainstream repos, delivering a 5.7 point improvement at 1.6× the cost of gpt-5.6-sol (high).

2,001–8,000 stars task score34 tasks / 30 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on fairly established repos, which model should I use?

gpt-5.6-sol leads on fairly established repos, while terra can be a good bargain if I’m willing to trade off 6.6% score for a 53.3% cost reduction.

>8,000 stars task score36 tasks / 27 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMclaude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Shade 02 · Scale

Shades of SWE by repo size

When I work on small repos, which model should I use?

Reasoning longer sometimes hurts performance in small repos. gpt-5.6-sol (high) hits the sweet spot of perf and cost here.

≤200 files task score35 tasks / 31 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on medium repos, which model should I use?

opus 5 leads in medium repos, but make sure to set reasoning to medium as well, otherwise your CC usage may drain away 1.9× faster without perf gains.

201–600 files task score39 tasks / 28 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on large repos, which model should I use?

kimi k3 appears on the pareto frontier for large repos. If you want the extra 3.7 points, go with opus 5 (xhigh), but be ready for a 1.8× increase in cost.

>600 files task score39 tasks / 32 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXkimi-k3MAXgpt-5.6-solMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Shade 03 · Ecosystem

Shades of SWE by language

When I work on TypeScript repos, which model should I use?

Most frontier models perform similarly well in TypeScript. Not surprising since TS is the most popular language on GitHub. gpt-5.6-sol delivers the best perf. Personally, I’d go with luna since it’s within 9.4% of sol but 5.9× cheaper.

TypeScript task score35 tasks / 28 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMgpt-5.6-terraMEDIUMgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on Go repos, which model should I use?

Performance is more spread out here. Go for opus 5 (xhigh) to get the best perf. luna is not a bad choice at 93.2% cheaper with a slight 8.4% lower perf.

Go task score34 tasks / 29 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗claude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

When I work on Python repos, which model should I use?

gpt-5.6-sol and opus are both fine at Python or you can also code by hand once in a while to keep that Two Sum solution fresh in your mind.

Python task score34 tasks / 28 repos0%20%40%60%80%$0$5.00$10$15Avg cost per taskmost efficient ↗gpt-5.6-solMEDIUMclaude-opus-5MAXgpt-5.6-lunaMEDIUMdeepseek-v4-flashMAX

Rust and JavaScript are omitted because each has only 5 tasks, which is too few for stable slice-level estimates.

Find Your Shade of SWE

  1. 01

    Let a router choose

    Try a model router like Cursor Router, Devin Fusion, or OpenRouter’s Auto Router, just to name a few. Picking models and reasoning efforts by hand is tricky. Let another model do it for you.

  2. 02

    Benchmark your repo

    Try converting your repo into a custom benchmark with services like Vals-Smith. Measure first, then optimize.

  3. 03

    Benchmark your workload

    Try making a benchmark for workloads your organization cares about. The react-doctor creators made ReactBench to measure agents on realistic React work. It takes more work, but pays off when the workload powers your organization.