One DeepSWE
Many Shades of SWE
...I wonder how many of the 113 tasks/91 repos are related to MY work as a SWE?
Shade 01 · Popularity
Shades of SWE by GitHub stars
When I work on niche repos, which model should I use?
deepseek-v4-flash performs surprisingly well on niche repos: 18.3× cheaper than gpt-5.6-sol (medium) at 1.2× the score.
When I work on more mainstream repos, which model should I use?
opus 5 works the best on more mainstream repos, delivering a 5.7 point improvement at 1.6× the cost of gpt-5.6-sol (high).
When I work on fairly established repos, which model should I use?
gpt-5.6-sol leads on fairly established repos, while terra can be a good bargain if I’m willing to trade off 6.6% score for a 53.3% cost reduction.
Shade 02 · Scale
Shades of SWE by repo size
When I work on small repos, which model should I use?
Reasoning longer sometimes hurts performance in small repos. gpt-5.6-sol (high) hits the sweet spot of perf and cost here.
When I work on medium repos, which model should I use?
opus 5 leads in medium repos, but make sure to set reasoning to medium as well, otherwise your CC usage may drain away 1.9× faster without perf gains.
When I work on large repos, which model should I use?
kimi k3 appears on the pareto frontier for large repos. If you want the extra 3.7 points, go with opus 5 (xhigh), but be ready for a 1.8× increase in cost.
Shade 03 · Ecosystem
Shades of SWE by language
When I work on TypeScript repos, which model should I use?
Most frontier models perform similarly well in TypeScript. Not surprising since TS is the most popular language on GitHub. gpt-5.6-sol delivers the best perf. Personally, I’d go with luna since it’s within 9.4% of sol but 5.9× cheaper.
When I work on Go repos, which model should I use?
Performance is more spread out here. Go for opus 5 (xhigh) to get the best perf. luna is not a bad choice at 93.2% cheaper with a slight 8.4% lower perf.
When I work on Python repos, which model should I use?
gpt-5.6-sol and opus are both fine at Python or you can also code by hand once in a while to keep that Two Sum solution fresh in your mind.
Rust and JavaScript are omitted because each has only 5 tasks, which is too few for stable slice-level estimates.
Find Your Shade of SWE
- 01
Let a router choose
Try a model router like Cursor Router, Devin Fusion, or OpenRouter’s Auto Router, just to name a few. Picking models and reasoning efforts by hand is tricky. Let another model do it for you.
- 02
Benchmark your repo
Try converting your repo into a custom benchmark with services like Vals-Smith. Measure first, then optimize.
- 03
Benchmark your workload
Try making a benchmark for workloads your organization cares about. The react-doctor creators made ReactBench to measure agents on realistic React work. It takes more work, but pays off when the workload powers your organization.