Global Pareto
113 tasks · 91 repositories · 51 configurations
The reference
DeepSWE summarizes performance across 113 software-engineering tasks. But the global view assumes every kind of repository matters equally.
113 tasks · 91 repositories · 51 configurations
91 repositories. Dot size is task count; color is primary language.
Shade 01 · Popularity
Repository popularity separates DeepSWE into three equally readable views of the cost–performance frontier.
43 tasks · 34 repositories
34 tasks · 30 repositories
36 tasks · 27 repositories
These slices are descriptive. Stars may also reflect repository age, language, domain, or project type.
Shade 02 · Scale
As codebases grow, agents must navigate and coordinate changes across more files. The frontier shifts with that scale.
35 tasks · 31 repositories
39 tasks · 28 repositories
39 tasks · 32 repositories
Shade 03 · Ecosystem
Language and ecosystem change the mix of work in the benchmark—and the models that offer the best tradeoff.
35 tasks · 28 repositories
34 tasks · 29 repositories
34 tasks · 28 repositories
Rust and JavaScript are omitted because each has only 5 tasks, which is too few for stable slice-level estimates.
Your repository
Paste a public GitHub repository. We’ll inspect its language, size, and stars, then use similar DeepSWE repositories to draft a personalized Pareto frontier.
The result is a draft, not a measured benchmark run. We’ll show the repositories contributing to the estimate and how much support it has.