SWE-bench repository
Official benchmark code, data, and evaluation documentation.
Compute College
Learn what SWE-bench measures, why it matters for AI coding agents, and how software-engineering benchmarks connect to AI compute demand.
SWE-bench is a software-engineering benchmark that evaluates whether AI systems can resolve real GitHub issues by producing changes to real repositories that satisfy evaluation tests.
Memory trick: SWE-bench is closer to “fix this repo issue” than “write this function.”
Repository-level repair is closer to deployed coding-agent work than short completions. If such workflows become reliable, developers can generate longer, repeated inference demand for debugging, patching, and validation.
A task can give an agent a code repository plus an issue description, then evaluate whether the submitted patch resolves the problem under its tests. Tool access and agent scaffold affect both score and cost.
Example figures are illustrative calculations, not current quoted market prices.
Current example
The official SWE-bench repository describes a benchmark for resolving real-world GitHub issues and its variants, including SWE-bench Verified. Its September 2026 update also made SWE-bench Multimodal v2 available as open source. Last checked: Sep 16, 2026.
Official benchmark code, data, and evaluation documentation.
No leaderboard performance claim is made here; consult the official benchmark configuration before comparing systems.
Do not compare SWE-bench values without checking the subset, scaffold, tools, test-time compute, and evaluation date.
Practical takeaway
Use SWE-bench as capability evidence, then measure your own repository tasks by cost per accepted patch and engineer review burden.
Decision check: do two SWE-bench numbers you are comparing share subset, scaffold, tools, and test-time compute?
Compute College
Follow model releases as AI compute market signals in the ComputeTape Market Brief.
Compute College track
Step 11 of 25: What is swe bench