← Benchmarks
SWE-bench Pro: A multi-language software engineering benchmark with 731 instances covering Python, JavaScript/TypeScript, and Go. Evaluates AI systems' ability to resolve real-world bugs and implement features across diverse production codebases.
Open benchmark problems affecting this task set.
Problems with a fix underway upstream.
Problems fixed upstream and recorded here.
Health numbers count affected task rows; benchmark-wide problems count once.