← Benchmarks
Evaluates agents on ML tasks across computer vision, RL, tabular ML, and game theory.
Open benchmark problems affecting this task set.
Problems with a fix underway upstream.
Problems fixed upstream and recorded here.
Health numbers count affected task rows; benchmark-wide problems count once.