openai/mle-bench resource
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering observed · 2026-08-28
Health v2 · maintenance only
58/100
- Activity 79
- Release rhythm 35
- Longevity 49
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 694
- days_rel: n/a
- days_push: 131
- n_releases_24m: 0
Adoption not part of the score
1720 stars · 257 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
MLE-bench is an open-source benchmark from OpenAI that measures how well AI agents perform machine learning engineering tasks, built from 75 Kaggle competitions with human baselines. It includes dataset construction code, evaluation/grading logic, and agent scaffolds for evaluating frontier LLMs.
Use cases
- evaluate how well an LLM agent does machine learning engineering
- benchmark AI agents on Kaggle-style ML competitions
- compare agent scaffolds like AIDE on ML tasks
- measure whether an agent can train models and prepare datasets autonomously
- research contamination and resource scaling for ML agents
- reproduce the MLE-bench leaderboard results for a new agent
When to choose
- you are researching or benchmarking LLM agents on real-world ML engineering tasks
- you want standardized Kaggle-derived tasks with human baselines and grading
- you need open-source evaluation harness code and reference agent scaffolds
When to avoid
- you need a general coding benchmark like SWE-bench rather than ML-specific tasks
- you want a production tool for running ML pipelines rather than an evaluation benchmark
- you require an actively accepting leaderboard submissions (currently paused)
Facets
dataset · maturity active
benchmarking agent-framework machine-learning testing machine-learning artificial-intelligence developer-tools tutorials python cross-platform kaggle evaluation llm-agents ml-engineering leaderboard openai ai-agents docker linux
5 sources
- readme: https://github.com/openai/mle-bench · fetched 2026-08-28 · f75375122f4c
- homepage: https://openai.com/index/mle-bench/ · fetched 2026-08-29 · 688b58243c42
- site_page: https://openai.com/about · fetched 2026-08-29 · de4b627b33e3
- site_page: https://developers.openai.com/api/docs · fetched 2026-08-29 · d617c5215f2a
- site_page: https://developers.openai.com/ · fetched 2026-08-29 · 7d240c906018
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| openai/mle-bench | main | 58 |
For agents
markdown · JSON · MCP: product_card(name="openai/mle-bench")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem