Ross ROSS = Recommend OSS · open-source software intelligence for agents

Cloud-CV/EvalAI

:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI observed · 2026-08-28

github.com/Cloud-CV/EvalAI · homepage · Python · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

67/100

  • Activity 99
  • Release rhythm 8
  • Longevity 100

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 3604
  • days_rel: n/a
  • days_push: 10
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2039 stars · 983 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

EvalAI is an open-source platform for evaluating and comparing machine learning and AI algorithms at scale. It provides a central leaderboard and submission interface with remote, map-reduce-backed evaluation to make benchmark results reproducible.

Use cases

  • host an AI challenge with a public leaderboard
  • evaluate ML model submissions against a test set
  • reproduce benchmark results from research papers
  • run remote evaluation of large-scale ML challenges
  • compare algorithms on standardized dataset splits
  • manage private and public leaderboards for a competition

When to choose

  • you need to host an ML/AI competition with submissions and leaderboards
  • you want reproducible, standardized evaluation of algorithms at scale
  • you need custom evaluation phases, splits, and metrics
  • you want a self-hosted alternative to closed challenge platforms

When to avoid

  • you only need simple unit testing of ML code rather than challenge evaluation
  • you want a lightweight single-model benchmark harness without a web platform
  • you cannot operate a Django/Angular web service with Docker infrastructure

Facets

service · maturity active

machine-learning benchmarking api-framework web-framework self-hosted machine-learning artificial-intelligence data-science analytics web-development python self-hosted cross-platform ai-challenges leaderboard evaluation-platform reproducible-research django angularjs ml-evaluation competition-hosting docker web-server

2 sources

Member repositories

RepositoryRoleHealth v2
Cloud-CV/EvalAImain67

For agents

markdown · JSON · MCP: product_card(name="Cloud-CV/EvalAI")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem