Ross ROSS = Recommend OSS · open-source software intelligence for agents

huggingface/evaluate

🤗 Evaluate: A library for easily evaluating machine learning models and datasets. observed · 2026-08-28

github.com/huggingface/evaluate · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

74/100

  • Activity 91
  • Release rhythm 36
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 69
  • age_days: 1617
  • days_rel: 349
  • days_push: 58
  • n_releases_24m: 4

Full methodology

Adoption not part of the score

2477 stars · 336 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Hugging Face's library for easily evaluating machine learning models and datasets with dozens of standardized metrics, comparisons, and measurements. It supports NLP, computer vision, and audio tasks across frameworks like PyTorch, TensorFlow, and JAX, and lets users share custom evaluation modules on the Hub.

Use cases

  • compute accuracy or BLEU scores for my model predictions
  • evaluate a hugging face transformer model on a dataset
  • compare performance of two machine learning models
  • find standard metrics for machine translation evaluation
  • measure properties of a dataset before training
  • share a custom evaluation metric on the hugging face hub
  • evaluate text classification and question answering pipelines

When to choose

  • you need standardized, well-tested ML metrics across NLP, vision, and audio tasks
  • you work in the Hugging Face ecosystem and want Hub-integrated evaluation modules
  • you want framework-agnostic metrics usable with NumPy, PyTorch, TensorFlow, or JAX
  • you need an Evaluator to run model + dataset + metric pipelines end to end

When to avoid

  • you specifically need LLM evaluation - Hugging Face recommends the newer LightEval library instead
  • you need a general-purpose unit testing framework rather than ML metrics
  • you require cutting-edge actively developed features, since this library is in maintenance mode

Facets

library · maturity maintenance

machine-learning benchmarking testing data-science machine-learning computer-vision data-science python cross-platform evaluation-metrics hugging-face model-evaluation metrics nlp-metrics dataset-evaluation natural-language-processing

10 sources

Member repositories

RepositoryRoleHealth v2
huggingface/evaluatemain74

For agents

markdown · JSON · MCP: product_card(name="huggingface/evaluate")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem