Ross ROSS = Recommend OSS · open-source software intelligence for agents

embeddings-benchmark/mteb

MTEB: State-of-the-art evaluation of embeddings across languages and modalities observed · 2026-08-28

github.com/embeddings-benchmark/mteb · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

95/100

  • Activity 99
  • Release rhythm 87
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 0
  • age_days: 1611
  • days_rel: 8
  • days_push: 7
  • n_releases_24m: 522

Full methodology

Adoption not part of the score

3406 stars · 673 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

MTEB (Massive Text Embedding Benchmark) is a Python library and CLI for benchmarking embedding models across 1000+ tasks, languages, and modalities including text, image, and audio. It provides a standard evaluation harness, task selection, and a public leaderboard for comparing embedding and retrieval systems.

Use cases

  • evaluate embedding model quality on standard benchmarks
  • compare sentence embedding models on retrieval tasks
  • benchmark text embeddings across multiple languages
  • test multimodal image-text embedding models
  • run embedding evaluation from the command line
  • generate model card metadata from benchmark results
  • find the best embedding model for classification or reranking

When to choose

  • you need standardized, reproducible evaluation of embedding or retrieval models
  • you want to compare your model against the MTEB leaderboard
  • you need multilingual or multimodal (image/audio) embedding benchmarks
  • you want a CLI or Python API for batch-running benchmark tasks

When to avoid

  • you need to train or fine-tune embedding models rather than evaluate them
  • you need a production embedding inference service
  • you only need simple cosine similarity utilities without benchmarking

Facets

library · maturity active

benchmarking machine-learning nlp search-engine cli machine-learning data-science python cli cross-platform embeddings text-embedding semantic-search information-retrieval sentence-transformers evaluation leaderboard multilingual reranking benchmark-suite natural-language-processing search multimodal

5 sources

Member repositories

RepositoryRoleHealth v2
embeddings-benchmark/mtebmain95

For agents

markdown · JSON · MCP: product_card(name="embeddings-benchmark/mteb")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem