embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities observed · 2026-08-28
Health v2 · maintenance only
95/100
- Activity 99
- Release rhythm 87
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0
- age_days: 1611
- days_rel: 8
- days_push: 7
- n_releases_24m: 522
Adoption not part of the score
3406 stars · 673 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
MTEB (Massive Text Embedding Benchmark) is a Python library and CLI for benchmarking embedding models across 1000+ tasks, languages, and modalities including text, image, and audio. It provides a standard evaluation harness, task selection, and a public leaderboard for comparing embedding and retrieval systems.
Use cases
- evaluate embedding model quality on standard benchmarks
- compare sentence embedding models on retrieval tasks
- benchmark text embeddings across multiple languages
- test multimodal image-text embedding models
- run embedding evaluation from the command line
- generate model card metadata from benchmark results
- find the best embedding model for classification or reranking
When to choose
- you need standardized, reproducible evaluation of embedding or retrieval models
- you want to compare your model against the MTEB leaderboard
- you need multilingual or multimodal (image/audio) embedding benchmarks
- you want a CLI or Python API for batch-running benchmark tasks
When to avoid
- you need to train or fine-tune embedding models rather than evaluate them
- you need a production embedding inference service
- you only need simple cosine similarity utilities without benchmarking
Facets
library · maturity active
benchmarking machine-learning nlp search-engine cli machine-learning data-science python cli cross-platform embeddings text-embedding semantic-search information-retrieval sentence-transformers evaluation leaderboard multilingual reranking benchmark-suite natural-language-processing search multimodal
5 sources
- readme: https://github.com/embeddings-benchmark/mteb · fetched 2026-08-28 · e37a62d2258f
- homepage: https://docs.mteb.org · fetched 2026-08-29 · 028c96e9e4dc
- site_page: https://docs.mteb.org/installation · fetched 2026-08-29 · f3893b66ce04
- site_page: https://docs.mteb.org/get_started/usage/get_started · fetched 2026-08-29 · 7f469a4ec9c8
- site_page: https://docs.mteb.org/get_started/usage/cli · fetched 2026-08-29 · 423d2fe904d4
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| embeddings-benchmark/mteb | main | 95 |
For agents
markdown · JSON · MCP: product_card(name="embeddings-benchmark/mteb")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem