# Separius/awesome-sentence-embedding

A curated list of pretrained sentence and word embedding models

Repository: https://github.com/Separius/awesome-sentence-embedding
Canonical: https://ross.abutalabs.com/products/awesome-sentence-embedding
Language: Python
License: GPL-3.0
License Family: copyleft
Topics: wordembedding, word-embeddings, sentence-embeddings, nlp, pretrained-embedding, awesome-list, pretrained-models, unsupervised-learning, sentence-representations, contextualized-representation, awesome, natural-language, subword-models, embedding-models, cross-lingual, bert, pretrained-language-model, language-model
Archived: true
Last push: 2021-04-23T09:28:48+00:00

## Health v2 (maintenance only)
Score: 10/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2823, "days_push": 1958, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, archived
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2289, forks 263 (observed 2026-08-28T04:06:34.791735+00:00)

## What it is
A curated awesome-list of pretrained sentence and word embedding models, organized into tables covering word embeddings, contextualized embeddings, encoders, pooling methods, and evaluation. It links to papers, training code, and pretrained model downloads rather than providing software itself.

## Use cases
- find pretrained sentence embedding models
- compare word embedding papers and implementations
- discover contextualized embedding models like BERT variants
- look up cross-lingual embedding resources
- research sentence representation methods
- find pretrained word2vec and subword models

## When to choose
- you need a survey of embedding models with links to pretrained weights
- you are researching sentence representation techniques
- you want a structured comparison of embedding papers by date and citations

## When to avoid
- you need a runnable library or tool rather than a reference list
- you need up-to-date coverage of recent embedding models, as the list was last updated in 2021
- you want production-ready embedding inference code

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: nlp, machine-learning
- domain: machine-learning, awesome-lists
- platform: cross-platform
- tags: awesome-list, sentence-embeddings, word-embeddings, pretrained-models, curated-list, bert, cross-lingual, natural-language-processing

## Member repositories
- Separius/awesome-sentence-embedding (main) score 10

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:34.791735+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:40:43.309493+00:00, confidence not recorded.
  - readme: https://github.com/Separius/awesome-sentence-embedding (fetched 2026-08-28T04:06:34.791735+00:00, sha 6880bd60d4c3)
- Data as of 2026-08-30T08:39:29.467469+00:00.
