# Embedding/Chinese-Word-Vectors

100+ Chinese Word Vectors 上百种预训练中文词向量

Repository: https://github.com/Embedding/Chinese-Word-Vectors
Canonical: https://ross.abutalabs.com/products/chinese-word-vectors
Language: Python
License: Apache-2.0
License Family: permissive
Topics: chinese, chinese-word-segmentation, embeddings, word-embeddings, vectors-trained, embedding
Last push: 2023-10-30T14:44:50+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3158, "days_push": 1038, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 12225, forks 2323 (observed 2026-08-28T04:10:52.078153+00:00)

## What it is
A collection of 100+ pre-trained Chinese word vector embeddings trained with different representations, context features, and corpora, plus the CA8 analogical reasoning dataset and an evaluation toolkit. Vectors are provided in text format for easy use in downstream NLP tasks.

## Use cases
- download pretrained chinese word embeddings
- find word vectors for chinese text classification
- evaluate quality of chinese word embeddings
- get embeddings for chinese nlp downstream tasks
- chinese word analogy reasoning dataset
- compare dense and sparse chinese word vectors

## When to choose
- you need ready-made Chinese word embeddings for NLP tasks
- you want to benchmark or evaluate Chinese word vectors
- you need a Chinese analogical reasoning evaluation dataset

## When to avoid
- you need embeddings for languages other than Chinese
- you want contextual embeddings like BERT rather than static word vectors
- you need a maintained library rather than pre-trained data files

## Facets
- artifact type: dataset
- maturity: stable
- function: nlp, machine-learning, data-science
- domain: machine-learning, localization
- platform: python, cross-platform
- tags: word-embeddings, chinese-nlp, pretrained-vectors, word2vec, evaluation-toolkit, analogical-reasoning, natural-language-processing

## Member repositories
- Embedding/Chinese-Word-Vectors (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:52.078153+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:14:15.548175+00:00, confidence not recorded.
  - readme: https://github.com/Embedding/Chinese-Word-Vectors (fetched 2026-08-28T04:10:52.078153+00:00, sha 755f32463ad0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
