# TencentGameMate/chinese_speech_pretrain

chinese speech pretrained models

Repository: https://github.com/TencentGameMate/chinese_speech_pretrain
Canonical: https://ross.abutalabs.com/products/chinese_speech_pretrain
Language: Shell
License Family: other
Last push: 2024-08-23T03:14:03+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 1561, "days_push": 740, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1215, forks 89 (observed 2026-08-28T04:04:01.103801+00:00)

## What it is
A collection of Chinese speech pretrained models (wav2vec 2.0 and HuBERT, BASE and LARGE) trained by Tencent on 10,000 hours of WenetSpeech data using Fairseq. Checkpoints are hosted on Hugging Face and Baidu Pan, with ESPnet-based ASR recipes demonstrating their use as feature extractors for Conformer speech recognition.

## Use cases
- pretrain wav2vec2 or hubert models on chinese speech
- use chinese speech representations as features for asr
- improve mandarin speech recognition with low-resource fine-tuning
- download chinese wav2vec2 checkpoints for huggingface transformers
- benchmark chinese speech recognition on aishell and wenetspeech
- extract self-supervised speech embeddings for downstream audio tasks

## When to choose
- you need chinese-language speech pretrained models for asr or audio feature extraction
- you want wav2vec2/hubert checkpoints trained on large-scale mandarin data
- you are fine-tuning speech recognition on limited labeled chinese audio

## When to avoid
- you need pretrained models for non-chinese languages
- you want a ready-to-use end-to-end speech recognition product rather than pretrained features
- you cannot work with fairseq or espnet tooling

## Facets
- artifact type: dataset
- maturity: maintenance
- function: speech-recognition, machine-learning, transformers, audio-processing
- domain: speech-processing, machine-learning, artificial-intelligence
- platform: python
- tags: wav2vec2, hubert, self-supervised-pretraining, fairseq, chinese-asr, wenetspeech, pretrained-models, huggingface, natural-language-processing, gpu, linux

## Member repositories
- TencentGameMate/chinese_speech_pretrain (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:01.103801+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:17:14.107152+00:00, confidence not recorded.
  - readme: https://github.com/TencentGameMate/chinese_speech_pretrain (fetched 2026-08-28T04:04:01.103801+00:00, sha 18750fd41ff3)
- Data as of 2026-08-30T08:39:29.467469+00:00.
