# thunlp/THUOCL

THUOCL（THU Open Chinese Lexicon）中文词库

Repository: https://github.com/thunlp/THUOCL
Canonical: https://ross.abutalabs.com/products/thuocl
License: MIT
License Family: permissive
Topics: nlp, chinese
Last push: 2023-04-03T01:21:32+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2842, "days_push": 1249, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1116, forks 213 (observed 2026-08-28T04:03:38.827662+00:00)

## What it is
THUOCL (THU Open Chinese Lexicon) is a set of high-quality open Chinese word lists curated by Tsinghua University's NLP lab, covering domains like IT, finance, medicine, law, food, and place names. Each entry includes document frequency statistics, and the lexicons are designed to improve Chinese word segmentation, particularly when paired with the THULAC toolkit.

## Use cases
- improve chinese word segmentation accuracy
- get domain-specific chinese vocabulary lists
- find chinese word frequency statistics
- build a chinese nlp custom dictionary
- segment medical or legal chinese text better
- download chinese idiom and place name word lists

## When to choose
- you need curated, human-verified Chinese domain vocabularies with DF values
- you use THULAC or another segmenter and want domain-specific dictionaries
- you need Chinese lexicons for IT, finance, medicine, law, or other domains

## When to avoid
- you need a segmentation tool or runtime library rather than static word lists
- you need non-Chinese language lexicons
- you need frequently updated vocabularies, as updates are infrequent

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, parser
- domain: -
- platform: cross-platform
- tags: chinese-lexicon, word-segmentation, word-frequency, text-corpus, thulac, natural-language-processing, chinese-language

## Member repositories
- thunlp/THUOCL (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:38.827662+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:42:06.459704+00:00, confidence not recorded.
  - readme: https://github.com/thunlp/THUOCL (fetched 2026-08-28T04:03:38.827662+00:00, sha 3a8b1fcb89c0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
