shangjingbo1226/AutoPhrase
AutoPhrase: Automated Phrase Mining from Massive Text Corpora observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 0
- Release rhythm 35
- Longevity 100
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 3562
- days_rel: n/a
- days_push: 1679
- n_releases_24m: 0
Adoption not part of the score
1202 stars · 271 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
AutoPhrase is a C++/Java tool for automated phrase mining that extracts quality phrases from massive text corpora with minimal human effort, leveraging existing knowledge bases for distant supervision. It supports multiple languages including English, Spanish, and Chinese with automatic language detection.
Use cases
- extract quality phrases from large text corpora
- mine compound words and multi-word expressions automatically
- build a domain phrase lexicon from raw documents
- segment text into phrases for downstream NLP tasks
- mine key phrases from Chinese or Spanish text
- preprocess corpora by identifying salient phrases
When to avoid
- you need a maintained library with active development and recent fixes
- you want a pure Python or pip-installable solution
- your corpus is small and simple keyword extraction would suffice
Facets
library · maturity maintenance
nlp parser data-science data-science cpp phrase-mining quality-phrases text-mining multi-language lexicon-extraction natural-language-processing linux macos docker
1 source
- readme: https://github.com/shangjingbo1226/AutoPhrase · fetched 2026-08-28 · 46dd2723238d
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| shangjingbo1226/AutoPhrase | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="shangjingbo1226/AutoPhrase")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem