# neologd/mecab-ipadic-neologd

Neologism dictionary based on the language resources on the Web for mecab-ipadic

Repository: https://github.com/neologd/mecab-ipadic-neologd
Canonical: https://ross.abutalabs.com/products/mecab-ipadic-neologd
Language: Shell
License: NOASSERTION
License Family: other
Topics: mecab-ipadic, named-entities, dictionary, furigana, neologism-dictionary, mecab, language-resources, japanese-language
Last push: 2023-12-27T06:46:17+00:00

## Health v2 (maintenance only)
Score: 23/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 8, longevity 100
- inputs: {"age_days": 4195, "days_push": 980, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2789, forks 287 (observed 2026-08-28T04:07:22.020458+00:00)

## What it is
A customized system dictionary for the MeCab Japanese morphological analyzer, containing millions of neologisms and named entities extracted from web language resources. It supplements the default ipadic dictionary to correctly tokenize new words, proper nouns, and expressions found in web documents.

## Use cases
- tokenize japanese social media posts with new slang words
- extract named entities from japanese web text
- improve mecab segmentation of recent product names and person names
- analyze japanese documents containing emojis and kaomoji
- get furigana readings for newly coined japanese words

## When to choose
- your mecab-based pipeline fails to correctly tokenize modern japanese words and named entities
- you process web-crawled japanese text with slang, brand names, or person names
- you need furigana readings for words missing from the default ipadic dictionary

## When to avoid
- you need strictly accurate named entity classification, since categories are coarse and noisy
- you cannot use UTF-8 versions of MeCab
- you lack the 2GB+ RAM required to build the dictionary
- you need a regularly updated dictionary, as releases are infrequent

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, parser
- domain: -
- platform: cross-platform, cli
- tags: mecab, japanese-tokenization, neologism-dictionary, furigana, named-entities, ipadic, natural-language-processing, japanese-language

## Member repositories
- neologd/mecab-ipadic-neologd (main) score 23

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:22.020458+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T08:15:30.365557+00:00, confidence not recorded.
  - readme: https://github.com/neologd/mecab-ipadic-neologd (fetched 2026-08-28T04:07:22.020458+00:00, sha ff767da3bf82)
- Data as of 2026-08-30T08:39:29.467469+00:00.
