Ross ROSS = Recommend OSS · open-source software intelligence for agents

neologd/mecab-ipadic-neologd resource

Neologism dictionary based on the language resources on the Web for mecab-ipadic observed · 2026-08-28

github.com/neologd/mecab-ipadic-neologd · Shell · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

23/100

  • Activity 0
  • Release rhythm 8
  • Longevity 100

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 4195
  • days_rel: n/a
  • days_push: 980
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2789 stars · 287 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A customized system dictionary for the MeCab Japanese morphological analyzer, containing millions of neologisms and named entities extracted from web language resources. It supplements the default ipadic dictionary to correctly tokenize new words, proper nouns, and expressions found in web documents.

Use cases

  • tokenize japanese social media posts with new slang words
  • extract named entities from japanese web text
  • improve mecab segmentation of recent product names and person names
  • analyze japanese documents containing emojis and kaomoji
  • get furigana readings for newly coined japanese words

When to choose

  • your mecab-based pipeline fails to correctly tokenize modern japanese words and named entities
  • you process web-crawled japanese text with slang, brand names, or person names
  • you need furigana readings for words missing from the default ipadic dictionary

When to avoid

  • you need strictly accurate named entity classification, since categories are coarse and noisy
  • you cannot use UTF-8 versions of MeCab
  • you lack the 2GB+ RAM required to build the dictionary
  • you need a regularly updated dictionary, as releases are infrequent

Facets

dataset · maturity maintenance

nlp parser cross-platform cli mecab japanese-tokenization neologism-dictionary furigana named-entities ipadic natural-language-processing japanese-language

1 source

Member repositories

RepositoryRoleHealth v2
neologd/mecab-ipadic-neologdmain23

For agents

markdown · JSON · MCP: product_card(name="neologd/mecab-ipadic-neologd")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem