function: nlp
1557 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Libr-AI/OpenFactVerification Loki is an open-source Python tool that automates fact verification by decomposing texts into claims, retrieving evidence via search, and u… | 15 | 1153 | active |
| PyThaiNLP/pythainlp PyThaiNLP is a Python library for Thai natural language processing, offering tokenization, POS tagging, transliteration, soundex, spell cor… | 98 | 1149 | active |
| brightmart/albert_zh A repository providing pre-trained ALBERT models for Chinese language, implemented in TensorFlow with PyTorch and Keras conversions. It inc… | 32 | 3982 | maintenance |
| ttttccxxui/DataInfra-RedactionEverything A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla… | 60 | 1147 | active |
| snipsco/snips-nlu Snips NLU is a Python library (with a Rust core) that extracts structured meaning from natural language text by detecting user intents and … | 23 | 3973 | maintenance |
| suifengqjn/videoWater AI快剪 (videoWater) is a desktop application for fully automated batch video editing, built in Go. It bundles clipping, merging, watermarking… | 32 | 1143 | active |
| jkwill87/mnamer mnamer is a Python command-line utility that parses media filenames for metadata, queries online providers like TVDb, TvMaze, TMDb, and OMD… | 93 | 1142 | active |
| AIScientists-Dev/academic-humanizer A skill/plugin for AI coding agents (Claude Code, Codex, MorphMind) that edits AI-assisted academic drafts to remove generic AI-writing tel… | 54 | 1142 | active |
| facebookresearch/StarSpace StarSpace is a general-purpose neural model from Facebook Research that learns entity embeddings for classification, retrieval, ranking, an… | 10 | 3952 | maintenance |
| AndyTheFactory/newspaper4k Newspaper4k is a Python library and CLI for scraping and curating news articles, extracting text, titles, authors, publish dates, and metad… | 84 | 1140 | active |
| vec2text/vec2text A Python library for text embedding inversion: training and running models that reconstruct text sequences from their sentence embeddings. … | 57 | 1136 | active |
| video-db/call.md Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr… | 59 | 1132 | active |
| Lingua Lingua is a language detection library for Python (with a Rust core) that identifies which of 75 languages a text is written in. It is desi… | 87 | 1129 | active |
| mrphrazer/reverser_ai ReverserAI is a Binary Ninja plugin that provides automated reverse engineering assistance using locally-hosted large language models runni… | 68 | 1127 | active |
| cookjohn/zotero-mcp A Zotero plugin with an integrated MCP server that lets AI assistants like Claude access and operate on a local Zotero library via the Mode… | 85 | 1120 | active |
| WojciechMula/pyahocorasick A Python module (implemented as a C extension with a pure-Python fallback) implementing the Aho-Corasick algorithm for fast multi-pattern s… | 76 | 1120 | stable |
| NTMC-Community/MatchZoo MatchZoo is a Python toolkit for designing, comparing, and sharing deep text matching models. It provides a unified data pipeline, pre-buil… | 23 | 3849 | maintenance |
| bevibing/tutor-skills A pair of Claude Code skills that convert documents (PDF, MD, HTML, EPUB) or source codebases into structured Obsidian study vaults with in… | 46 | 1114 | active |
| index-labs/readpilot Read Pilot is a web application that analyzes online articles and generates Q&A cards from them using OpenAI models. It is built with Next.… | 63 | 1112 | active |
| AdaptiveMotorControlLab/CEBRA CEBRA is a Python library for self-supervised learning of consistent latent embeddings from high-dimensional time-series recordings, using … | 86 | 1111 | active |
| textlint-ja/textlint-rule-preset-ai-writing A textlint rule preset that detects AI-generated writing patterns in Japanese text and suggests more natural phrasing. It can also run as a… | 74 | 1110 | active |
| fxy2311-youyou/expression-trainer A local desktop app (Electron) that trains spoken expression skills by combining fully offline real-time speech recognition (Sherpa-ONNX) w… | 55 | 1105 | active |
| dominostars/playtranslate PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for … | 81 | 1102 | active |
| yangheng95/PyABSA PyABSA is a PyTorch-based library providing state-of-the-art models for aspect-based sentiment analysis, including aspect term extraction, … | 67 | 1102 | active |
| wtetsu/mouse-dictionary Mouse Dictionary is a super fast browser extension that shows dictionary definitions instantly when you select or hover over words on web p… | 72 | 1096 | active |
| wilpel/caveman-compression A Python tool that compresses text for LLM contexts by stripping predictable grammar while preserving factual content, reducing token usage… | 41 | 1092 | active |
| greyblake/whatlang-rs Whatlang is a lightweight Rust library for detecting the natural language of a text, supporting 70 languages plus script recognition. It is… | 49 | 1087 | stable |
| QwenLM/Qwen2.5-Math Qwen2.5-Math is a series of math-specialized large language models (1.5B/7B/72B base and instruct variants plus a 72B reward model) built o… | 23 | 1087 | active |
| olivia-ai/olivia Olivia is an open-source chatbot written in Go that uses a neural network for natural language understanding, aiming to be a free alternati… | 10 | 3717 | maintenance |
| dtsola/xiaoyaosearch XiaoyaoSearch is a cross-platform desktop application (Electron + Python/FastAPI) that lets users find local files using AI-powered semanti… | 70 | 1083 | active |
| brianmario/charlock_holmes A Ruby library for character encoding detection built on top of ICU. It detects the encoding of arbitrary text or binary content and can tr… | 63 | 1083 | stable |
| jaraco/inflect A Python library that accurately generates English plurals, singular nouns, ordinals, indefinite articles, present participles, and word-ba… | 60 | 1083 | active |
| csurfer/rake-nltk rake-nltk is a Python library implementing the Rapid Automatic Keyword Extraction (RAKE) algorithm on top of NLTK. It determines key phrase… | 32 | 1083 | stable |
| zsyggg/paper-craft-skills A collection of Claude Code / Codex agent skills that turn academic papers into method figures, visual slide decks, and in-depth HTML artic… | 53 | 1080 | active |
| CJ-Chen/TBtools-II TBtools-II is a stand-alone bioinformatics platform with a user-friendly GUI and command-line tools for analyzing high-throughput sequencin… | 92 | 1074 | active |
| JimmyLefevre/kb A collection of single-header, permissively-licensed C/C++ libraries in the stb style. Its main library, kb_text_shape.h, provides Unicode … | 61 | 1074 | active |
| soxoj/socid-extractor socid_extractor is a Python library and CLI that extracts structured account metadata and stable internal identifiers (usernames, UIDs, GAI… | 92 | 1073 | active |
| FalkorDB/QueryWeaver QueryWeaver is an open-source Text2SQL tool that converts plain-English questions into SQL queries using graph-powered schema understanding… | 85 | 1073 | active |
| tianma8023/XposedSmsCode An Xposed module for Android that recognizes and parses SMS verification codes when a new message arrives, copying them to the clipboard an… | 23 | 1072 | active |
| ufal/whisper_streaming A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a … | 53 | 3672 | maintenance |
| mskayyali/nodepad nodepad is a spatial note-taking web application where notes are placed on a canvas and AI quietly classifies them, infers connections, and… | 56 | 1071 | active |
| rockbenben/subtitle-translator A free, browser-based batch subtitle translation tool for .srt, .ass, .vtt, and .lrc files that strips timing locally and sends only dialog… | 83 | 1070 | active |
| xmflswood/pinyin-match A JavaScript library for fast pinyin-based matching of Chinese text, supporting polyphonic characters, traditional Chinese, and initial-let… | 57 | 1070 | active |
| facebookresearch/LASER LASER is a Python library from Facebook Research for computing multilingual, language-agnostic sentence embeddings supporting 200+ language… | 10 | 3660 | maintenance |
| wzc570738205/smartParsePro smartParsePro is an offline JavaScript library for parsing Chinese addresses, extracting province, city, county, street, detail address, na… | 66 | 1068 | active |
| infrost/DeeplxFile DeeplxFile is a free, cross-platform desktop file translation tool built on Deeplx and Playwright that supports unlimited file sizes and ve… | 24 | 1068 | active |
| princeton-nlp/SimCSE SimCSE is a Python library and research codebase implementing simple contrastive learning for sentence embeddings, with pre-trained unsuper… | 23 | 3654 | maintenance |
| THUDM/GLM GLM is a general language model pretrained with an autoregressive blank-filling objective, released with pretrained checkpoints and fine-tu… | 32 | 3652 | maintenance |
| Makememo/MemoAI MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and… | 96 | 1059 | active |
| sylvainhalle/textidote TeXtidote is a Java command-line tool that performs spelling, grammar, and style checking on LaTeX documents (and Markdown files). It strip… | 80 | 1056 | active |
| THU-BPM/MarkLLM MarkLLM is an open-source Python toolkit for watermarking large language model outputs, implementing multiple LLM watermarking algorithms w… | 65 | 1054 | active |
| translate-tools/linguist Linguist is a privacy-first browser extension for Chrome and Firefox that translates web pages, selected text, subtitles, and messages, wit… | 98 | 1053 | active |
| InternRobotics/PointLLM PointLLM is a multimodal large language model that understands colored 3D point clouds of objects, built on a point cloud encoder fused wit… | 65 | 1051 | active |
| mikiarlo3/ai-copywriter A portable Markdown-based agent skill that writes marketing copy (headlines, descriptions, microcopy, subject lines) with a human tone whil… | 55 | 1048 | active |
| soulverteam/SoulverCore SoulverCore is a Swift framework providing a natural language math engine that evaluates expressions like '65 kg in pounds' or '$25k over 1… | 95 | 1047 | active |
| Kieirra/murmure Murmure is a privacy-first, open-source desktop speech-to-text application that transcribes voice entirely on-device using NVIDIA's Parakee… | 85 | 1047 | active |
| ChenYCL/chrome-extension-udemy-translate A Chrome extension that translates video subtitles on any website in real time, with support for custom DOM selectors per site. It supports… | 53 | 1046 | active |
| stay-leave/weibo-public-opinion-analysis A Python project for Weibo public opinion analysis that combines a web crawler, LDA topic modeling, sentiment analysis, and spatiotemporal … | 32 | 1045 | active |
| appsfolder/livebridge LiveBridge is a Flutter Android app with native Kotlin logic that converts regular notifications into Android 16+ Live Updates, providing a… | 74 | 1044 | active |
| Shawn1993/cnn-text-classification-pytorch A PyTorch implementation of Kim's CNN architecture for sentence classification, reproducing results from the paper 'Convolutional Neural Ne… | 65 | 1043 | active |
| thuiar/MMSA MMSA is a unified Python framework for multimodal sentiment analysis, supporting 15 MSA models and datasets like MOSI, MOSEI, and CH-SIMS. … | 23 | 1041 | active |
| gaboolic/rime-shuangpin-fuzhuma Moqi Yinxing is an open-source Rime input method configuration providing double-pinyin (shuangpin) schemes with shape-based auxiliary codes… | 85 | 1039 | active |
| D2I-CUHKSZ/MicroWorld MicroWorld is a lightweight Python engine that turns multi-modal event materials (documents, images, videos, graph signals) into structured… | 53 | 1039 | active |
| antimatter15/ocrad.js Ocrad.js is a pure-JavaScript port of the Ocrad OCR engine, compiled to JavaScript via Emscripten, that converts scanned images of text bac… | 32 | 3517 | maintenance |
| m-damien/VisualStoryWriting A web application that automatically visualizes a story's chronology, characters, and their movements, letting writers edit the story by ma… | 32 | 1031 | active |
| bilibili/Index-1.9B Index-1.9B is a family of lightweight 1.9-billion-parameter multilingual language models from Bilibili's Index team, released in base, chat… | 68 | 1030 | active |
| Felix3322/PotPlayer_ChatGPT_Translate A PotPlayer plugin that integrates OpenAI-compatible AI APIs (including ChatGPT and Ollama) to translate video subtitles in real time. It u… | 84 | 1029 | active |
| chatmcp/mcp-server-chatsum An MCP server that queries and summarizes chat messages stored in a local chat database, exposed as a tool for MCP-compatible clients like … | 21 | 1029 | active |
| jfilter/clean-text A Python package for cleaning and normalizing messy text, especially user-generated content from the web and social media. It fixes unicode… | 81 | 1027 | active |
| orange2ai/renwei-writing An open-source AI agent skill (a prompt/instruction pack) called 'Renwei Writing' that guides AI editors to revise text while preserving th… | 52 | 1017 | active |
| clab/dynet DyNet is a C++ neural network toolkit with Python bindings, designed for efficient CPU/GPU training of networks with dynamic per-instance s… | 23 | 3436 | maintenance |
| hezarai/hezar Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta… | 78 | 1013 | active |
| ownthink/Jiagu Jiagu is a Python library for Chinese natural language processing built on deep learning models trained on large-scale corpora. It provides… | 32 | 3425 | maintenance |
| wang-rui/phishguard-scaffold PhishGuard is a Python research framework that jointly performs phishing detection and dissemination control on social media using LLaMA-ba… | 47 | 1009 | active |
| voquill/voquill Voquill is an open-source, cross-platform AI voice dictation app that lets users dictate into any desktop application, with AI-powered tran… | 77 | 1008 | active |
| google-research/inksight InkSight is a Google Research system that converts photos of offline handwritten text into digital ink strokes using a ViT and mT5 encoder-… | 65 | 1006 | active |
| huawei-noah/noah-research A collection of research code subprojects released by Huawei Noah's Ark Lab, each in its own directory. It is not an official Huawei produc… | 76 | 1004 | active |
| letiantian/TextRank4ZH A Python implementation of the TextRank algorithm tailored for Chinese text, extracting keywords, key phrases, and extractive summaries usi… | 41 | 3393 | maintenance |
| google-research/albert Official TensorFlow implementation and pretrained checkpoints of ALBERT, a lite version of BERT for self-supervised learning of language re… | 10 | 3278 | maintenance |
| mojombo/chronic Chronic is a pure Ruby natural language date and time parser that converts phrases like 'tomorrow at 6:45pm' or '3rd wednesday in november'… | 32 | 3254 | maintenance |
| facebookresearch/MUSE MUSE is a Python library from Facebook AI Research for creating and aligning multilingual word embeddings, using both supervised (bilingual… | 10 | 3244 | maintenance |
| farizrahman4u/seq2seq A sequence-to-sequence learning add-on library for Keras, providing modular encoder-decoder layers and ready-made Seq2Seq models. It suppor… | 32 | 3169 | maintenance |
| cemoody/lda2vec A Python library implementing lda2vec, a hybrid topic model that combines word2vec word embeddings with LDA-style interpretable document to… | 32 | 3169 | maintenance |
| FudanNLP/fastNLP fastNLP is a lightweight, modularized and extensible NLP framework in Python that reduces engineering boilerplate such as data processing l… | 23 | 3141 | maintenance |
| twitter/twitter-text A collection of official Twitter libraries and conformance tests for parsing and tokenizing Tweet text. It determines character counts and … | 23 | 3139 | maintenance |
| matheuss/google-translate-api A free and unlimited Node.js library that provides an unofficial API for Google Translate, using the same servers as translate.google.com. … | 23 | 3135 | maintenance |
| dbiir/UER-py UER-py is a PyTorch framework for pre-training transformer language models (BERT, GPT-2, T5, ELMo, etc.) and fine-tuning them on downstream… | 32 | 3112 | maintenance |
| salesforce/CodeT5 Official research release of CodeT5 and CodeT5+ open code large language models from Salesforce Research for code understanding and generat… | 10 | 3093 | maintenance |
| pluja/whishper Whishper is a self-hosted, 100% local audio transcription and subtitling suite with a web UI, powered by FasterWhisper. It transcribes audi… | 61 | 3066 | maintenance |
| bigscience-workshop/promptsource PromptSource is a Python toolkit and repository for creating, sharing, and applying natural language prompts to Hugging Face datasets. It h… | 23 | 3031 | maintenance |
| yangjianxin1/GPT2-chitchat A GPT2-based Chinese chitchat dialogue model project built on HuggingFace transformers, including training, preprocessing, and interactive … | 32 | 2996 | maintenance |
| kotartemiy/newscatcher A Python package that programmatically collects normalized news articles from thousands of news websites, filterable by topic, country, and… | 32 | 2987 | maintenance |
| facebookresearch/XLM PyTorch implementation of Cross-lingual Language Model Pretraining (XLM) from Facebook AI Research, covering MLM, CLM, and TLM objectives p… | 10 | 2920 | maintenance |
| jbesomi/texthero Texthero is a Python toolkit for text preprocessing, representation, and visualization, designed to work on top of Pandas Series and DataFr… | 23 | 2907 | maintenance |
| EdgeTranslate/EdgeTranslate Edge Translate is a browser extension for translating selected text and entire web pages, supporting Chrome, Firefox, Edge, and QQ Browser.… | 23 | 2902 | maintenance |
| huggingface/neuralcoref NeuralCoref is a spaCy pipeline extension that annotates and resolves coreference clusters using a neural network, with a pre-trained Engli… | 23 | 2892 | maintenance |
| SCUTlihaoyu/open-chat-video-editor An open-source Python tool that automatically generates short videos from a short text prompt or a web URL, producing narration, background… | 29 | 2813 | maintenance |
| GerevAI/gerev Gerev is an AI-powered, self-hostable enterprise search engine that indexes workplace data sources like Slack, Confluence, Jira, and Google… | 30 | 2807 | maintenance |
| TeamHG-Memex/eli5 ELI5 is a Python library for debugging, inspecting, and explaining machine learning classifiers and regressors. It supports scikit-learn, X… | 66 | 2798 | maintenance |
| microsoft/CodeBERT A collection of pre-trained code models from Microsoft, including CodeBERT and successors like GraphCodeBERT and UniXcoder, usable via Hugg… | 32 | 2785 | maintenance |