function: nlp
1557 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| roshan-research/hazm Hazm is a Python library for natural language processing on Persian (Farsi) text, offering normalization, tokenization, stemming, lemmatiza… | 76 | 1417 | active |
| oracle/tribuo Tribuo is a Java machine learning library from Oracle Labs providing classification, regression, clustering, anomaly detection, and multi-l… | 59 | 1417 | active |
| jrzaurin/pytorch-widedeep A PyTorch library for multimodal deep learning that combines tabular data with text and images using Wide and Deep model architectures. It … | 62 | 1416 | active |
| honeyandme/RAGQnASystem A medical intelligent question-answering system combining knowledge-graph RAG with large language models, built on the DiseaseKG dataset wi… | 62 | 1415 | active |
| btford/write-good A naive linter for English prose that detects passive voice, weasel words, and other weak writing patterns. It works as both a Node.js libr… | 37 | 5086 | maintenance |
| argilla-io/argilla Argilla is an open-source collaboration tool for AI engineers and domain experts to build, annotate, and curate high-quality datasets for N… | 79 | 5085 | maintenance |
| zstar1003/ragflow-plus Ragflow-Plus is a fork of Ragflow, an open-source RAG (retrieval-augmented generation) platform, enhanced for practical Chinese-language us… | 64 | 1413 | active |
| explosion/spacy-transformers A spaCy v3 extension package that provides pipeline components for using pretrained transformer models like BERT, RoBERTa, XLNet, and GPT-2… | 71 | 1409 | stable |
| joewongjc/type4me Type4Me is a macOS voice input method that transcribes speech in real time (streaming, local or cloud ASR engines) and refines text with LL… | 81 | 1404 | active |
| alibaba/Logics-Parsing Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul… | 54 | 1402 | active |
| explosion/spacy-llm spacy-llm is a Python library that integrates Large Language Models into spaCy NLP pipelines via a serializable llm component. It provides … | 67 | 1394 | active |
| DualSubs/Universal DualSubs: Universal is a proxy-script-based solution that adds bilingual and translated subtitle options to HLS streaming platforms (e.g., … | 84 | 1392 | active |
| Ferry-200/coriander_player Coriander Player is a Windows local music player built with Flutter (Dart), Rust, and C (BASS library), featuring Material You theming. It … | 44 | 1392 | active |
| winkjs/wink-nlp WinkNLP is a developer-friendly JavaScript library for natural language processing with a lean, dependency-free codebase of ~10Kb minified … | 63 | 1388 | active |
| lamm-mit/PDF2Audio A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge… | 30 | 1383 | active |
| k1995/github-i18n-plugin A Tampermonkey userscript that translates the GitHub.com interface into Chinese (and Japanese), localizing menus, titles, buttons, and othe… | 71 | 1382 | active |
| textstat/textstat Textstat is a Python library that calculates statistical features from text, including readability scores like Flesch Reading Ease, Flesch-… | 71 | 1378 | stable |
| CIRCL/AIL-framework AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T… | 67 | 1378 | active |
| dorianbrown/rank_bm25 A Python library implementing several BM25 ranking algorithms (Okapi BM25, BM25L, BM25+) for scoring and ranking documents against a query.… | 59 | 1378 | stable |
| fukuball/jieba-php A PHP library for Chinese word segmentation (jieba), ported from the Python jieba project and now independently maintained. It supports pre… | 48 | 1376 | active |
| minimaxir/textgenrnn A Python 3 library built on Keras/TensorFlow for easily training char-rnn style neural networks that generate text from any dataset in a fe… | 23 | 4922 | maintenance |
| ssnangua/ColorTxt ColorTxt is a local desktop novel reader for TXT and common ebook formats (epub, mobi, pdf, etc.) that colorizes content with custom highli… | 80 | 1374 | active |
| nagisanzenin/engram Engram is a Claude Code plugin (portable to other agentic platforms) that turns your AI coding agent into a personal tutor for the human us… | 79 | 1374 | active |
| macanv/BERT-BiLSTM-CRF-NER A TensorFlow implementation of named entity recognition that fine-tunes Google BERT with a BiLSTM-CRF model, primarily targeting Chinese te… | 32 | 4906 | maintenance |
| emalderson/ThePhish ThePhish is an automated phishing email analysis web application built on TheHive, Cortex, and MISP. It extracts observables from email hea… | 32 | 1368 | stable |
| thunlp/OpenPrompt OpenPrompt is a PyTorch-based open-source framework for prompt-learning, providing a standard, flexible pipeline of templates and verbalize… | 23 | 4890 | maintenance |
| SUSYUSTC/MathTranslate MathTranslate is a Python tool that translates LaTeX documents, especially scientific papers from arXiv, between any languages while keepin… | 41 | 1363 | active |
| pemistahl/lingua-go Lingua is a Go library for highly accurate natural language detection, supporting 75 languages. It works well on short and mixed-language t… | 25 | 1359 | active |
| rockingdingo/deepnlp DeepNLP is a deep learning NLP pipeline implemented on TensorFlow, distributed as a Python package, which has evolved into the DeepNLP AI S… | 32 | 1358 | active |
| jenssegers/agent A PHP user agent parser that identifies browsers, operating systems, devices, robots, and accept languages, built on Mobile Detect with add… | 23 | 4850 | maintenance |
| huggingface/swift-transformers A Swift Package providing a transformers-like API for Swift apps, including fast tokenization, chat templating, and reliable model download… | 91 | 1350 | active |
| raznem/parsera Parsera is a lightweight Python library for scraping websites using LLMs, letting users define elements to extract with natural-language de… | 54 | 1350 | active |
| nyrahealth/CrisperWhisper CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers,… | 89 | 1349 | active |
| airbnb/aerosolve Aerosolve is a machine learning library from Airbnb built for human-friendly, interpretable modeling on the JVM. It provides a thrift-based… | 45 | 4808 | maintenance |
| natasha/natasha Natasha is a Python library that solves basic NLP tasks for the Russian language, including tokenization, sentence segmentation, morphology… | 67 | 1347 | active |
| JoySafety/JoySafety JoySafety is an open-source large language model safety framework from JD.com, written in Java, providing prompt injection detection, conte… | 47 | 1344 | active |
| a-real-ai/pywinassistant PyWinAssistant is an open-source agentic framework that acts as a Computer-Using-Agent, operating Windows 10/11 graphical user interfaces e… | 19 | 1343 | active |
| segment-any-text/wtpsplit wtpsplit is a Python toolkit for segmenting text into sentences or other semantic units using the SaT and WtP deep learning models. It prov… | 86 | 1333 | active |
| DeepL/deepl-python The official Python client library for the DeepL language translation API, providing a DeepLClient for translating text and documents via D… | 94 | 1330 | stable |
| agemagician/ProtTrans ProtTrans provides state-of-the-art pre-trained Transformer language models for protein sequences, trained on thousands of GPUs and hundred… | 33 | 1324 | active |
| sockysec/Telerecon Telerecon is a Python-based OSINT reconnaissance framework for researching and investigating Telegram. It scrapes user profiles, messages, … | 28 | 1324 | active |
| CBIhalsen/PolyglotPDF PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting bot… | 42 | 1315 | active |
| ndl-lab/ndlocr-lite NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma… | 77 | 1309 | active |
| albermax/innvestigate iNNvestigate is a Python toolbox providing a common interface and out-of-the-box implementations of many neural network explanation methods… | 30 | 1309 | active |
| ZachSaucier/Just-Read Just Read is a customizable reader mode browser extension that reformats article pages into a clean, distraction-free reading view with cus… | 77 | 1308 | active |
| ttop32/MouseTooltipTranslator A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports … | 95 | 1300 | active |
| stenolabs/stenoai Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire… | 84 | 1299 | active |
| okfn-brasil/serenata-de-amor Operação Serenata de Amor is an open-source data science project that uses machine learning to audit Brazilian congresspeople's public expe… | 32 | 4603 | maintenance |
| JuliaStrings/utf8proc utf8proc is a small, clean C library providing Unicode normalization, case-folding, grapheme segmentation, and other operations for UTF-8 e… | 80 | 1293 | stable |
| xinychen/transdim transdim is a Python/Jupyter Notebook project providing machine learning models for transportation data imputation and spatiotemporal time … | 63 | 1289 | active |
| onvoyage-ai/gtm-engineer-skills A collection of Claude Code skills that automate go-to-market and AI-search optimization workflows, from brand research and keyword cluster… | 54 | 1288 | active |
| ictnlp/StreamSpeech StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p… | 37 | 1287 | active |
| xunbu/docutranslate DocuTranslate is a lightweight local document translation tool powered by large language models, supporting formats such as pdf, docx, xlsx… | 80 | 1285 | active |
| inukshuk/anystyle AnyStyle is a fast machine-learning-based parser that splits bibliographic references into structured segments like author, title, and publ… | 42 | 1285 | active |
| greymd/ojichat ojichat is a Go CLI tool that generates humorous text mimicking the style of messages middle-aged Japanese men ('ojisan') send on LINE or e… | 23 | 1275 | active |
| microsoft/BioGPT BioGPT is Microsoft's domain-specific generative Transformer language model pre-trained on biomedical text, with implementation code and pr… | 32 | 4488 | maintenance |
| MilaNLProc/contextualized-topic-models A Python library implementing Contextualized Topic Models (CTM), which combine pre-trained contextual embeddings like BERT with neural topi… | 47 | 1269 | active |
| amaiya/ktrain ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul… | 25 | 1268 | active |
| huichen/wukong Wukong is a highly customizable full-text search engine library written in Go, with efficient indexing, Chinese word segmentation via the s… | 23 | 4474 | maintenance |
| facebookresearch/DrQA DrQA is a PyTorch implementation of a system for open-domain question answering that combines document retrieval over Wikipedia with a neur… | 10 | 4468 | maintenance |
| thunlp/OpenNRE OpenNRE is an open-source Python toolkit for neural relation extraction, extracting relation triples between entities from plain text. It u… | 32 | 4467 | maintenance |
| mvdan/xurls A Go library and CLI tool that extracts URLs from arbitrary text using regular expressions built from TLD lists. It offers Relaxed and Stri… | 67 | 1265 | active |
| cclank/news-aggregator-skill A Python-based agent skill that aggregates news from 44+ sources (tech, finance, AI, international) and generates AI-summarized daily brief… | 53 | 1263 | active |
| MemeMeow-Studio/MemeMeow MemeMeow is a self-hosted meme/sticker management and retrieval application that lets users find images by describing the desired scene in … | 65 | 1260 | active |
| stair-lab/kg-gen kg-gen is a Python library that extracts knowledge graphs from arbitrary plain text or conversation messages using LLMs, with model routing… | 63 | 1260 | active |
| Fictionarry/ER-NeRF ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr… | 24 | 1260 | stable |
| mjpost/sacrebleu SacreBLEU is a Python library and CLI tool for computing shareable, comparable, and reproducible BLEU, chrF, and TER scores for machine tra… | 76 | 1258 | active |
| tatuylonen/wiktextract A Python package and CLI tool that parses Wiktionary XML dump files and extracts structured dictionary data (glosses, translations, pronunc… | 76 | 1251 | active |
| arunsupe/semantic-grep w2vgrep is a grep-like command-line tool that finds words semantically similar to a query using word2vec embeddings, with familiar grep opt… | 14 | 1246 | active |
| MrGeDiao/shuorenhua A Chinese-first rewrite skill that removes AI-generated tone (template phrasing, performative language, translationese) from text while pre… | 81 | 1245 | active |
| unum-cloud/UForm UForm is a compact multimodal AI library providing tiny image-text embedding models (64-768 dimensions, Matryoshka-style) and small generat… | 55 | 1244 | active |
| fpgaminer/joycaption JoyCaption is an open, free, and uncensored image captioning Visual Language Model (VLM) with released weights and training scripts. It gen… | 53 | 1244 | active |
| apache/lucene-solr The former shared repository for Apache Lucene and Apache Solr, open-source search software. Lucene and Solr have split into separate top-l… | 69 | 4363 | maintenance |
| thetahealth/mirobody Mirobody is an open-source, AI-native health data engine that collects readings from lab reports, wearables, and genomics, standardizes the… | 84 | 1229 | active |
| Heavrnl/TelegramForwarder A self-hosted Telegram message forwarder built on Telethon that copies messages from multiple source chats to target chats with keyword/reg… | 52 | 1229 | active |
| iDC-NEU/YiGraph YiGraph is an LLM-driven agent system for autonomous graph data analytics built on the Analytics-Augmented Generation (AAG) framework. It e… | 61 | 1227 | active |
| OStudi/short-video-generator-AI An open-source Python tool that turns YouTube videos into ready-to-post vertical short videos by automatically detecting highlights, adding… | 57 | 1224 | active |
| nfstream/nfstream NFStream is a multiplatform Python framework for fast, flexible network flow data analysis from live interfaces or pcap files. It provides … | 81 | 1218 | stable |
| lishix520/academic-paper-skills A set of Claude Code skills that provide a structured pipeline for planning and writing academic papers, with a strategist skill for venue … | 43 | 1217 | active |
| persian-tools/persian-tools A comprehensive, zero-dependency TypeScript toolkit with 27+ utilities for Persian (Farsi) text, numbers, and Iranian-specific validation s… | 72 | 1210 | active |
| taranis-ai/taranis-ai Taranis AI is a self-hosted open-source OSINT platform that collects news articles from web sources and uses NLP/AI to enrich, cluster, and… | 95 | 1207 | active |
| opensemanticsearch/open-semantic-search An open-source integrated search server and ETL framework for processing, analyzing, and exploring large document collections. It combines … | 40 | 1203 | active |
| juliasilge/tidytext tidytext is an R package that applies tidy data principles to text mining, providing functions like unnest_tokens to convert text to and fr… | 66 | 1201 | stable |
| deepseek-ai/DeepSeek-VL DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf… | 25 | 4175 | maintenance |
| jncraton/languagemodels A Python library providing simple building blocks for running large language models locally with as little as 512MB of RAM. It offers instr… | 60 | 1192 | active |
| deusyu/translate-book An agent skill for Codex, Claude Code, and OpenClaw that translates entire books in PDF, DOCX, or EPUB format into any language. It convert… | 58 | 1192 | active |
| huggingface/Math-Verify A Python library from Hugging Face for robustly parsing and verifying mathematical expressions, designed to evaluate Large Language Model o… | 50 | 1186 | active |
| chakki-works/seqeval seqeval is a Python library for evaluating sequence labeling tasks such as named-entity recognition, part-of-speech tagging, and semantic r… | 23 | 1184 | stable |
| goodroot/hyprwhspr hyprwhspr is a native Linux system-wide speech-to-text dictation application supporting local models (Whisper, Parakeet, Cohere) with optio… | 85 | 1182 | active |
| mlfoundations/open_flamingo OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima… | 23 | 4118 | maintenance |
| huangkiki/dailypaper-skills A set of Claude Code skills that automates a daily research paper pipeline: fetching new papers from HuggingFace Daily, Trending, and arXiv… | 59 | 1181 | active |
| alchaincyf/x-mentor-skill An agent skill (SKILL.md package) that distills the methodologies of six top X (Twitter) creators plus open-source X algorithm weight data … | 59 | 1180 | active |
| grangier/python-goose Python-Goose is a Python library that extracts the main body text, metadata, top image, and embedded videos from news article web pages. It… | 64 | 4106 | maintenance |
| yanyiwu/simhash A C++ header-only library that computes Simhash fingerprints for Chinese documents, using CppJieba for tokenization and keyword extraction.… | 65 | 1170 | stable |
| matthiasn/lotti Lotti is a private, local-first logbook app for journaling, task management, time tracking, habits, and health data, with a staff of person… | 98 | 1169 | active |
| SuperBruceJia/EEG-DL EEG-DL is a deep learning library built on TensorFlow for classifying EEG signals, supporting many architectures including CNNs, RNNs, GCNs… | 47 | 1167 | active |
| DavidVentura/offline-translator An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off… | 88 | 1166 | active |
| thunlp/OpenKE OpenKE is an open-source PyTorch-based toolkit for knowledge graph embedding (knowledge representation learning), with C++ accelerated data… | 32 | 4047 | maintenance |
| liujuntao123/smart-mermaid Smart Mermaid is an AI-powered web application that converts natural language text into Mermaid diagram code and renders it as visual chart… | 41 | 1154 | active |
| baidu/lac LAC (Lexical Analysis of Chinese) is Baidu's deep-learning-based Chinese lexical analysis toolkit that jointly performs word segmentation, … | 23 | 4001 | maintenance |