function: nlp
1557 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| OneMoreGres/ScreenTranslator Screen Translator is a desktop utility that captures a selected region of the screen, performs OCR on it, and sends the recognized text to … | 68 | 1279 | maintenance |
| kaiyinzhou/BERT-NER A Python implementation that fine-tunes Google's BERT for named entity recognition using the CoNLL-2003 dataset, built on TensorFlow. It pr… | 32 | 1277 | maintenance |
| dragnet-org/dragnet Dragnet is a Python library that uses machine learning models to extract the main article content, and optionally user comments, from HTML … | 36 | 1274 | maintenance |
| google-research/deduplicate-text-datasets A Rust implementation of ExactSubstr deduplication for language model training datasets, with Python scripts for running deduplication and … | 10 | 1270 | maintenance |
| sdkcarlos/artyom.js Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech… | 23 | 1269 | maintenance |
| summanlp/textrank A Python 3 implementation of the TextRank algorithm for extractive text summarization and keyword extraction. It is the library published o… | 23 | 1269 | maintenance |
| HendrikStrobelt/LSTMVis LSTMVis is a visual analysis toolbox for inspecting hidden state dynamics of LSTM (and other RNN) networks. It provides a browser-based int… | 23 | 1266 | maintenance |
| machinalis/quepy Quepy is a Python framework that transforms natural language questions into database queries, currently supporting SPARQL and MQL. It uses … | 32 | 1264 | maintenance |
| linkedin/detext DeText is a deep neural text understanding framework from LinkedIn for NLP ranking, classification, and language generation tasks. It uses … | 23 | 1263 | maintenance |
| UKPLab/EasyNMT EasyNMT is a Python library providing easy access to state-of-the-art neural machine translation for 100+ languages using models like Opus-… | 23 | 1260 | maintenance |
| andrewyng/translation-agent A Python demonstration library by Andrew Ng implementing an agentic machine translation workflow: an LLM translates text, reflects on its o… | 24 | 5804 | experimental |
| pasky/speedread A terminal-based speed-reading tool that displays text as per-word RSVP (rapid serial visual presentation) aligned on optimal reading point… | 23 | 1257 | maintenance |
| maciejkula/glove-python A toy Python/Cython implementation of the GloVe algorithm for training dense word vector embeddings by factorizing the log of the word co-o… | 32 | 1253 | maintenance |
| taskrabbit/react-native-parsed-text A React Native library that parses text strings and renders matching segments (URLs, phone numbers, emails, or custom regex patterns) as se… | 23 | 1250 | maintenance |
| kamalkraj/BERT-NER A PyTorch library for training and running named entity recognition (NER) models based on Google's BERT, evaluated on the CoNLL-2003 datase… | 32 | 1249 | maintenance |
| XiangLi1999/Diffusion-LM Diffusion-LM is the official research code for the paper 'Diffusion-LM Improves Controllable Text Generation', implementing a diffusion-bas… | 32 | 1245 | maintenance |
| jakesnell/prototypical-networks Reference PyTorch implementation of Prototypical Networks for few-shot classification from the NeurIPS 2017 paper. It includes training and… | 32 | 1235 | maintenance |
| yuanxiaosc/Entity-Relation-Extraction A TensorFlow and BERT based pipeline for joint entity and relation extraction, built as a solution for the 2019 Language and Intelligence C… | 32 | 1230 | maintenance |
| jiangqizheng/BlueSea BlueSea is a browser extension for English learning that offers word selection translation, word highlighting, word danmaku (bullet comment… | 32 | 1228 | maintenance |
| liuhuanyong/ComplexEventExtraction A Python library for Chinese compound event extraction that identifies conditional, causal, sequential, and adversative events using explic… | 32 | 1227 | maintenance |
| cameron/squirt Squirt is a speed-reading bookmarklet that displays web article text one word at a time in a Spritz-style RSVP reader. It automatically ext… | 32 | 1226 | maintenance |
| bheinzerling/bpemb BPEmb is a collection of pre-trained subword embeddings in 275 languages based on Byte-Pair Encoding, trained on Wikipedia, distributed as … | 32 | 1224 | maintenance |
| tylin/coco-caption The official evaluation code for the Microsoft COCO image captioning benchmark, implementing metrics such as BLEU, METEOR, ROUGE-L, CIDEr, … | 32 | 1224 | maintenance |
| iqiyi/FASPell FASPell is a Python-based Chinese spell checker based on the DAE-Decoder paradigm, published at the EMNLP 2019 W-NUT workshop. It detects a… | 32 | 1223 | maintenance |
| facebookresearch/BLINK BLINK is a Python entity linking library from Facebook Research that resolves mentions in text to Wikipedia entities using a two-stage bi-e… | 10 | 1210 | maintenance |
| leizongmin/node-segment A pure JavaScript Chinese word segmentation (tokenization) library for Node.js, based on the Pangu segmentation dictionary and algorithms. … | 32 | 1206 | maintenance |
| shangjingbo1226/AutoPhrase AutoPhrase is a C++/Java tool for automated phrase mining that extracts quality phrases from massive text corpora with minimal human effort… | 32 | 1202 | maintenance |
| google-research/tapas TAPAS is Google Research's implementation of transformer-based models for question answering over tables, including pretrained checkpoints … | 10 | 1202 | maintenance |
| epfml/sent2vec Sent2vec is a C++ library with a Cython/Python interface that trains unsupervised distributed representations of sentences and short texts,… | 23 | 1201 | maintenance |
| huggingface/hmtl HMTL is a Hierarchical Multi-Task Learning model for NLP that jointly trains Named Entity Recognition, Entity Mention Detection, Relation E… | 32 | 1195 | maintenance |
| kexinhuang12345/DeepPurpose DeepPurpose is a PyTorch-based deep learning library for molecular modeling, supporting drug-target interaction (DTI), drug property, drug-… | 23 | 1184 | maintenance |
| google/budou Budou is a Python library and CLI tool that automatically organizes CJK (Chinese, Japanese, Korean) text into semantically meaningful chunk… | 10 | 1183 | maintenance |
| ChestnutHeng/Wudao-dict Wudao-dict is a command-line English-Chinese dictionary based on Youdao, offering offline and online lookups with definitions, phrases, and… | 32 | 1182 | maintenance |
| pymorphy2/pymorphy2 pymorphy2 is a Python library providing morphological analysis (POS tagging and inflection) for Russian and Ukrainian languages. It include… | 32 | 1175 | maintenance |
| YaoFANGUK/video-subtitle-generator A Python application with both GUI and CLI interfaces that generates subtitle files (SRT) from video or audio using local Whisper-based spe… | 32 | 1175 | maintenance |
| IndieKKY/bilibili-subtitle A browser extension (哔哔君) that displays a clickable subtitle list panel for Bilibili videos, enabling jump-to-timestamp navigation, subtitl… | 52 | 1174 | maintenance |
| soulteary/docker-prompt-generator A Docker-based web application that uses language models to generate and expand prompts for image generation tools like MidJourney and Stab… | 30 | 1158 | maintenance |
| HKUST-Aerial-Robotics/GVINS GVINS is a C++/ROS nonlinear optimization system that tightly fuses GNSS raw measurements (pseudorange and Doppler) with visual and inertia… | 32 | 1155 | maintenance |
| magese/ik-analyzer-solr A Java library that adapts the IK Chinese word segmentation analyzer for Apache Solr 7.x-8.x, distributed via Maven. It merges multiple Chi… | 23 | 1154 | maintenance |
| strib/scigen SCIgen is an automatic generator of random, syntactically valid but meaningless computer science research papers, originally from MIT CSAIL… | 32 | 1153 | maintenance |
| uber-research/PPLM PPLM (Plug and Play Language Model) is a research implementation for controlled text generation that steers the topic and attributes of GPT… | 32 | 1153 | maintenance |
| pyloque/fastscan FastScan is a JavaScript library implementing the Aho-Corasick algorithm for fast multi-pattern text search, commonly used for sensitive/ba… | 32 | 1148 | maintenance |
| tech-srl/code2vec Official TensorFlow implementation of the code2vec model from the POPL'2019 paper, which learns distributed vector representations of code … | 32 | 1147 | maintenance |
| ZhuiyiTechnology/roformer RoFormer is an MLM pre-trained language model built on rotary position embeddings (RoPE), a relative position encoding method with strong t… | 32 | 1146 | maintenance |
| AimeeLee77/keyword_extraction A Python project implementing Chinese text keyword extraction using three methods: TF-IDF, TextRank, and Word2Vec word clustering. It inclu… | 32 | 1145 | maintenance |
| anderscui/jieba.NET jieba.NET is a C# port of the popular jieba Chinese word segmentation library, supporting .NET Framework and .NET Core. It offers precise, … | 32 | 1143 | maintenance |
| patil-suraj/question_generation An open-source study and library for neural question generation using pre-trained seq2seq transformer models like T5 via Hugging Face trans… | 32 | 1141 | maintenance |
| Uahh/Slscq Slscq is a Chinese 'shenlun' (civil-service exam essay) generator written in C++. Given a topic and word count, it randomly assembles a pla… | 32 | 1136 | maintenance |
| harvardnlp/pytorch-struct Torch-Struct is a PyTorch library of tested, GPU-accelerated implementations of core structured prediction algorithms such as CRFs, HMMs, H… | 23 | 1133 | maintenance |
| fighting41love/cocoNLP cocoNLP is a Python library for Chinese information extraction from unstructured text. It extracts emails, phone numbers (with carrier and … | 32 | 1128 | maintenance |
| kohlschutter/boilerpipe boilerpipe is a Java library for removing boilerplate (ads, navigation, headers) from HTML pages and extracting the main full text content.… | 32 | 1127 | maintenance |
| xudaolong/CodeVar An Alfred workflow that translates Chinese phrases into English variable names in multiple naming conventions (camelCase, PascalCase, snake… | 23 | 1125 | maintenance |
| rhysd/vim-grammarous A Vim plugin that provides grammar checking by integrating with LanguageTool, which it can download automatically. It highlights grammar er… | 32 | 1122 | maintenance |
| microsoft/MASS MASS is Microsoft's PyTorch implementation of Masked Sequence to Sequence Pre-training for language generation tasks. It provides pre-train… | 10 | 1115 | maintenance |
| whyliam/whyliam.workflows.youdao An Alfred workflow for macOS that translates words and phrases between English and Chinese using Youdao's translation/dictionary service. I… | 50 | 1114 | maintenance |
| google-deepmind/dramatron Dramatron is a research tool from DeepMind that uses pre-trained large language models to hierarchically co-write theatre scripts and scree… | 32 | 1112 | maintenance |
| synesthesiam/voice2json voice2json is a collection of command-line tools for offline speech-to-text and intent recognition on Linux, supporting 18 languages via en… | 10 | 1105 | maintenance |
| RUCAIBox/TextBox TextBox 2.0 is a Python/PyTorch library providing a unified pipeline for applying pre-trained language models to text generation tasks. It … | 23 | 1097 | maintenance |
| maelfabien/Multimodal-Emotion-Recognition A real-time multimodal emotion recognition web app built with Flask that analyzes emotions from text, audio, and video inputs using deep le… | 32 | 1089 | maintenance |
| exorde-labs/exorde-client The Exorde client is a Python CLI worker node for the Exorde Network, a decentralized protocol where participants scrape social media and w… | 60 | 1085 | maintenance |
| PrincetonML/SIF A Python research library implementing the Smooth Inverse Frequency (SIF) weighting scheme for computing sentence embeddings, from the ICLR… | 32 | 1085 | maintenance |
| datumbox/datumbox-framework Datumbox is an open-source Machine Learning framework written in Java that enables rapid development of ML and statistical applications. It… | 32 | 1084 | maintenance |
| WangRongsheng/XrayGLM XrayGLM is the first Chinese multimodal medical large language model that generates radiology report summaries from chest X-ray images, bui… | 29 | 1082 | maintenance |
| openai/automated-interpretability OpenAI's code and tools for automatically generating, simulating, and scoring explanations of neuron behavior in language models, based on … | 10 | 1081 | maintenance |
| uber/queryparser A Haskell library for parsing and analyzing SQL queries in Vertica, Hive, and Presto dialects into a shared abstract syntax tree. It provid… | 32 | 1077 | maintenance |
| jiesutd/YEDDA YEDDA is a lightweight desktop GUI tool for manually annotating text spans with entity, chunk, or event labels, built with Python's tkinter… | 32 | 1071 | maintenance |
| OFA-Sys/ONE-PEACE ONE-PEACE is a general multimodal representation model that jointly encodes vision, audio, and language modalities without initializing fro… | 29 | 1060 | maintenance |
| chenyuntc/PyTorchText A PyTorch implementation of multiple text classification models (TextCNN, TextRNN/LSTM, RCNN, FastText, inception CNN) that won 1st place i… | 32 | 1057 | maintenance |
| atilika/kuromoji Kuromoji is a self-contained, easy-to-use Japanese morphological analyzer written in Java, supporting word segmentation, part-of-speech tag… | 32 | 1056 | maintenance |
| microsoft/Oscar Oscar is Microsoft's research code for object-semantics aligned cross-modal pre-training of vision-language models, with VinVL providing im… | 10 | 1053 | maintenance |
| microsoft/Cognitive-Samples-IntelligentKiosk A UWP sample application from Microsoft showcasing hands-free kiosk-style demos built on Azure Cognitive Services (Face, Computer Vision, T… | 10 | 1052 | maintenance |
| thunlp/OpenDelta OpenDelta is a Python library for parameter-efficient tuning (delta tuning) of pretrained language models, letting users attach small train… | 23 | 1046 | maintenance |
| facebookresearch/cc_net CCNet is a Python pipeline from Facebook AI Research for downloading, deduplicating, and cleaning Common Crawl web data into high-quality m… | 10 | 1045 | maintenance |
| pykaldi/pykaldi PyKaldi is a Python scripting layer providing wrappers for the C++ APIs of the Kaldi speech recognition toolkit and OpenFst library. It ena… | 47 | 1039 | maintenance |
| 1e0ng/simhash A Python implementation of the Simhash algorithm for near-duplicate detection and similarity estimation of text. It provides a small, focus… | 32 | 1039 | maintenance |
| bytedance/godlp godlp is a Go library from ByteDance for sensitive data discovery and de-identification (data loss prevention). It detects sensitive inform… | 10 | 1038 | maintenance |
| client9/libinjection A C library that tokenizes and analyzes input strings to detect SQL injection (SQLi) attacks using fingerprint matching. It has bindings fo… | 32 | 1032 | maintenance |
| dandelionsllm/pandallm Panda is an open-source project for overseas Chinese large language models, providing PandaLLM model weights (continued pretraining of LLaM… | 30 | 1031 | maintenance |
| BeautyyuYanli/full-mark-composition-generator A humorous web application that generates satirical 'full-mark' Chinese essays by randomly filling templates with jargon, famous quotes, an… | 23 | 1030 | maintenance |
| mimno/Mallet MALLET (MAchine Learning for LanguagE Toolkit) is a Java-based package for statistical natural language processing, including document clas… | 86 | 1028 | maintenance |
| turtlesoupy/this-word-does-not-exist A project that trains a GPT-2 variant to invent fake English words with generated definitions and example sentences, powering the thiswordd… | 72 | 1023 | maintenance |
| hankcs/AhoCorasickDoubleArrayTrie A Java library implementing the Aho-Corasick multi-pattern string matching algorithm on top of a Double Array Trie, achieving O(n) matching… | 23 | 1016 | maintenance |
| luozhouyang/python-string-similarity A Python 3 library implementing a dozen string similarity and distance algorithms, including Levenshtein variants, Jaro-Winkler, longest co… | 23 | 1016 | maintenance |
| ggeop/Python-ai-assistant Jarvis is a Python voice-controlled AI assistant for Linux that recognizes speech, responds conversationally, and executes commands like op… | 23 | 1014 | maintenance |
| kakaobrain/kogpt KakaoBrain's KoGPT, a Korean Generative Pre-trained Transformer (GPT) model with 6B parameters, distributed via Hugging Face with inference… | 23 | 1011 | maintenance |
| naver/splade SPLADE is a research library from NAVER for training, indexing, and retrieval with sparse neural search models based on BERT. It learns spa… | 23 | 1007 | maintenance |
| varunshenoy/GraphGPT GraphGPT is a web application that converts unstructured natural language text into a knowledge graph using GPT-3, visualizing entities and… | 31 | 4426 | experimental |
| LeeSureman/Flat-Lattice-Transformer Reference implementation of the ACL 2020 paper FLAT: Chinese NER Using Flat-Lattice Transformer, built on PyTorch and FastNLP. It trains fl… | 32 | 1003 | maintenance |
| UdaraJay/Pile Pile is an open-source desktop app for reflective journaling that keeps entries stored locally. It optionally integrates AI (OpenAI GPT-4 o… | 18 | 3113 | experimental |
| Priler/jarvis JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech… | 53 | 2909 | experimental |
| b-nnett/goose Goose is a local-first iOS companion app for WHOOP 5.0 fitness bands, built with SwiftUI and a Rust core that parses Bluetooth packet data … | 10 | 2718 | experimental |
| Nutlope/llama-ocr An npm library that performs OCR by sending images to Llama 3.2 Vision models via Together AI and returns structured Markdown. It supports … | 63 | 2431 | experimental |
| facebookresearch/large_concept_model Official PyTorch implementation of Meta's Large Concept Models (LCM), which perform language modeling by autoregressively predicting senten… | 23 | 2375 | experimental |
| gvzdv/claudish-to-english A Claude Code plugin that rewrites assistant messages into plain English using a local LLM via ollama, or alternatively the codex CLI, Anth… | 67 | 2287 | experimental |
| mshumer/gpt-investor An experimental AI agent notebook that uses Claude 3 models to analyze stocks in a given industry, gathering financial data, news, and anal… | 25 | 2265 | experimental |
| deepklarity/jupyter-text2code A proof-of-concept Jupyter Notebook extension that converts English queries into relevant Python code using sentence embedding models. It s… | 54 | 2082 | experimental |
| Klotzkette/claude-fuer-deutsches-recht An experimental collection of Claude skills, sub-agents, and workflows adapted for German legal practice, covering areas like employment, c… | 76 | 1507 | experimental |
| obsei/obsei Obsei is an open-source, low-code, AI-powered automation framework for text analysis workflows. It collects unstructured data from sources … | 52 | 1426 | experimental |
| anc95/writely Writely is an open-source browser extension for Chrome, Firefox, and Edge that brings GPT-powered AI writing assistance to any editable web… | 36 | 1297 | experimental |
| nethical6/conversation-steganography A Go CLI tool that hides encrypted secret messages inside natural-looking chat text generated by a local LLM (GPT-2), enabling covert commu… | 55 | 1239 | experimental |