function: nlp
1557 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| DimiMikadze/orca Orca is an AI agent application for deep LinkedIn profile analysis that scrapes posts, comments, reactions, and interaction networks, then … | 75 | 1297 | active |
| deepseek-ai/DeepSeek-Prover-V2 DeepSeek-Prover-V2 is an open-source large language model for formal theorem proving in Lean 4, trained via reinforcement learning with sub… | 34 | 1297 | active |
| codefuse-ai/codefuse-chatbot CodeFuse-ChatBot is an open-source AI assistant for the software development lifecycle, combining a multi-agent scheduling framework with D… | 27 | 1290 | active |
| SWI-Prolog/swipl-devel SWI-Prolog is a comprehensive, open-source (BSD-2) implementation of the Prolog logic programming language, implemented in C and Prolog wit… | 100 | 1278 | stable |
| neo4j/neo4j-graphrag-python The official Neo4j first-party Python library for building graph retrieval-augmented generation (GraphRAG) applications. It provides retrie… | 93 | 1275 | active |
| flutter-ml/google_ml_kit_flutter A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa… | 76 | 1274 | active |
| Renumics/spotlight Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe… | 93 | 1272 | active |
| jordanrendric/claude-video-vision A Claude Code plugin with an MCP server that gives Claude the ability to watch and understand videos by extracting frames via ffmpeg and tr… | 58 | 1265 | active |
| myreader-io/myGPTReader myGPTReader is a Slack bot powered by ChatGPT that reads and summarizes webpages, documents (eBooks, PDF, DOCX), and YouTube videos, and su… | 60 | 4419 | maintenance |
| TencentCloudADP/youtu-graphrag Youtu-GraphRAG is a Python framework for graph-based retrieval-augmented generation that unifies graph schema construction, community detec… | 57 | 1253 | active |
| minsight-ai-info/AI-Search-Hub AI Search Hub is an open-source Skill that aggregates native AI search capabilities from platforms like Gemini, Grok, Doubao, and Yuanbao i… | 50 | 1250 | active |
| xtreme1-io/xtreme1 Xtreme1 is an open-source, self-hosted data labeling and annotation platform for multimodal training data, supporting images, 3D LiDAR poin… | 62 | 1234 | active |
| DroppedNeedle/DroppedNeedle DroppedNeedle is a self-hosted music request and discovery application with a built-in library and download engine that replaces Lidarr. It… | 85 | 1229 | active |
| MoonshotAI/Kimi-VL Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo… | 33 | 1224 | active |
| GML-MMGroup/GMTalker GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l… | 46 | 1217 | active |
| InternScience/GraphGen GraphGen is a Python framework for knowledge-graph-guided synthetic data generation for LLM training. It builds fine-grained knowledge grap… | 59 | 1212 | active |
| OpenTSLM/OpenTSLM OpenTSLM is a family of Time-Series Language Models that integrate time series as a native modality into pretrained LLMs (Llama, Gemma), en… | 61 | 1211 | active |
| aTrainTranscription/aTrain aTrain is a desktop GUI application for offline transcription of speech recordings using Whisper-based machine learning models, with speake… | 79 | 1205 | active |
| modelscope/sirchmunk Sirchmunk is an agentic, embedding-free search engine that turns raw files into a self-evolving knowledge base in real time, without vector… | 83 | 1202 | active |
| superlinear-ai/raglite RAGLite is a Python toolkit for building Retrieval-Augmented Generation (RAG) pipelines backed by DuckDB or PostgreSQL, with hybrid keyword… | 85 | 1198 | active |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1195 | active |
| NVIDIA/audio-flamingo NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u… | 50 | 1182 | active |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1179 | active |
| ddlBoJack/emotion2vec Official PyTorch implementation of emotion2vec, a self-supervised pre-trained model for speech emotion representation. It provides code for… | 27 | 1179 | active |
| baichuan-inc/Baichuan2 Baichuan 2 is a family of open large language models (7B and 13B, Base and Chat variants with 4-bit quantized versions) trained by Baichuan… | 28 | 4084 | maintenance |
| magicrew/doc7 doc7 is a Go CLI tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using any… | 77 | 1173 | active |
| mlomb/chat-analytics A web app and CLI that takes chat exports from platforms like Telegram, Discord, and WhatsApp and generates a single interactive HTML repor… | 44 | 1167 | active |
| ALwrity/ALwrity ALwrity is an open-source, AI-first digital marketing platform built in Python that combines content strategy planning, multimodal AI conte… | 79 | 1153 | active |
| amazon-science/mm-cot Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw… | 31 | 3985 | maintenance |
| FlagOpen/RoboBrain2.5 RoboBrain 2.5 is an open-source embodied AI foundation model from BAAI that combines multimodal large language model capabilities with 3D s… | 50 | 1132 | active |
| NovaSearch-Team/RAG-Retrieval A Python library and toolkit for unified fine-tuning, inference, and distillation of RAG retrieval models, including embedding models, ColB… | 60 | 1127 | active |
| K-Dense-AI/k-dense-byok K-Dense BYOK is a free, open-source desktop application providing 'Kady', an AI research assistant (co-scientist) that runs locally and use… | 78 | 1126 | active |
| FlagAI-Open/FlagAI FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int… | 64 | 3869 | maintenance |
| microsoft/onnxruntime-genai ONNX Runtime GenAI is a C++ library with Python, C#, C/C++, and Java APIs for running generative AI models (LLMs, Whisper, vision-language … | 93 | 1116 | active |
| FutureUniant/Tailor Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea… | 37 | 1115 | active |
| easydoc-ai/easydoc EasyDoc is a multimodal document processing API that converts unstructured documents like PDFs into hierarchical, machine-readable JSON. It… | 35 | 1115 | active |
| amazon-science/RAGChecker RAGChecker is an automatic evaluation framework for diagnosing Retrieval-Augmented Generation (RAG) systems. It provides holistic and diagn… | 25 | 1110 | active |
| ycccccccy/echotrace EchoTrace is a fully local, privacy-focused desktop application for exporting, decrypting, and analyzing WeChat chat records. It generates … | 60 | 3782 | maintenance |
| rhymes-ai/Aria Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda… | 23 | 1087 | active |
| minghanqin/LangSplat Official implementation of LangSplat, a CVPR 2024 Highlight paper that constructs a 3D language field using 3D Gaussian Splatting with CLIP… | 47 | 1077 | active |
| Muesli-HQ/muesli Muesli is an open-source native macOS app that combines hotkey-driven AI dictation with local meeting transcription, running speech-to-text… | 82 | 1072 | active |
| titanwings/ex-skill A Python tool that generates Claude Code / OpenClaw agent 'skills' capturing a specific person's texting persona from chat history (WeChat,… | 62 | 1072 | active |
| AILab-CVC/UniRepLKNet UniRepLKNet is a large-kernel ConvNet architecture (CVPR 2024, TPAMI 2025) that provides universal perception across image, audio, video, p… | 43 | 1072 | stable |
| GanymedeNil/document.ai A universal local knowledge base solution that stores documents as vectors in a vector database and uses GPT3.5 to generate answers from re… | 30 | 3670 | maintenance |
| kagisearch/kite-public Kite is the open-source front-end for Kagi News, a web news application that delivers daily curated, LLM-summarized news stories from diver… | 59 | 1064 | active |
| mittagessen/kraken kraken is a turn-key OCR/HTR engine built on neural networks, optimized for historical and non-Latin script material. It provides trainable… | 99 | 1061 | active |
| InternRobotics/InternNav InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation… | 58 | 1061 | active |
| spider-ios/autox-release AutoX is a desktop social media operations tool that automates one-click publishing of videos to multiple platforms such as Douyin, TikTok,… | 65 | 1059 | active |
| sb-ai-lab/EmotiEffLib EmotiEffLib (formerly HSEmotion) is a lightweight library for facial emotion and engagement recognition in photos and videos, available in … | 65 | 1057 | active |
| alibaba/Alink Alink is a machine learning algorithm platform built on Apache Flink, developed by Alibaba's PAI team. It provides a large library of batch… | 23 | 3611 | maintenance |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1056 | active |
| CRui5in/paper-ppt-agent A self-hosted web application that uses a multi-agent AI pipeline (Strategist, Executor, Critic) to convert academic papers from PDF or LaT… | 79 | 1048 | active |
| wu-yc/LabClaw LabClaw is a modular library of 240 SKILL.md agent skills for biomedical AI research, covering biology, drug discovery, medicine, data scie… | 48 | 1048 | active |
| shang-zhu/violin Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi… | 52 | 1047 | active |
| agents-flex/agents-flex Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI… | 93 | 1046 | active |
| ml-tooling/ml-workspace ML Workspace is an all-in-one web-based IDE Docker image specialized for machine learning and data science. It bundles Jupyter, JupyterLab,… | 23 | 3544 | maintenance |
| ocropus-archive/DUP-ocropy OCRopy is a collection of Python-based tools for document analysis and OCR, covering binarization, page layout analysis, and text line reco… | 10 | 3465 | maintenance |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3454 | maintenance |
| EvolvingLMMs-Lab/Otter Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa… | 21 | 3436 | maintenance |
| braintrustdata/autoevals Autoevals is a Python and TypeScript library for automatically evaluating AI model outputs using LLM-as-a-judge, heuristic, and statistical… | 85 | 1009 | active |
| InsiderX-Pro/video-translator An OpenClaw agent skill (written in Shell/Python) that translates and dubs videos by submitting jobs to a remote video-translation service … | 54 | 1007 | active |
| octimot/StoryToolkitAI StoryToolkitAI is a desktop film editing tool that transcribes, indexes, and semantically searches video footage locally, using speech reco… | 65 | 1006 | active |
| PrithivirajDamodaran/FlashRank FlashRank is an ultra-light, fast Python library for re-ranking search and retrieval results using SoTA cross-encoders and listwise LLM-bas… | 58 | 1003 | active |
| agenmod/immortal-skill An open-source 'digital immortality' framework that distills a person's persona from chat logs and documents across 12+ platforms (WeChat, … | 49 | 1002 | active |
| ZJUI-AI4H/Hulu-Med Hulu-Med is a family of open-source transparent generalist medical vision-language models ranging from 4B to 235B parameters, covering text… | 62 | 1001 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| catalyst-team/catalyst Catalyst is a high-level PyTorch framework for deep learning research and development, focused on reproducibility, rapid experimentation, a… | 64 | 3382 | maintenance |
| resemble-ai/Resemblyzer Resemblyzer is a Python package that uses a deep learning voice encoder to convert speech audio into 256-dimensional voice embeddings. Thes… | 23 | 3300 | maintenance |
| bytedance/lightseq LightSeq is a high-performance CUDA-based library for training and inference of sequence models like Transformer, BERT, GPT, and BART, with… | 10 | 3295 | maintenance |
| huawei-noah/Pretrained-Language-Model A collection of pretrained language models and optimization techniques from Huawei Noah's Ark Lab, including PanGu-α (200B-parameter Chines… | 32 | 3165 | maintenance |
| Conchylicultor/DeepQA DeepQA is a TensorFlow implementation of Google's 'A Neural Conversational Model', a seq2seq RNN-based deep learning chatbot. It supports t… | 32 | 2910 | maintenance |
| cskefu/cskefu CSKeFu (Chunsong Customer Service) is an open-source enterprise customer service / contact center platform providing agent workbenches, con… | 59 | 2894 | maintenance |
| alisen39/TrWebOCR TrWebOCR is an open-source offline Chinese OCR service built on the Tr project, exposing both a web UI and HTTP API for text recognition. I… | 23 | 2878 | maintenance |
| tensorflow/lingvo Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach… | 72 | 2864 | maintenance |
| readbeyond/aeneas aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a… | 65 | 2863 | maintenance |
| yanqiangmiffy/Chinese-LangChain A Chinese-language local knowledge base question-answering application built on ChatGLM-6B and LangChain, with a Gradio web UI. It supports… | 30 | 2824 | maintenance |
| illuin-tech/colpali ColPali Engine is the Python library for training and running inference with ColVision visual document retrieval models such as ColPali, Co… | 91 | 2798 | maintenance |
| microsoft/NUWA Microsoft's official research repository for the NUWA family of multimodal generative models, a unified 3D transformer pipeline for visual … | 10 | 2791 | maintenance |
| ctripcorp/C-OCR C-OCR is Ctrip's in-house OCR project focused on recognizing travel-related documents such as ID cards, passports, train tickets, and visas… | 32 | 2476 | maintenance |
| zai-org/CogVLM2 CogVLM2 is an open-source multi-modal vision-language model family built on Meta-Llama-3-8B-Instruct, offering image and video understandin… | 28 | 2433 | maintenance |
| microsoft/DialoGPT DialoGPT is a large-scale pretrained dialogue response generation model from Microsoft, based on GPT-2 and trained on 147M multi-turn Reddi… | 10 | 2420 | maintenance |
| OpenBMB/CPM-Bee CPM-Bee is a fully open-source, commercially usable 10-billion-parameter bilingual (Chinese-English) foundation large language model traine… | 71 | 2406 | maintenance |
| allenai/RL4LMs RL4LMs is a modular Python library from AllenAI for fine-tuning language models with reinforcement learning to align with human preferences… | 32 | 2394 | maintenance |
| MetaGLM/FinGLM FinGLM is an open, community-driven financial LLM project centered on a dialog-based question-answering system that analyzes Chinese listed… | 27 | 2258 | maintenance |
| kingyiusuen/image-to-latex A PyTorch application that converts images of LaTeX math equations into LaTeX code using a ResNet-18 encoder and Transformer decoder traine… | 32 | 2159 | maintenance |
| facebookresearch/chameleon Repository for Meta Chameleon, an early-fusion token-based mixed-modal foundation model that understands and generates interleaved images a… | 10 | 2103 | maintenance |
| ai-forever/ru-gpts A repository of Russian GPT-3 language models (ruGPT3XL/Large/Medium/Small and ruGPT2Large) with usage and fine-tuning examples. It provide… | 32 | 2087 | maintenance |
| vaguileradiaz/tinfoleak tinfoleak is an open-source Python tool for OSINT/SOCMINT analysis of Twitter accounts, extracting structured intelligence such as user act… | 32 | 1980 | maintenance |
| r9y9/deepvoice3_pytorch A PyTorch implementation of Deep Voice 3 and related convolutional neural network-based text-to-speech synthesis models. It includes traini… | 23 | 1975 | maintenance |
| Ucas-HaoranWei/Vary Official ECCV 2024 implementation of Vary, a method for scaling up the vision vocabulary of large vision-language models. It provides train… | 26 | 1889 | maintenance |
| facebookresearch/DPR DPR is a set of tools and pretrained models for dense passage retrieval in open-domain question answering, based on the EMNLP 2020 paper fr… | 10 | 1869 | maintenance |
| huggingface/transfer-learning-conv-ai A clean, commented PyTorch codebase for training a dialog/chatbot agent by transfer learning from OpenAI GPT/GPT-2 language models. It repr… | 32 | 1754 | maintenance |
| Lightning-Universe/lightning-bolts Lightning Bolts is a toolbox of pre-built models, callbacks, and datasets that extend PyTorch Lightning for AI/ML research and production. … | 10 | 1751 | maintenance |
| magenta/mt3 MT3 is a multi-instrument automatic music transcription model built on the T5X framework, converting audio recordings into multitrack MIDI.… | 73 | 1744 | maintenance |
| Lightning-Universe/lightning-flash Lightning Flash is a high-level PyTorch library built on PyTorch Lightning that provides ready-made 'recipes' for over 15 AI tasks across 7… | 10 | 1722 | maintenance |
| imcaspar/gpt2-ml GPT2-ML is a TensorFlow-based GPT-2 training and inference project supporting multiple languages, with a focus on Chinese. It provides 1.5B… | 23 | 1701 | maintenance |
| dbashford/textract A Node.js library that extracts plain text from many document formats including HTML, PDF, DOC/DOCX, XLS/XLSX, CSV, PPTX, RTF, EPUB, and im… | 58 | 1694 | maintenance |
| invictus717/MetaTransformer Meta-Transformer is a research framework for unified multimodal learning that maps inputs from 12 modalities (text, images, point clouds, a… | 19 | 1647 | maintenance |
| mmz-001/knowledge_gpt KnowledgeGPT is a Streamlit web application that lets users upload documents (primarily PDFs) and ask questions about them, receiving answe… | 10 | 1635 | maintenance |
| jina-ai/thinkgpt ThinkGPT is a Python library implementing Chain of Thought techniques for LLMs, providing memory, self-refinement, knowledge compression, a… | 30 | 1583 | maintenance |