function: nlp
1557 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| tensorflow/tensor2tensor Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak… | 10 | 17464 | maintenance |
| ucbepic/docetl DocETL is a Python library and CLI for building LLM-powered data processing and ETL pipelines over structured and unstructured data using d… | 86 | 3995 | active |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3980 | active |
| gusye1234/nano-graphrag nano-graphrag is a lightweight, ~1100-line Python implementation of Microsoft's GraphRAG, designed to be small, fast, and easy to read or h… | 54 | 3974 | active |
| OSU-NLP-Group/HippoRAG HippoRAG is a Python RAG framework inspired by human long-term memory that combines LLMs, knowledge graphs, and Personalized PageRank to in… | 59 | 3967 | active |
| UditAkhourii/adhd ADHD is a TypeScript skill for coding agents built on the Claude and Codex Agent SDKs that implements parallel divergent ideation via tree-… | 65 | 3953 | active |
| yuanzhongqiao/printfilm Printfilm is a self-hostable AI short-drama (short film / motion comic) creation SaaS platform built on Next.js and Spring Boot. It provide… | 72 | 3913 | active |
| NVlabs/VILA VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d… | 57 | 3857 | active |
| circlemind-ai/fast-graphrag fast-graphrag is a Python library providing a streamlined, promptable GraphRAG framework for interpretable, high-precision retrieval workfl… | 44 | 3849 | active |
| VectifyAI/OpenKB OpenKB is an open-source CLI tool that compiles raw documents (PDF, Word, Markdown, HTML, and more) into a structured, interlinked wiki-sty… | 76 | 3847 | active |
| zly2006/zhihu-plus-plus Zhihu++ is an open-source third-party Android client for the Chinese Q&A platform Zhihu, built in Kotlin, that removes ads, promotional pos… | 85 | 3845 | active |
| morphik-org/morphik-core Morphik Core is an open-source, source-available multimodal retrieval engine for building RAG applications over unstructured data like PDFs… | 64 | 3708 | active |
| Mouseww/anything-analyzer An Electron-based all-in-one protocol analysis toolkit that captures HTTP(S) traffic from browsers, desktop apps, terminals, scripts, and m… | 77 | 3586 | active |
| whoiskatrin/chart-gpt Chart-GPT is a web application that generates charts from natural language text input using AI. Users describe what they want to visualize … | 42 | 3583 | active |
| sligter/LandPPT LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p… | 81 | 3575 | active |
| PKU-YuanGroup/Video-LLaVA Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into… | 27 | 3500 | active |
| NVlabs/Eagle Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L… | 64 | 3462 | active |
| antiboredom/videogrep Videogrep is a Python command line tool that searches through dialog in video or audio files using subtitle tracks or speech transcriptions… | 23 | 3461 | active |
| SamurAIGPT/llm-wiki-agent A coding-agent skill that turns source documents into a self-maintaining, interlinked markdown wiki. You drop files into a raw/ folder and … | 74 | 3458 | active |
| Grt1228/chatgpt-java An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati… | 21 | 3423 | active |
| OpenGVLab/Ask-Anything VideoChat/Ask-Anything is a family of multimodal chat models and demos that combine video understanding with large language models, letting… | 72 | 3346 | active |
| Kedreamix/Linly-Dubbing Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L… | 27 | 3331 | active |
| notdog1998/yourself-skill A Claude Code skill that builds a digital persona of yourself from chat logs, diaries, photos, and self-descriptions, structured as a Self … | 48 | 3324 | active |
| timerring/bilive BILIVE is a Python application that records Bilibili live streams and danmaku 24/7, then automatically renders danmaku and AI-generated sub… | 61 | 3275 | active |
| vladmandic/human Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod… | 48 | 3264 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3248 | active |
| imraywang/wewrite WeWrite is a Python-based AI agent skill that automates the full WeChat Official Account content pipeline: topic selection from trending ne… | 79 | 3182 | active |
| CatchTheTornado/text-extract-api A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC… | 45 | 3175 | active |
| Filimoa/open-parse Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG … | 64 | 3159 | active |
| liyown/ai-trend-publish TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s… | 82 | 3158 | active |
| SciSharp/BotSharp BotSharp is an open-source multi-agent AI framework written in C# for .NET, providing an agent abstraction layer, conversation state manage… | 79 | 3098 | active |
| blazickjp/arxiv-mcp-server A Model Context Protocol (MCP) server that lets AI agents search, download, and analyze arXiv papers, including reading original LaTeX sect… | 88 | 3076 | active |
| viperrcrypto/Siftly Siftly is a self-hosted, local-first web application for organizing Twitter/X bookmarks into a searchable, categorized knowledge base. It r… | 64 | 2982 | active |
| pharmapsychotic/clip-interrogator A Python library that combines OpenAI's CLIP and Salesforce's BLIP to reverse-engineer text prompts from images, optimized for use with tex… | 23 | 2982 | stable |
| OpenMOSS/MOSS MOSS is an open-source tool-augmented conversational large language model from Fudan University, released with base models, SFT models, plu… | 68 | 12230 | maintenance |
| mazzzystar/Queryable Queryable is an open-source iOS app that runs Apple's MobileCLIP (formerly OpenAI's CLIP) entirely on-device to search your photo album wit… | 62 | 2977 | active |
| Simon-He95/markstream-vue A family of streaming Markdown renderer components for AI chat and LLM token-stream UIs, with markstream-vue as the stable Vue 3/Nuxt/ViteP… | 82 | 2967 | stable |
| InternLM/InternLM-XComposer InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u… | 38 | 2925 | active |
| ogkalu2/comic-translate An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language… | 90 | 2911 | active |
| cdhigh/KindleEar KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep… | 73 | 2865 | active |
| VRSEN/OpenSwarm OpenSwarm is an open-source multi-agent system built on Agency Swarm that turns a single terminal prompt into complete deliverables like sl… | 77 | 2854 | active |
| superlinked/sie SIE (Superlinked Inference Engine) is an open-source, self-hosted inference server and production cluster that serves 100+ open models (emb… | 92 | 2830 | active |
| satijalab/seurat Seurat is an R toolkit for single-cell genomics developed by the Satija Lab, providing a complete pipeline for analyzing single-cell RNA-se… | 92 | 2790 | stable |
| imanoop7/Ollama-OCR A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports… | 26 | 2780 | active |
| kha-white/manga-ocr Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to… | 90 | 2758 | stable |
| apple/turicreate Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj… | 10 | 11159 | maintenance |
| protectai/vulnhuntr Vulnhuntr is a Python CLI tool that uses large language models combined with static code analysis to autonomously discover exploitable vuln… | 24 | 2747 | active |
| Audiveris/audiveris Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat… | 97 | 2727 | active |
| naiveHobo/InvoiceNet InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN… | 32 | 2694 | active |
| JIA-Lab-research/LISA LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati… | 31 | 2674 | active |
| darkzOGx/youtube-automation-agent AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts… | 82 | 2666 | active |
| hgmzhn/manga-translator-ui A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,… | 80 | 2651 | active |
| aardio/ImTip ImTip is a lightweight (under 1 MB) Windows desktop assistant that shows real-time input method and keyboard state indicators at the text c… | 74 | 2610 | active |
| yazinsai/OpenOats OpenOats is a macOS meeting assistant that transcribes both sides of a call in real time using on-device speech recognition and surfaces re… | 77 | 2565 | active |
| X-PLUG/mPLUG-Owl mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and… | 36 | 2539 | active |
| microsoft/ResearchStudio ResearchStudio is a Microsoft collection of AI agent skills that cover the entire research lifecycle, from an under-specified research dire… | 59 | 2523 | active |
| InternLM/HuixiangDou HuixiangDou is an LLM-based professional knowledge assistant designed for group chat scenarios, using a three-stage pipeline of preprocess,… | 46 | 2499 | active |
| pingcap/ossinsight OSSInsight is a web analytics platform that analyzes over 10 billion GitHub events to provide rankings, trends, and comparisons of open sou… | 67 | 2498 | active |
| Live-GalGame/LiveGalGame LiveGalGame is a playful app that overlays a visual-novel (galgame) interface onto real-life conversations, providing real-time speech-to-t… | 58 | 2490 | active |
| Zafer-Liu/Data-Analysis-Agent An LLM-powered conversational data analysis agent that lets users connect data sources (Excel/CSV files, databases, Feishu tables) and ask … | 81 | 2442 | active |
| nz-m/SocialEcho SocialEcho is a full-featured social networking platform built on the MERN stack (MongoDB, Express.js, React.js, Node.js) with automated co… | 31 | 2435 | active |
| wangshub/Douyin-Bot A Python bot that automates the Douyin (TikTok China) mobile app via ADB, taking screenshots and calling a face-recognition API to auto-lik… | 32 | 9631 | maintenance |
| Alan AI SDK Alan AI SDK is a set of client libraries for embedding Alan AI's conversational AI agents and intelligent app layer into web, iOS, Android,… | 93 | 2431 | active |
| Zleap-AI/SAG SAG is an open-source retrieval architecture and knowledge base application that replaces both traditional RAG and GraphRAG with event-enti… | 83 | 2425 | active |
| X-PLUG/mPLUG-DocOwl mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO… | 39 | 2411 | active |
| Cicada000/VV A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d… | 35 | 2375 | active |
| zai-org/GLM-V GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil… | 60 | 2370 | active |
| OpenGVLab/InternVideo InternVideo is a series of open-source video foundation models for multimodal video understanding, spanning generative and discriminative l… | 72 | 2368 | active |
| Natively-AI-assistant/natively-cluely-ai-assistant Natively is a free, source-available desktop AI meeting assistant and interview copilot that provides real-time transcription, AI-generated… | 82 | 2354 | active |
| Turbo1123/roubao Roubao is an open-source AI phone automation assistant for Android, built natively in Kotlin and powered by vision-language models. It runs… | 53 | 2325 | active |
| PKU-YuanGroup/MoE-LLaVA MoE-LLaVA is an open-source Mixture-of-Experts based sparse large vision-language model, released with the MoE-Tuning training strategy fro… | 31 | 2322 | active |
| facebookresearch/ImageBind A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud… | 54 | 9064 | maintenance |
| Anakin-Inc/anakin AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s… | 73 | 2302 | active |
| Xiangyu-CAS/xiaohongshu-ops-skill A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f… | 50 | 2291 | active |
| perkfly/ex-skill A Claude Code / OpenClaw skill generator that distills chat logs, photos, and social media exports into a persona 'Skill' that mimics a spe… | 48 | 2272 | active |
| smittix/intercept iNTERCEPT is a free, open-source, web-based signal intelligence (SIGINT) platform that unifies dozens of software-defined radio tools into … | 78 | 2269 | active |
| codedogQBY/ReadAny ReadAny is a local-first, AI-powered cross-platform e-book reader for desktop and mobile. It combines RAG-based chat with your books, hybri… | 81 | 2267 | active |
| guy-hartstein/company-research-agent A multi-agent company research application built with LangGraph and Tavily that generates comprehensive due-diligence reports on any compan… | 75 | 2250 | active |
| vasu-devs/JustHireMe JustHireMe is a local-first desktop workbench (Tauri frontend, Python backend) that scrapes job postings, ranks role fit against your profi… | 78 | 2236 | stable |
| apconw/Aix-DB Aix-DB is an AI-powered data analysis system (ChatBI) built on LangChain/LangGraph with an MCP Skills multi-agent architecture, converting … | 72 | 2234 | active |
| microsoft/LLaVA-Med LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su… | 40 | 2231 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| lkarlslund/Adalanche Adalanche is an open-source Active Directory attack graph visualizer and explorer written in Go. It collects data via LDAP and SYSVOL, anal… | 67 | 2195 | active |
| google/trax Trax is an end-to-end deep learning library built on JAX and TensorFlow that focuses on clear code and speed, developed and maintained by t… | 10 | 8306 | maintenance |
| PKU-YuanGroup/LLaVA-CoT LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea… | 47 | 2132 | active |
| 6551Team/opennews-mcp An MCP (Model Context Protocol) server that aggregates real-time news and market data from 85+ sources across news, exchange listings, on-c… | 59 | 2125 | active |
| coolight7/musicxx Musicxx (拟声) is a cross-platform audio and video player supporting local files, cloud drives (Baidu, Aliyun, 115), WebDAV, and NAS media se… | 71 | 2100 | active |
| walterlow/freecut FreeCut is a professional-grade, browser-based multi-track video editor requiring no installation or uploads, with all media and project fi… | 60 | 2096 | active |
| raphaelmansuy/edgequake EdgeQuake is a high-performance Graph-RAG framework written in Rust, inspired by LightRAG, that transforms documents (PDFs, markdown, text)… | 79 | 2078 | active |
| intel/openvino-plugins-ai-audacity A set of AI-enabled effects, generators, and analyzers for Audacity, powered by Intel OpenVINO running fully locally on CPU, GPU, or NPU. I… | 73 | 2066 | active |
| WenjieDu/PyPOTS PyPOTS is a Python toolbox for machine learning and data mining on partially-observed time series with missing values. It integrates 50+ st… | 93 | 2051 | active |
| mli/autocut AutoCut is a Python CLI tool that automatically transcribes video audio into subtitles using Whisper, then cuts video segments based on whi… | 32 | 7790 | maintenance |
| QwenLM/Qwen3-Embedding Qwen3-Embedding is a series of text embedding and reranking models (0.6B, 4B, 8B) built on Qwen3 foundation models, with a Python repositor… | 38 | 2016 | active |
| zai-org/GLM-130B GLM-130B is an open bilingual (English and Chinese) 130-billion-parameter dense language model pre-trained with the General Language Model … | 32 | 7651 | maintenance |
| microsoft/msticpy msticpy is a Python library from Microsoft for security investigation and threat hunting in Jupyter notebooks. It provides data acquisition… | 92 | 1995 | active |
| showlab/Show-o Show-o is a research repository implementing a unified transformer model that combines autoregressive and discrete diffusion modeling for m… | 50 | 1973 | active |
| FireRedTeam/FireRedASR FireRedASR is a family of open-source industrial-grade automatic speech recognition models supporting Mandarin, Chinese dialects, and Engli… | 52 | 1971 | active |
| run-llama/notebookllama NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval… | 52 | 1967 | active |
| SaiAkhil066/CORTEX-AI-SUPER-RAG CORTEX RAG is a local-first, agentic retrieval-augmented generation application that lets users upload PDFs and ask questions with cited an… | 61 | 1962 | active |