Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: nlp

1557 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
QwenLM/Qwen2.5-Omni
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre…
314074active
tensorflow/tensor2tensor
Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak…
1017464maintenance
ucbepic/docetl
DocETL is a Python library and CLI for building LLM-powered data processing and ETL pipelines over structured and unstructured data using d…
863995active
QwenLM/Qwen3-Omni
Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an…
523980active
gusye1234/nano-graphrag
nano-graphrag is a lightweight, ~1100-line Python implementation of Microsoft's GraphRAG, designed to be small, fast, and easy to read or h…
543974active
OSU-NLP-Group/HippoRAG
HippoRAG is a Python RAG framework inspired by human long-term memory that combines LLMs, knowledge graphs, and Personalized PageRank to in…
593967active
UditAkhourii/adhd
ADHD is a TypeScript skill for coding agents built on the Claude and Codex Agent SDKs that implements parallel divergent ideation via tree-…
653953active
yuanzhongqiao/printfilm
Printfilm is a self-hostable AI short-drama (short film / motion comic) creation SaaS platform built on Next.js and Spring Boot. It provide…
723913active
NVlabs/VILA
VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d…
573857active
circlemind-ai/fast-graphrag
fast-graphrag is a Python library providing a streamlined, promptable GraphRAG framework for interpretable, high-precision retrieval workfl…
443849active
VectifyAI/OpenKB
OpenKB is an open-source CLI tool that compiles raw documents (PDF, Word, Markdown, HTML, and more) into a structured, interlinked wiki-sty…
763847active
zly2006/zhihu-plus-plus
Zhihu++ is an open-source third-party Android client for the Chinese Q&A platform Zhihu, built in Kotlin, that removes ads, promotional pos…
853845active
morphik-org/morphik-core
Morphik Core is an open-source, source-available multimodal retrieval engine for building RAG applications over unstructured data like PDFs…
643708active
Mouseww/anything-analyzer
An Electron-based all-in-one protocol analysis toolkit that captures HTTP(S) traffic from browsers, desktop apps, terminals, scripts, and m…
773586active
whoiskatrin/chart-gpt
Chart-GPT is a web application that generates charts from natural language text input using AI. Users describe what they want to visualize …
423583active
sligter/LandPPT
LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p…
813575active
PKU-YuanGroup/Video-LLaVA
Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into…
273500active
NVlabs/Eagle
Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L…
643462active
antiboredom/videogrep
Videogrep is a Python command line tool that searches through dialog in video or audio files using subtitle tracks or speech transcriptions…
233461active
SamurAIGPT/llm-wiki-agent
A coding-agent skill that turns source documents into a self-maintaining, interlinked markdown wiki. You drop files into a raw/ folder and …
743458active
Grt1228/chatgpt-java
An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati…
213423active
OpenGVLab/Ask-Anything
VideoChat/Ask-Anything is a family of multimodal chat models and demos that combine video understanding with large language models, letting…
723346active
Kedreamix/Linly-Dubbing
Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L…
273331active
notdog1998/yourself-skill
A Claude Code skill that builds a digital persona of yourself from chat logs, diaries, photos, and self-descriptions, structured as a Self …
483324active
timerring/bilive
BILIVE is a Python application that records Bilibili live streams and danmaku 24/7, then automatically renders danmaku and AI-generated sub…
613275active
vladmandic/human
Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod…
483264active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983248active
imraywang/wewrite
WeWrite is a Python-based AI agent skill that automates the full WeChat Official Account content pipeline: topic selection from trending ne…
793182active
CatchTheTornado/text-extract-api
A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC…
453175active
Filimoa/open-parse
Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG …
643159active
liyown/ai-trend-publish
TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s…
823158active
SciSharp/BotSharp
BotSharp is an open-source multi-agent AI framework written in C# for .NET, providing an agent abstraction layer, conversation state manage…
793098active
blazickjp/arxiv-mcp-server
A Model Context Protocol (MCP) server that lets AI agents search, download, and analyze arXiv papers, including reading original LaTeX sect…
883076active
viperrcrypto/Siftly
Siftly is a self-hosted, local-first web application for organizing Twitter/X bookmarks into a searchable, categorized knowledge base. It r…
642982active
pharmapsychotic/clip-interrogator
A Python library that combines OpenAI's CLIP and Salesforce's BLIP to reverse-engineer text prompts from images, optimized for use with tex…
232982stable
OpenMOSS/MOSS
MOSS is an open-source tool-augmented conversational large language model from Fudan University, released with base models, SFT models, plu…
6812230maintenance
mazzzystar/Queryable
Queryable is an open-source iOS app that runs Apple's MobileCLIP (formerly OpenAI's CLIP) entirely on-device to search your photo album wit…
622977active
Simon-He95/markstream-vue
A family of streaming Markdown renderer components for AI chat and LLM token-stream UIs, with markstream-vue as the stable Vue 3/Nuxt/ViteP…
822967stable
InternLM/InternLM-XComposer
InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u…
382925active
ogkalu2/comic-translate
An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language…
902911active
cdhigh/KindleEar
KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep…
732865active
VRSEN/OpenSwarm
OpenSwarm is an open-source multi-agent system built on Agency Swarm that turns a single terminal prompt into complete deliverables like sl…
772854active
superlinked/sie
SIE (Superlinked Inference Engine) is an open-source, self-hosted inference server and production cluster that serves 100+ open models (emb…
922830active
satijalab/seurat
Seurat is an R toolkit for single-cell genomics developed by the Satija Lab, providing a complete pipeline for analyzing single-cell RNA-se…
922790stable
imanoop7/Ollama-OCR
A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports…
262780active
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
apple/turicreate
Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj…
1011159maintenance
protectai/vulnhuntr
Vulnhuntr is a Python CLI tool that uses large language models combined with static code analysis to autonomously discover exploitable vuln…
242747active
Audiveris/audiveris
Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat…
972727active
naiveHobo/InvoiceNet
InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN…
322694active
JIA-Lab-research/LISA
LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati…
312674active
darkzOGx/youtube-automation-agent
AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts…
822666active
hgmzhn/manga-translator-ui
A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,…
802651active
aardio/ImTip
ImTip is a lightweight (under 1 MB) Windows desktop assistant that shows real-time input method and keyboard state indicators at the text c…
742610active
yazinsai/OpenOats
OpenOats is a macOS meeting assistant that transcribes both sides of a call in real time using on-device speech recognition and surfaces re…
772565active
X-PLUG/mPLUG-Owl
mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and…
362539active
microsoft/ResearchStudio
ResearchStudio is a Microsoft collection of AI agent skills that cover the entire research lifecycle, from an under-specified research dire…
592523active
InternLM/HuixiangDou
HuixiangDou is an LLM-based professional knowledge assistant designed for group chat scenarios, using a three-stage pipeline of preprocess,…
462499active
pingcap/ossinsight
OSSInsight is a web analytics platform that analyzes over 10 billion GitHub events to provide rankings, trends, and comparisons of open sou…
672498active
Live-GalGame/LiveGalGame
LiveGalGame is a playful app that overlays a visual-novel (galgame) interface onto real-life conversations, providing real-time speech-to-t…
582490active
Zafer-Liu/Data-Analysis-Agent
An LLM-powered conversational data analysis agent that lets users connect data sources (Excel/CSV files, databases, Feishu tables) and ask …
812442active
nz-m/SocialEcho
SocialEcho is a full-featured social networking platform built on the MERN stack (MongoDB, Express.js, React.js, Node.js) with automated co…
312435active
wangshub/Douyin-Bot
A Python bot that automates the Douyin (TikTok China) mobile app via ADB, taking screenshots and calling a face-recognition API to auto-lik…
329631maintenance
Alan AI SDK
Alan AI SDK is a set of client libraries for embedding Alan AI's conversational AI agents and intelligent app layer into web, iOS, Android,…
932431active
Zleap-AI/SAG
SAG is an open-source retrieval architecture and knowledge base application that replaces both traditional RAG and GraphRAG with event-enti…
832425active
X-PLUG/mPLUG-DocOwl
mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO…
392411active
Cicada000/VV
A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d…
352375active
zai-org/GLM-V
GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil…
602370active
OpenGVLab/InternVideo
InternVideo is a series of open-source video foundation models for multimodal video understanding, spanning generative and discriminative l…
722368active
Natively-AI-assistant/natively-cluely-ai-assistant
Natively is a free, source-available desktop AI meeting assistant and interview copilot that provides real-time transcription, AI-generated…
822354active
Turbo1123/roubao
Roubao is an open-source AI phone automation assistant for Android, built natively in Kotlin and powered by vision-language models. It runs…
532325active
PKU-YuanGroup/MoE-LLaVA
MoE-LLaVA is an open-source Mixture-of-Experts based sparse large vision-language model, released with the MoE-Tuning training strategy fro…
312322active
facebookresearch/ImageBind
A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud…
549064maintenance
Anakin-Inc/anakin
AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s…
732302active
Xiangyu-CAS/xiaohongshu-ops-skill
A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f…
502291active
perkfly/ex-skill
A Claude Code / OpenClaw skill generator that distills chat logs, photos, and social media exports into a persona 'Skill' that mimics a spe…
482272active
smittix/intercept
iNTERCEPT is a free, open-source, web-based signal intelligence (SIGINT) platform that unifies dozens of software-defined radio tools into …
782269active
codedogQBY/ReadAny
ReadAny is a local-first, AI-powered cross-platform e-book reader for desktop and mobile. It combines RAG-based chat with your books, hybri…
812267active
guy-hartstein/company-research-agent
A multi-agent company research application built with LangGraph and Tavily that generates comprehensive due-diligence reports on any compan…
752250active
vasu-devs/JustHireMe
JustHireMe is a local-first desktop workbench (Tauri frontend, Python backend) that scrapes job postings, ranks role fit against your profi…
782236stable
apconw/Aix-DB
Aix-DB is an AI-powered data analysis system (ChatBI) built on LangChain/LangGraph with an MCP Skills multi-agent architecture, converting …
722234active
microsoft/LLaVA-Med
LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su…
402231active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992203active
lkarlslund/Adalanche
Adalanche is an open-source Active Directory attack graph visualizer and explorer written in Go. It collects data via LDAP and SYSVOL, anal…
672195active
google/trax
Trax is an end-to-end deep learning library built on JAX and TensorFlow that focuses on clear code and speed, developed and maintained by t…
108306maintenance
PKU-YuanGroup/LLaVA-CoT
LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea…
472132active
6551Team/opennews-mcp
An MCP (Model Context Protocol) server that aggregates real-time news and market data from 85+ sources across news, exchange listings, on-c…
592125active
coolight7/musicxx
Musicxx (拟声) is a cross-platform audio and video player supporting local files, cloud drives (Baidu, Aliyun, 115), WebDAV, and NAS media se…
712100active
walterlow/freecut
FreeCut is a professional-grade, browser-based multi-track video editor requiring no installation or uploads, with all media and project fi…
602096active
raphaelmansuy/edgequake
EdgeQuake is a high-performance Graph-RAG framework written in Rust, inspired by LightRAG, that transforms documents (PDFs, markdown, text)…
792078active
intel/openvino-plugins-ai-audacity
A set of AI-enabled effects, generators, and analyzers for Audacity, powered by Intel OpenVINO running fully locally on CPU, GPU, or NPU. I…
732066active
WenjieDu/PyPOTS
PyPOTS is a Python toolbox for machine learning and data mining on partially-observed time series with missing values. It integrates 50+ st…
932051active
mli/autocut
AutoCut is a Python CLI tool that automatically transcribes video audio into subtitles using Whisper, then cuts video segments based on whi…
327790maintenance
QwenLM/Qwen3-Embedding
Qwen3-Embedding is a series of text embedding and reranking models (0.6B, 4B, 8B) built on Qwen3 foundation models, with a Python repositor…
382016active
zai-org/GLM-130B
GLM-130B is an open bilingual (English and Chinese) 130-billion-parameter dense language model pre-trained with the General Language Model …
327651maintenance
microsoft/msticpy
msticpy is a Python library from Microsoft for security investigation and threat hunting in Jupyter notebooks. It provides data acquisition…
921995active
showlab/Show-o
Show-o is a research repository implementing a unified transformer model that combines autoregressive and discrete diffusion modeling for m…
501973active
FireRedTeam/FireRedASR
FireRedASR is a family of open-source industrial-grade automatic speech recognition models supporting Mandarin, Chinese dialects, and Engli…
521971active
run-llama/notebookllama
NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval…
521967active
SaiAkhil066/CORTEX-AI-SUPER-RAG
CORTEX RAG is a local-first, agentic retrieval-augmented generation application that lets users upload PDFs and ask questions with cited an…
611962active

← prev page 13 / 16 next →