Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: machine-learning

5378 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ianarawjo/ChainForge
ChainForge is an open-source visual programming environment for battle-testing prompts to LLMs, built on ReactFlow and Flask. It lets users…
733026active
deepnote/deepnote
Deepnote is an open-source, AI-first drop-in replacement for Jupyter notebooks with a human-readable YAML format, block-based architecture …
793000active
NVIDIA/NeMo-Retriever
NVIDIA's scalable document content and metadata extraction library (also known as NVIDIA Ingest) that splits documents, classifies and extr…
862973active
signerless/llm-checker
A Node.js CLI tool that scans your hardware (GPU, VRAM, CPU, memory) and recommends which LLM or small language models you can run locally,…
842951active
KevinWang676/Bark-Voice-Cloning
A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S…
672947active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
532934active
michaelfeil/infinity
Infinity is a high-throughput, low-latency serving engine that exposes text-embedding, reranking, CLIP, CLAP and ColPali models via an Open…
632933active
ax-llm/ax
Ax is a TypeScript-first LLM programming framework implementing the DSPy model, providing typed signatures, agents, flows, and optimizers l…
942891active
jhj0517/Whisper-WebUI
A Gradio-based web interface for OpenAI's Whisper models that generates subtitles from files, YouTube videos, or microphone input. It suppo…
632870active
bytedeco/javacpp-presets
JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP…
862850active
whylabs/whylogs
whylogs is an open-source Python data logging library for machine learning models and data pipelines. It profiles datasets with privacy-pre…
342831active
openalpr/openalpr
OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete…
2311456maintenance
google-research/kubric
Kubric is a data generation pipeline from Google Research for creating semi-realistic synthetic multi-object videos with rich annotations l…
602813active
RoboTwin-Platform/RoboTwin
RoboTwin 2.0 is a scalable benchmark and data-generation platform for bimanual (dual-arm) robotic manipulation, built on simulated digital …
692808active
mll-lab-nu/RAGEN
RAGEN is a Python framework for training reasoning LLM agents with multi-turn reinforcement learning using the StarPO algorithm. It also pr…
652789active
ModelTC/LightX2V
LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-…
652777active
FluidInference/FluidAudio
A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot…
812727active
CVCUDA/CV-CUDA
CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs…
932723active
freedmand/semantra-python
Semantra is a command-line tool that semantically indexes text and PDF documents and launches a local web application for querying them by …
302711active
xdit-project/xDiT
xDiT is a scalable inference engine for Diffusion Transformers (DiTs) that enables parallel deployment across multiple GPUs and machines. I…
922706active
UKGovernmentBEIS/inspect_ai
Inspect is an open-source Python framework from the UK AI Security Institute for running large language model evaluations, with composable …
722696active
spotify/scio
Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and strea…
962628active
NVlabs/LongLive
LongLive is an NVIDIA research framework providing parallel training and inference infrastructure for real-time long video generation, usin…
602592active
Memento-Teams/Memento
Memento is a Python framework for building LLM agents that continually improve from experience via memory-based case-based reasoning, witho…
382568active
YeQing17-2026/OmniAgent
OmniAgent is an open-source Python agent framework that self-evolves across skills, context, memory, and its underlying model during intera…
572557active
AIGC-Audio/AudioGPT
AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP…
3010168maintenance
microsoft/foundry-local
Foundry Local is Microsoft's end-to-end local AI runtime and SDK suite (C#, JavaScript, Python, Rust) for running optimized models entirely…
872538active
facebookresearch/nougat
Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py…
2310071maintenance
janhq/ichigo
Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to…
492489active
xuebinqin/U-2-Net
Official PyTorch implementation of U^2-Net, a nested U-structure deep network for salient object detection, published in Pattern Recognitio…
329856maintenance
microsoft/PIKE-RAG
PIKE-RAG is a Microsoft framework for building Retrieval-Augmented Generation systems that extract, understand, and apply specialized domai…
312481active
mirage-project/mirage
Mirage Persistent Kernel (MPK) is a compiler and runtime that transforms multi-GPU LLM inference into a single fused megakernel, reducing i…
832474active
data-infra/cube-studio
CubeStudio is an open-source, cloud-native, all-in-one AI platform covering the full machine learning lifecycle (MLOps/MaaS/LLMOps), includ…
792473active
Zleap-AI/SAG
SAG is an open-source retrieval architecture and knowledge base application that replaces both traditional RAG and GraphRAG with event-enti…
842460active
nz-m/SocialEcho
SocialEcho is a full-featured social networking platform built on the MERN stack (MongoDB, Express.js, React.js, Node.js) with automated co…
312435active
snapotter-hq/SnapOtter
SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities…
802412active
ailia-ai/ailia-models
A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,…
772389active
showlab/Paper2Video
Paper2Video is a Python pipeline that automatically generates academic presentation videos from scientific papers, taking a paper PDF, a sp…
492373active
eikek/docspell
Docspell is a self-hosted document management system (DMS) for organizing scanned papers, e-mails, and other files. It automates metadata t…
682317active
claimed-framework/claimed
CLAIMED (C3) is a component compiler that turns Jupyter notebooks, Python scripts, and R scripts into portable containerized AI components.…
882305active
togethercomputer/OpenChatKit
OpenChatKit is an open-source toolkit from Together, LAION, and Ontocord.ai for building and fine-tuning chat language models, including in…
308983maintenance
casadi/casadi
CasADi is an open-source symbolic framework for gradient-based numerical optimization, implementing forward and reverse mode automatic diff…
912284stable
langfengQ/verl-agent
verl-agent is an extension of the veRL framework for training LLM and VLM agents via reinforcement learning, featuring step-independent mul…
562276active
PrismML-Eng/Bonsai-demo
A demo repository for running PrismML's Bonsai family of 1-bit and ternary-quantized language models locally via llama.cpp and MLX. It prov…
592272active
IBM/AssetOpsBench
AssetOpsBench is an open-source benchmark and framework from IBM for building, orchestrating, and evaluating domain-specific AI agents for …
642256active
0xShug0/audio.cpp
audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and …
812249active
wxyhgk/retain-pdf
RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/…
822238active
DanOps-1/Gpt-Agreement-Payment
A Python toolkit that reverse-engineers and replays the end-to-end ChatGPT Plus/Team/Pro subscription payment flow (Stripe Checkout, PayPal…
532238active
dstackai/dstack
dstack is an open-source, vendor-agnostic control plane for GPU provisioning and orchestration that works across GPU clouds, Kubernetes, an…
962236active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992234active
TommyLemon/APIAuto
APIAuto is a web-based HTTP API tool combining machine-learning-powered no-code automated testing, code generation with static checks, and …
862218active
MVIG-SJTU/AlphaPose
AlphaPose is an open-source real-time multi-person full-body pose estimation and tracking system built on PyTorch. It detects human keypoin…
328600maintenance
unrealcv/unrealcv
UnrealCV is an open-source Unreal Engine plugin that connects computer vision research to virtual worlds by exposing a command API and Pyth…
882209active
TimmyOVO/deepseek-ocr.rs
A Rust implementation of the DeepSeek-OCR inference stack with multiple OCR/VLM backends (DeepSeek-OCR, PaddleOCR-VL, DotsOCR), DSQ quantiz…
602184active
intel/intel-extension-for-transformers
Intel's toolkit for accelerating transformer-based GenAI/LLM workloads on Intel platforms, offering state-of-the-art compression (e.g., INT…
102173active
boson-ai/higgs-audio
Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s…
568339maintenance
HUANGCHIHHUNGLeo/claude-real-video
A Python CLI tool and agent skill that lets LLMs like Claude actually watch videos by extracting scene-aware, deduplicated keyframes plus a…
802116active
GAIR-NLP/daVinci-MagiHuman
daVinci-MagiHuman is an open-source 15B-parameter single-stream transformer foundation model that jointly generates synchronized audio and …
492115active
A-EDev/Flow
Flow is an open-source, privacy-respecting YouTube and YouTube Music client for Android built with Kotlin and Jetpack Compose (Material 3).…
842108active
LeCAR-Lab/ASAP
ASAP is a two-stage framework for training agile humanoid whole-body skills by aligning simulation and real-world physics. It pre-trains mo…
472107active
QwenLM/Qwen2-Audio
Official repository for Qwen2-Audio, a 7B-parameter large audio-language model from Alibaba Cloud that accepts audio inputs and responds to…
312099active
raphaelmansuy/edgequake
EdgeQuake is a high-performance Graph-RAG framework written in Rust, inspired by LightRAG, that transforms documents (PDFs, markdown, text)…
792084active
visomaster/VisoMaster
VisoMaster is a Python-based desktop application for AI-powered face swapping and face editing in images and videos. It supports multiple s…
272068active
Plachtaa/VALL-E-X
An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint…
107930maintenance
gocrane/crane
Crane is a FinOps platform for cloud resource analytics and cost optimization in Kubernetes clusters. It provides cost insight, optimizatio…
232054active
microsoft/foundry-dev-tools
Microsoft Foundry Toolkit (formerly AI Toolkit) is a Visual Studio Code extension for building, testing, and deploying AI agents and models…
802050active
ainfosec/FISSURE
FISSURE is an open-source RF and reverse engineering framework built around software-defined radios, supporting signal detection, classific…
772039active
HKUDS/MiniRAG
MiniRAG is an extremely simple retrieval-augmented generation framework designed to work with small, open-source language models. It uses s…
392008active
praat/praat.github.io
Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro…
991971stable
showlab/computer_use_ootb
An out-of-the-box desktop GUI agent that lets vision-language models like Claude 3.5 Computer Use, ShowUI, and UI-TARS control Windows and …
321958active
chn-lee-yumi/MaterialSearch
MaterialSearch is a self-hosted semantic search tool that indexes local photos and videos using a CLIP multimodal model, letting users find…
731957active
kenforthewin/atomic
Atomic is a self-hosted, local-first personal knowledge base that turns markdown notes into a semantically-connected, AI-augmented knowledg…
771945active
google-deepmind/lab
DeepMind Lab is a customisable 3D learning environment built on Quake III Arena (ioquake3) that provides navigation and puzzle-solving task…
237372maintenance
LCAV/pyroomacoustics
Pyroomacoustics is a Python package for audio signal processing in indoor scenarios, combining a fast C++ room acoustics simulator (image s…
901930stable
chthollyphile/folia-major
Folia is an online music player focused on immersive full-screen lyrics animations, supporting NetEase Cloud Music, KuGou, Navidrome, and l…
801910active
oficcejo/aiagents-stock
A Python multi-AI-agent stock analysis and monitoring system for mainland China A-shares, simulating a team of securities analysts to produ…
531902active
baetyl/baetyl
Baetyl is an open-source edge computing framework from Linux Foundation Edge that extends cloud computing, data, and services to edge devic…
231900active
AutoArk/EVA-OS
EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining …
681890active
flybirdxx/ComfyUI-Qwen-TTS
A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice…
541884active
MontrealCorpusTools/Montreal-Forced-Aligner
Montreal Forced Aligner is a command line utility for time-aligning orthographic transcriptions and pronunciation dictionary entries to aud…
981876active
we0091234/Chinese_license_plate_detection_recognition
A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports …
701869active
AutoFigure
AutoFigure-Edit is a Python application that converts scientific paper method sections into fully editable SVG figures using large language…
551867active
gaomingqi/Track-Anything
Track-Anything is an interactive tool for video object tracking and segmentation built on Segment Anything, XMem, and E2FGVI. Users specify…
556997maintenance
YangLing0818/RPG-DiffusionMaster
Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and …
271844active
HKUDS/VideoAgent
VideoAgent is an all-in-one agentic framework for video understanding, editing, and remaking, built on multi-agent orchestration with over …
601838active
ROCm/FastFlowLM
FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama…
861836active
undertheseanlp/underthesea
Underthesea is an open-source Python toolkit for Vietnamese natural language processing that has evolved into an agentic AI toolkit with mu…
941803active
dreadl0ck/netcap
Netcap is a Go framework that converts network packets into structured, type-safe Protocol Buffer audit records for security monitoring, fo…
921803active
HybridRobotics/berkeley-humanoid-lite
Berkeley Humanoid Lite is the open-source codebase for a sub-$5,000 3D-printed humanoid robot platform from UC Berkeley. It includes Isaac …
551802active
SmartFlowAI/EmoLLM
EmoLLM is a series of open-source large language models fine-tuned for mental health understanding and support, built on models like Intern…
591781active
neuml/paperai
paperai is an AI application for medical and scientific papers that runs bulk LLM inference and RAG pipelines over article repositories to …
671779active
Emu Series
Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid…
561779active
beam-cloud/beta9
Beam (beta9) is an open-source serverless runtime for AI workloads, providing GPU inference endpoints, isolated sandboxes for running untru…
901766active
signerlabs/Klee
Klee is a native macOS AI chat application that runs large language models entirely on-device using Apple's MLX framework on Apple Silicon.…
551761active
x007xyz/flycut-caption
FlyCut Caption is an AI-powered video subtitle editing tool built as a React component and web/desktop app, offering speech recognition wit…
781757active
0xSero/turboquant
TurboQuant is a Python library implementing near-optimal KV cache quantization for LLM inference, compressing keys to 3-bit and values to 2…
591755active
facebookresearch/metaseq
Metaseq is a PyTorch codebase from Meta AI for training and working with large-scale Open Pre-trained Transformers (OPT), forked from fairs…
106547maintenance
Feather-2/Burner-X
Paper Burner X is a browser-based AI workstation for processing, translating, and analyzing academic documents like PDFs, DOCX, PPTX, and E…
511752active
NVIDIA-NeMo/Curator
NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for …
861751active
lf-edge/ekuiper
LF Edge eKuiper is a lightweight stream processing engine for IoT edge devices, offering SQL-based and graph-based rule engine for real-tim…
991733active

← prev page 51 / 54 next →