Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: machine-learning

5378 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
apache/tvm
Apache TVM is an open machine learning compiler framework that takes pre-trained models and compiles them into optimized, deployable module…
9013691active
CompVis/stable-diffusion
The original reference implementation of Stable Diffusion, a latent text-to-image diffusion model trained on LAION-5B data with a CLIP text…
3273347maintenance
Lightning-AI/litgpt
LitGPT is a Python library providing from-scratch, hackable implementations of 20+ open-source large language models with recipes for pretr…
9313629active
divamgupta/diffusionbee-stable-diffusion-ui
DiffusionBee is a free macOS desktop application that runs Stable Diffusion locally with a one-click installer and no technical setup. It p…
2313579active
microsoft/TRELLIS
TRELLIS is Microsoft's large-scale 3D asset generation model that creates high-quality 3D assets from text or image prompts. It uses a unif…
6113510active
Physical-Intelligence/openpi
Open-source repository from Physical Intelligence containing vision-language-action (VLA) models for robotics, including π₀, π₀-FAST, and π…
6613494active
NVIDIA/TensorRT
NVIDIA TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, comprising an inference compiler that converts train…
9413293stable
QwenLM/Qwen3-TTS
Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,…
4813113active
Vaibhavs10/insanely-fast-whisper
A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A…
4913049active
modelscope/DiffSynth-Studio
DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio …
7913003active
PaddlePaddle/PaddleFormers
PaddleFormers is a Transformers-style library built on PaddlePaddle providing a model zoo of 100+ large language models and vision-language…
9212986active
zai-org/CogVideo
CogVideo/CogVideoX is an open-source family of text-to-video and image-to-video generation models from Zhipu AI (THUDM), with inference and…
4512977active
jacobgil/pytorch-grad-cam
A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A…
7612958active
deepseek-ai/FlashMLA
FlashMLA is DeepSeek's library of optimized CUDA attention kernels implementing Multi-head Latent Attention (MLA), including dense and toke…
6212872active
ShiqiYu/libfacedetection
An open-source C++ library for CNN-based face detection in images, with the model embedded as static C source so it has no external depende…
6312784stable
google-research/vision_transformer
Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp…
7512683stable
sapientinc/HRM
Official PyTorch implementation of the Hierarchical Reasoning Model (HRM), a 27M-parameter recurrent architecture with high-level and low-l…
5212619active
YaoFANGUK/video-subtitle-remover
An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai…
7112553active
bmaltais/kohya_ss
A Gradio-based GUI and CLI wrapper around Kohya's Stable Diffusion training scripts for fine-tuning diffusion image generation models. It s…
9512548active
Tencent-Hunyuan/HunyuanVideo
HunyuanVideo is Tencent's open-source framework for large-scale video generation, providing PyTorch model definitions, pre-trained weights,…
6212476active
ace-step/ACE-Step-1.5
ACE-Step 1.5 is an open-source music generation foundation model combining a language model planner with a Diffusion Transformer to create …
7912421active
axolotl-ai-cloud/axolotl
Axolotl is a free, open-source, config-driven framework for fine-tuning large language models, supporting SFT, preference learning (DPO/KTO…
9512408active
Farama-Foundation/Gymnasium
Gymnasium is a Python library providing a standard API for single-agent reinforcement learning environments, maintained by the Farama Found…
8812408stable
DayBreak-u/chineseocr_lite
An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s…
7012339active
cupy/cupy
CuPy is a NumPy/SciPy-compatible array library for GPU-accelerated computing with Python, running on NVIDIA CUDA or AMD ROCm. It acts as a …
9512278stable
guoyww/AnimateDiff
Official implementation of AnimateDiff, a plug-and-play motion modeling module that turns personalized text-to-image diffusion models (e.g.…
2912227active
xmu-xiaoma666/External-Attention-pytorch
A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame…
6512183active
datalab-to/chandra
Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO…
7112171active
PKU-YuanGroup/Open-Sora-Plan
Open-Sora Plan is an open-source effort to reproduce OpenAI's Sora text-to-video model, providing training and inference code for video gen…
5112155active
FlagOpen/FlagEmbedding
FlagEmbedding is the official Python toolkit for BAAI's BGE family of embedding models and rerankers, covering inference, evaluation, and f…
8712086active
google/sentencepiece
SentencePiece is a fast, lightweight unsupervised text tokenizer and detokenizer for neural network-based text generation systems, implemen…
8612043stable
Autoware
Autoware is the world's leading open-source, production-ready software stack for autonomous driving, built on ROS 2 and hosted by the Autow…
9312016stable
instantX-research/InstantID
InstantID is a tuning-free, zero-shot identity-preserving image generation method built on diffusion models, generating customized images i…
2611987active
Tongyi-MAI/Z-Image
Z-Image is a 6B-parameter text-to-image generation foundation model family built on a single-stream diffusion transformer, with a distilled…
4611944active
nerfstudio-project/nerfstudio
Nerfstudio is a Python library and CLI toolkit providing a simple, modular API for creating, training, and testing Neural Radiance Fields (…
3811934active
facebookresearch/seamless_communication
A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre…
7111842active
ostris/ai-toolkit
An all-in-one open-source training toolkit for finetuning diffusion models (image and video) on consumer-grade hardware. It supports many r…
7311838active
speechbrain/speechbrain
SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained…
8311785active
ludwig-ai/ludwig
Ludwig is a declarative, low-code deep learning framework for training, fine-tuning, and deploying AI models — from LLMs to tabular, image,…
9911745active
qubvel-org/segmentation_models.pytorch
A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar…
7011706stable
NVIDIA/cosmos
NVIDIA Cosmos is an open platform of omnimodal world foundation models, datasets, and tools for building Physical AI systems such as robots…
7211641active
cleanlab/cleanlab
Cleanlab is a Python library for data-centric AI that automatically detects issues in ML datasets, such as label errors, outliers, duplicat…
6211636stable
milesial/Pytorch-UNet
A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva…
2311613active
facebookresearch/sam3
Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It …
6311487active
rerun-io/rerun
Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo…
9911362active
THU-MIG/yolov10
YOLOv10 is a real-time end-to-end object detection model family that removes NMS post-processing via consistent dual assignments and optimi…
2011336active
kornia/kornia
Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri…
8611327active
salesforce/LAVIS
LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including …
6111262active
facebookresearch/dinov3
Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t…
5911249active
Weights & Biases
Weights & Biases (wandb) is a Python SDK and platform for tracking, visualizing, and managing machine learning experiments, including metri…
9911239active
ultralytics/yolov5
Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif…
6757929maintenance
Numba
Numba is an open-source NumPy-aware JIT compiler that translates a subset of Python and NumPy code into fast machine code using LLVM. It su…
9411129stable
huggingface/tokenizers
Hugging Face Tokenizers is a fast, Rust-based library implementing state-of-the-art tokenization algorithms (BPE, WordPiece, Unigram) with …
9310997stable
ageitgey/face_recognition
A Python library and command-line tool providing a simple API for face detection, facial landmark extraction, and face recognition, built o…
6356684maintenance
NopeCHA
NopeCHA is an AI-powered CAPTCHA solving service distributed as a browser extension (Chrome/Firefox/Edge) plus Python and Node.js client li…
9310968active
thu-ml/tianshou
Tianshou is a modular, high-performance deep reinforcement learning library built on pure PyTorch and Gymnasium. It offers both low-level h…
7210943active
triton-inference-server/server
NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo…
9810939stable
Lightricks/LTX-Video
Official repository for LTX-Video, a DiT-based open-weights video generation model from Lightricks that generates high-fidelity video (up t…
4810907active
google/dopamine
Dopamine is a research framework from Google for fast prototyping of reinforcement learning algorithms, built around a small, easily readab…
5610900active
NaturalNode/natural
Natural is a general natural language processing library for Node.js offering tokenizing, stemming, part-of-speech tagging, sentiment analy…
6410880active
microsoft/TRELLIS.2
TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f…
5710869active
cumulo-autumn/StreamDiffusion
StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti…
1710806active
Tabnine
Tabnine is an AI code completion and coding assistant that provides all-language autocompletion across major IDEs (VS Code, JetBrains, Subl…
5010774active
doccano/doccano
Doccano is an open-source, web-based text annotation tool for building labeled datasets for machine learning. It supports text classificati…
6610758active
karpathy/minbpe
A minimal, clean Python implementation of the byte-level Byte Pair Encoding (BPE) algorithm used for tokenization in modern LLMs like GPT-4…
2510691stable
lucidrains/denoising-diffusion-pytorch
A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d…
8710679active
autogluon/autogluon
AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth…
9010617active
Megvii-BaseDetection/YOLOX
YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo…
3410587stable
facebookresearch/xformers
xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient …
8810542active
bigscience-workshop/petals
Petals is a Python library that lets you run and fine-tune large language models (Llama 3.1, Mixtral, Falcon, BLOOM) on a BitTorrent-style …
2310521active
IDEA-Research/GroundingDINO
Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v…
2110515stable
pyannote/pyannote-audio
pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ…
9510475active
niedev/RTranslator
RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M…
8510355active
zyddnys/manga-image-translator
A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru…
6510345active
vwxyzjn/cleanrl
CleanRL is a deep reinforcement learning library providing high-quality, single-file implementations of algorithms like PPO, DQN, DDPG, TD3…
5810326active
NVIDIA/cutlass
CUTLASS is NVIDIA's collection of CUDA C++ template abstractions and Python DSLs for implementing high-performance GEMM and related linear …
9910317active
lllyasviel/Fooocus
Fooocus is an offline, open-source image generation application built on Stable Diffusion XL with a Gradio interface. It simplifies text-to…
4452550maintenance
open-mmlab/Amphion
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s…
5110271active
OpenBMB/MiniCPM
MiniCPM is a family of small, state-of-the-art on-device language models from OpenBMB, with MiniCPM5-1B being a dense 1B Transformer for lo…
7510250active
Netflix/metaflow
Metaflow is a human-centric Python framework from Netflix for building, managing, and deploying real-life AI/ML and data science systems. I…
9510245stable
xai-org/grok-1
xAI's open release of the Grok-1 open-weights model (314B-parameter Mixture-of-Experts LLM) with JAX example code for loading and running i…
2552189maintenance
OpenGVLab/InternVL
InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to…
3710146active
stanfordnlp/CoreNLP
Stanford CoreNLP is a Java suite of natural language processing tools that annotates raw text with tokenization, POS tags, named entities, …
7410102stable
TencentARC/PhotoMaker
PhotoMaker is a personalized text-to-image generation method that encodes multiple reference face photos into a stacked ID embedding to gen…
2610088stable
freemocap/freemocap
FreeMoCap is a free, open-source, markerless motion capture system that uses ordinary cameras (webcams, GoPros, smartphones) to record and …
9810085active
deepseek-ai/DeepEP
DeepEP is a high-performance GPU communication library for expert parallelism (EP) in MoE training and inference, providing high-throughput…
5710066active
snakers4/silero-vad
Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d…
8610061stable
EpistasisLab/tpot
TPOT (Tree-based Pipeline Optimization Tool) is a Python automated machine learning library that optimizes scikit-learn machine learning pi…
4410052active
m87-labs/moondream
Moondream is an open-weight family of small, efficient vision language models (2B to 9B MoE) that perform image captioning, visual question…
6110014active
opendatalab/PDF-Extract-Kit
PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection…
259993active
yzhao062/pyod
PyOD is the most comprehensive Python library for anomaly detection, offering 60+ detectors across tabular, time series, graph, text, image…
989977stable
sktime/sktime
sktime is a unified Python framework for machine learning with time series, offering 500+ models behind a single scikit-learn-compatible AP…
989965stable
OpenMined/PySyft
PySyft is a Python library that lets data scientists run computations on private data that stays on the data owner's server, with results s…
749957active
OpenRLHF/OpenRLHF
OpenRLHF is a high-performance, production-ready open-source RLHF framework built on Ray + vLLM + DeepSpeed for scalable reinforcement lear…
899956active
facebookresearch/pytorch3d
PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides…
749954active
espnet/espnet
ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance…
889941active
open-mmlab/mmsegmentation
MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat…
239930stable
huggingface/accelerate
Hugging Face Accelerate is a Python library that lets you run the same PyTorch training and inference code on any device or distributed con…
959838stable
pycaret/pycaret
PyCaret is an open-source, low-code AutoML library for Python that wraps scikit-learn to automate training, tuning, and comparison of model…
839834active
gorse-io/gorse
Gorse is an AI-powered open-source recommender system engine written in Go that ingests items, users, and interaction feedback and automati…
979808active

← prev page 3 / 54 next →