function: machine-learning
5378 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| apache/tvm Apache TVM is an open machine learning compiler framework that takes pre-trained models and compiles them into optimized, deployable module… | 90 | 13691 | active |
| CompVis/stable-diffusion The original reference implementation of Stable Diffusion, a latent text-to-image diffusion model trained on LAION-5B data with a CLIP text… | 32 | 73347 | maintenance |
| Lightning-AI/litgpt LitGPT is a Python library providing from-scratch, hackable implementations of 20+ open-source large language models with recipes for pretr… | 93 | 13629 | active |
| divamgupta/diffusionbee-stable-diffusion-ui DiffusionBee is a free macOS desktop application that runs Stable Diffusion locally with a one-click installer and no technical setup. It p… | 23 | 13579 | active |
| microsoft/TRELLIS TRELLIS is Microsoft's large-scale 3D asset generation model that creates high-quality 3D assets from text or image prompts. It uses a unif… | 61 | 13510 | active |
| Physical-Intelligence/openpi Open-source repository from Physical Intelligence containing vision-language-action (VLA) models for robotics, including π₀, π₀-FAST, and π… | 66 | 13494 | active |
| NVIDIA/TensorRT NVIDIA TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, comprising an inference compiler that converts train… | 94 | 13293 | stable |
| QwenLM/Qwen3-TTS Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,… | 48 | 13113 | active |
| Vaibhavs10/insanely-fast-whisper A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A… | 49 | 13049 | active |
| modelscope/DiffSynth-Studio DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio … | 79 | 13003 | active |
| PaddlePaddle/PaddleFormers PaddleFormers is a Transformers-style library built on PaddlePaddle providing a model zoo of 100+ large language models and vision-language… | 92 | 12986 | active |
| zai-org/CogVideo CogVideo/CogVideoX is an open-source family of text-to-video and image-to-video generation models from Zhipu AI (THUDM), with inference and… | 45 | 12977 | active |
| jacobgil/pytorch-grad-cam A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A… | 76 | 12958 | active |
| deepseek-ai/FlashMLA FlashMLA is DeepSeek's library of optimized CUDA attention kernels implementing Multi-head Latent Attention (MLA), including dense and toke… | 62 | 12872 | active |
| ShiqiYu/libfacedetection An open-source C++ library for CNN-based face detection in images, with the model embedded as static C source so it has no external depende… | 63 | 12784 | stable |
| google-research/vision_transformer Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp… | 75 | 12683 | stable |
| sapientinc/HRM Official PyTorch implementation of the Hierarchical Reasoning Model (HRM), a 27M-parameter recurrent architecture with high-level and low-l… | 52 | 12619 | active |
| YaoFANGUK/video-subtitle-remover An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai… | 71 | 12553 | active |
| bmaltais/kohya_ss A Gradio-based GUI and CLI wrapper around Kohya's Stable Diffusion training scripts for fine-tuning diffusion image generation models. It s… | 95 | 12548 | active |
| Tencent-Hunyuan/HunyuanVideo HunyuanVideo is Tencent's open-source framework for large-scale video generation, providing PyTorch model definitions, pre-trained weights,… | 62 | 12476 | active |
| ace-step/ACE-Step-1.5 ACE-Step 1.5 is an open-source music generation foundation model combining a language model planner with a Diffusion Transformer to create … | 79 | 12421 | active |
| axolotl-ai-cloud/axolotl Axolotl is a free, open-source, config-driven framework for fine-tuning large language models, supporting SFT, preference learning (DPO/KTO… | 95 | 12408 | active |
| Farama-Foundation/Gymnasium Gymnasium is a Python library providing a standard API for single-agent reinforcement learning environments, maintained by the Farama Found… | 88 | 12408 | stable |
| DayBreak-u/chineseocr_lite An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s… | 70 | 12339 | active |
| cupy/cupy CuPy is a NumPy/SciPy-compatible array library for GPU-accelerated computing with Python, running on NVIDIA CUDA or AMD ROCm. It acts as a … | 95 | 12278 | stable |
| guoyww/AnimateDiff Official implementation of AnimateDiff, a plug-and-play motion modeling module that turns personalized text-to-image diffusion models (e.g.… | 29 | 12227 | active |
| xmu-xiaoma666/External-Attention-pytorch A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame… | 65 | 12183 | active |
| datalab-to/chandra Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO… | 71 | 12171 | active |
| PKU-YuanGroup/Open-Sora-Plan Open-Sora Plan is an open-source effort to reproduce OpenAI's Sora text-to-video model, providing training and inference code for video gen… | 51 | 12155 | active |
| FlagOpen/FlagEmbedding FlagEmbedding is the official Python toolkit for BAAI's BGE family of embedding models and rerankers, covering inference, evaluation, and f… | 87 | 12086 | active |
| google/sentencepiece SentencePiece is a fast, lightweight unsupervised text tokenizer and detokenizer for neural network-based text generation systems, implemen… | 86 | 12043 | stable |
| Autoware Autoware is the world's leading open-source, production-ready software stack for autonomous driving, built on ROS 2 and hosted by the Autow… | 93 | 12016 | stable |
| instantX-research/InstantID InstantID is a tuning-free, zero-shot identity-preserving image generation method built on diffusion models, generating customized images i… | 26 | 11987 | active |
| Tongyi-MAI/Z-Image Z-Image is a 6B-parameter text-to-image generation foundation model family built on a single-stream diffusion transformer, with a distilled… | 46 | 11944 | active |
| nerfstudio-project/nerfstudio Nerfstudio is a Python library and CLI toolkit providing a simple, modular API for creating, training, and testing Neural Radiance Fields (… | 38 | 11934 | active |
| facebookresearch/seamless_communication A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre… | 71 | 11842 | active |
| ostris/ai-toolkit An all-in-one open-source training toolkit for finetuning diffusion models (image and video) on consumer-grade hardware. It supports many r… | 73 | 11838 | active |
| speechbrain/speechbrain SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained… | 83 | 11785 | active |
| ludwig-ai/ludwig Ludwig is a declarative, low-code deep learning framework for training, fine-tuning, and deploying AI models — from LLMs to tabular, image,… | 99 | 11745 | active |
| qubvel-org/segmentation_models.pytorch A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar… | 70 | 11706 | stable |
| NVIDIA/cosmos NVIDIA Cosmos is an open platform of omnimodal world foundation models, datasets, and tools for building Physical AI systems such as robots… | 72 | 11641 | active |
| cleanlab/cleanlab Cleanlab is a Python library for data-centric AI that automatically detects issues in ML datasets, such as label errors, outliers, duplicat… | 62 | 11636 | stable |
| milesial/Pytorch-UNet A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva… | 23 | 11613 | active |
| facebookresearch/sam3 Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It … | 63 | 11487 | active |
| rerun-io/rerun Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo… | 99 | 11362 | active |
| THU-MIG/yolov10 YOLOv10 is a real-time end-to-end object detection model family that removes NMS post-processing via consistent dual assignments and optimi… | 20 | 11336 | active |
| kornia/kornia Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri… | 86 | 11327 | active |
| salesforce/LAVIS LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including … | 61 | 11262 | active |
| facebookresearch/dinov3 Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t… | 59 | 11249 | active |
| Weights & Biases Weights & Biases (wandb) is a Python SDK and platform for tracking, visualizing, and managing machine learning experiments, including metri… | 99 | 11239 | active |
| ultralytics/yolov5 Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif… | 67 | 57929 | maintenance |
| Numba Numba is an open-source NumPy-aware JIT compiler that translates a subset of Python and NumPy code into fast machine code using LLVM. It su… | 94 | 11129 | stable |
| huggingface/tokenizers Hugging Face Tokenizers is a fast, Rust-based library implementing state-of-the-art tokenization algorithms (BPE, WordPiece, Unigram) with … | 93 | 10997 | stable |
| ageitgey/face_recognition A Python library and command-line tool providing a simple API for face detection, facial landmark extraction, and face recognition, built o… | 63 | 56684 | maintenance |
| NopeCHA NopeCHA is an AI-powered CAPTCHA solving service distributed as a browser extension (Chrome/Firefox/Edge) plus Python and Node.js client li… | 93 | 10968 | active |
| thu-ml/tianshou Tianshou is a modular, high-performance deep reinforcement learning library built on pure PyTorch and Gymnasium. It offers both low-level h… | 72 | 10943 | active |
| triton-inference-server/server NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo… | 98 | 10939 | stable |
| Lightricks/LTX-Video Official repository for LTX-Video, a DiT-based open-weights video generation model from Lightricks that generates high-fidelity video (up t… | 48 | 10907 | active |
| google/dopamine Dopamine is a research framework from Google for fast prototyping of reinforcement learning algorithms, built around a small, easily readab… | 56 | 10900 | active |
| NaturalNode/natural Natural is a general natural language processing library for Node.js offering tokenizing, stemming, part-of-speech tagging, sentiment analy… | 64 | 10880 | active |
| microsoft/TRELLIS.2 TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f… | 57 | 10869 | active |
| cumulo-autumn/StreamDiffusion StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti… | 17 | 10806 | active |
| Tabnine Tabnine is an AI code completion and coding assistant that provides all-language autocompletion across major IDEs (VS Code, JetBrains, Subl… | 50 | 10774 | active |
| doccano/doccano Doccano is an open-source, web-based text annotation tool for building labeled datasets for machine learning. It supports text classificati… | 66 | 10758 | active |
| karpathy/minbpe A minimal, clean Python implementation of the byte-level Byte Pair Encoding (BPE) algorithm used for tokenization in modern LLMs like GPT-4… | 25 | 10691 | stable |
| lucidrains/denoising-diffusion-pytorch A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d… | 87 | 10679 | active |
| autogluon/autogluon AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth… | 90 | 10617 | active |
| Megvii-BaseDetection/YOLOX YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo… | 34 | 10587 | stable |
| facebookresearch/xformers xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient … | 88 | 10542 | active |
| bigscience-workshop/petals Petals is a Python library that lets you run and fine-tune large language models (Llama 3.1, Mixtral, Falcon, BLOOM) on a BitTorrent-style … | 23 | 10521 | active |
| IDEA-Research/GroundingDINO Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v… | 21 | 10515 | stable |
| pyannote/pyannote-audio pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ… | 95 | 10475 | active |
| niedev/RTranslator RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M… | 85 | 10355 | active |
| zyddnys/manga-image-translator A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru… | 65 | 10345 | active |
| vwxyzjn/cleanrl CleanRL is a deep reinforcement learning library providing high-quality, single-file implementations of algorithms like PPO, DQN, DDPG, TD3… | 58 | 10326 | active |
| NVIDIA/cutlass CUTLASS is NVIDIA's collection of CUDA C++ template abstractions and Python DSLs for implementing high-performance GEMM and related linear … | 99 | 10317 | active |
| lllyasviel/Fooocus Fooocus is an offline, open-source image generation application built on Stable Diffusion XL with a Gradio interface. It simplifies text-to… | 44 | 52550 | maintenance |
| open-mmlab/Amphion Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s… | 51 | 10271 | active |
| OpenBMB/MiniCPM MiniCPM is a family of small, state-of-the-art on-device language models from OpenBMB, with MiniCPM5-1B being a dense 1B Transformer for lo… | 75 | 10250 | active |
| Netflix/metaflow Metaflow is a human-centric Python framework from Netflix for building, managing, and deploying real-life AI/ML and data science systems. I… | 95 | 10245 | stable |
| xai-org/grok-1 xAI's open release of the Grok-1 open-weights model (314B-parameter Mixture-of-Experts LLM) with JAX example code for loading and running i… | 25 | 52189 | maintenance |
| OpenGVLab/InternVL InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to… | 37 | 10146 | active |
| stanfordnlp/CoreNLP Stanford CoreNLP is a Java suite of natural language processing tools that annotates raw text with tokenization, POS tags, named entities, … | 74 | 10102 | stable |
| TencentARC/PhotoMaker PhotoMaker is a personalized text-to-image generation method that encodes multiple reference face photos into a stacked ID embedding to gen… | 26 | 10088 | stable |
| freemocap/freemocap FreeMoCap is a free, open-source, markerless motion capture system that uses ordinary cameras (webcams, GoPros, smartphones) to record and … | 98 | 10085 | active |
| deepseek-ai/DeepEP DeepEP is a high-performance GPU communication library for expert parallelism (EP) in MoE training and inference, providing high-throughput… | 57 | 10066 | active |
| snakers4/silero-vad Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d… | 86 | 10061 | stable |
| EpistasisLab/tpot TPOT (Tree-based Pipeline Optimization Tool) is a Python automated machine learning library that optimizes scikit-learn machine learning pi… | 44 | 10052 | active |
| m87-labs/moondream Moondream is an open-weight family of small, efficient vision language models (2B to 9B MoE) that perform image captioning, visual question… | 61 | 10014 | active |
| opendatalab/PDF-Extract-Kit PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection… | 25 | 9993 | active |
| yzhao062/pyod PyOD is the most comprehensive Python library for anomaly detection, offering 60+ detectors across tabular, time series, graph, text, image… | 98 | 9977 | stable |
| sktime/sktime sktime is a unified Python framework for machine learning with time series, offering 500+ models behind a single scikit-learn-compatible AP… | 98 | 9965 | stable |
| OpenMined/PySyft PySyft is a Python library that lets data scientists run computations on private data that stays on the data owner's server, with results s… | 74 | 9957 | active |
| OpenRLHF/OpenRLHF OpenRLHF is a high-performance, production-ready open-source RLHF framework built on Ray + vLLM + DeepSpeed for scalable reinforcement lear… | 89 | 9956 | active |
| facebookresearch/pytorch3d PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides… | 74 | 9954 | active |
| espnet/espnet ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance… | 88 | 9941 | active |
| open-mmlab/mmsegmentation MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat… | 23 | 9930 | stable |
| huggingface/accelerate Hugging Face Accelerate is a Python library that lets you run the same PyTorch training and inference code on any device or distributed con… | 95 | 9838 | stable |
| pycaret/pycaret PyCaret is an open-source, low-code AutoML library for Python that wraps scikit-learn to automate training, tuning, and comparison of model… | 83 | 9834 | active |
| gorse-io/gorse Gorse is an AI-powered open-source recommender system engine written in Go that ingests items, users, and interaction feedback and automati… | 97 | 9808 | active |