function: machine-learning
5378 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Mega4alik/ollm oLLM is a lightweight Python library for large-context LLM inference built on Hugging Face Transformers and PyTorch. It offloads weights an… | 60 | 2788 | active |
| numz/ComfyUI-SeedVR2_VideoUpscaler The official ComfyUI integration of ByteDance's SeedVR2 model for high-quality video and image upscaling, provided as custom nodes. It can … | 54 | 2786 | active |
| huggingface/setfit SetFit is a Python library for efficient, prompt-free few-shot fine-tuning of Sentence Transformers for text classification. It achieves hi… | 64 | 2784 | active |
| rhasspy/piper Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr… | 10 | 11281 | maintenance |
| Geniusay/ChopperBot ChopperBot is a fully automated AI bot that monitors popular live streams across platforms like Douyu, Huya, Bilibili, Douyin, and Twitch, … | 45 | 2779 | active |
| FasterDecoding/Medusa Medusa is a framework that accelerates LLM text generation by adding multiple decoding heads to an existing model, avoiding the need for a … | 18 | 2770 | active |
| ideogram-oss/ideogram4 Ideogram 4 is an open-weight text-to-image foundation model trained from scratch, with inference code and weights released in Python. It fe… | 54 | 2766 | active |
| autodistill/autodistill Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab… | 29 | 2763 | active |
| NVlabs/stylegan2 The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit… | 32 | 11184 | maintenance |
| AutoArk/GPA GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an… | 54 | 2762 | active |
| NVIDIA/FastPhotoStyle FastPhotoStyle is NVIDIA's official PyTorch implementation of the ECCV 2018 paper 'A Closed-form Solution to Photorealistic Image Stylizati… | 23 | 11177 | maintenance |
| Stan Stan is a C++ probabilistic programming library for full Bayesian inference via NUTS/Hamiltonian Monte Carlo, approximate inference via ADV… | 86 | 2761 | stable |
| kha-white/manga-ocr Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to… | 90 | 2758 | stable |
| apple/turicreate Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj… | 10 | 11159 | maintenance |
| echohive42/AI-reads-books-page-by-page A Python script that reads PDF books page by page, using AI to extract knowledge points and generate progressive summaries at configurable … | 61 | 2756 | active |
| huggingface/swift-coreml-diffusers A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff… | 54 | 2756 | active |
| voxelmorph/voxelmorph VoxelMorph is a Python library for learning-based image registration and alignment, using unsupervised deep learning to model deformations … | 76 | 2748 | active |
| pyro-ppl/numpyro NumPyro is a lightweight probabilistic programming library providing a NumPy backend for Pyro, built on JAX for automatic differentiation a… | 89 | 2746 | active |
| artidoro/qlora QLoRA is the official implementation of the QLoRA paper, an efficient finetuning approach that backpropagates through a frozen 4-bit quanti… | 29 | 10998 | maintenance |
| Bytez-com/docs Bytez is a serverless model inference platform offering a unified API with a single key to run 175k+ open and closed-source AI models (LLMs… | 56 | 2718 | active |
| prophesier/diff-svc Diff-SVC is a deep learning project that performs singing voice conversion using diffusion models, transforming input singing audio into a … | 62 | 2717 | active |
| lengstrom/fast-style-transfer A TensorFlow implementation of fast neural style transfer that applies the style of famous paintings to photos and videos in real time. It … | 32 | 10962 | maintenance |
| mljs/ml ml.js is an umbrella library bundling the mljs organization's machine learning tools for JavaScript, including clustering, classification, … | 23 | 2715 | active |
| vllm-project/vllm-ascend vllm-ascend is a community-maintained hardware plugin that enables vLLM to run large language model inference on Huawei Ascend NPUs. It imp… | 83 | 2711 | active |
| bmild/nerf The official TensorFlow implementation of NeRF (Neural Radiance Fields), the ECCV 2020 paper representing scenes as neural radiance fields … | 39 | 10927 | maintenance |
| intel/neural-compressor Intel Neural Compressor is an open-source Python library providing state-of-the-art low-bit quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4),… | 92 | 2704 | active |
| yuweihao/MambaOut MambaOut is a PyTorch implementation of Gated CNN models from the CVPR 2025 paper 'MambaOut: Do We Really Need Mamba for Vision?', which qu… | 19 | 2704 | stable |
| magic-research/magic-animate MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image … | 44 | 10897 | maintenance |
| dscripka/openWakeWord openWakeWord is an open-source Python library for detecting wake words (or phrases) in streaming audio, with pre-trained models for common … | 49 | 2702 | active |
| TMElyralab/MusePose MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re… | 28 | 2701 | active |
| IDEA-Research/T-Rex T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete… | 48 | 2699 | active |
| facebookresearch/HyperAgents A research framework from Meta AI implementing self-referential, self-improving LLM agents that can optimize their own code for arbitrary c… | 57 | 2697 | active |
| SkyworkAI/SkyReels-V1 SkyReels V1 is an open-source human-centric video foundation model with Text-to-Video and Image-to-Video variants, fine-tuned from HunyuanV… | 25 | 2696 | active |
| secretflow/secretflow SecretFlow is a unified Python framework for privacy-preserving data analysis and machine learning. It layers cryptographic devices (MPC, H… | 68 | 2695 | active |
| roboflow/maestro maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2… | 62 | 2694 | active |
| naiveHobo/InvoiceNet InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN… | 32 | 2694 | active |
| intro-skipper/intro-skipper A Jellyfin server plugin that analyzes the audio of TV episodes to automatically detect intro and credit sequences. It provides media segme… | 94 | 2692 | active |
| baaivision/EVA EVA is a family of large-scale vision foundation models from BAAI, including masked image models (EVA-01/02) and scaled CLIP models (EVA-CL… | 23 | 2691 | active |
| openai/DALL-E The official PyTorch package for the discrete VAE (dVAE) component of OpenAI's DALL·E model. It does not include the transformer that gener… | 10 | 10834 | maintenance |
| qualcomm/aimet AIMET (AI Model Efficiency Toolkit) is a Python library from Qualcomm providing advanced quantization and compression techniques for traine… | 99 | 2688 | active |
| torinmb/mediapipe-touchdesigner A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac… | 86 | 2686 | active |
| KimMeen/Time-LLM Time-LLM is the official PyTorch implementation of an ICLR 2024 paper that reprograms frozen large language models (Llama, GPT-2, BERT) for… | 47 | 2685 | active |
| bytedance/InfiniteYou InfiniteYou (InfU) is a research framework from ByteDance for identity-preserved text-to-image generation built on Diffusion Transformers l… | 37 | 2685 | active |
| yuqinie98/PatchTST Official PyTorch implementation of PatchTST, an ICLR 2023 Transformer model for long-term time series forecasting based on patching and cha… | 32 | 2685 | stable |
| thewh1teagle/kokoro-onnx A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T… | 67 | 2680 | active |
| MrGiovanni/UNetPlusPlus Official implementation of UNet++, a nested U-Net architecture for medical image segmentation, in both Keras and PyTorch. It redesigns skip… | 77 | 2679 | stable |
| HiLab-git/SSL4MIS A benchmark and code collection of semi-supervised learning methods for medical image segmentation, re-implementing approaches like Mean Te… | 44 | 2676 | active |
| tryolabs/norfair Norfair is a lightweight, customizable Python library for real-time multi-object tracking that works with any detector outputting (x, y) co… | 31 | 2676 | stable |
| PowerHouseMan/ComfyUI-AdvancedLivePortrait A ComfyUI custom node implementing LivePortrait for fast facial expression editing and animation with real-time preview. It can edit expres… | 23 | 2675 | active |
| stochasticai/xTuring xTuring is a Python library for fine-tuning, evaluating, and running open-source large language models such as LLaMA, GPT-J, GPT-2, Qwen, a… | 52 | 2674 | active |
| JIA-Lab-research/LISA LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati… | 31 | 2674 | active |
| georgia-tech-db/evadb EvaDB is a Python database system for building AI-powered applications, exposing a SQL API that queries structured data, videos, and unstru… | 10 | 2673 | active |
| ZHZisZZ/dllm dLLM is a Python library that unifies training, inference, and evaluation of diffusion language models such as LLaDA and Dream. It builds o… | 59 | 2672 | active |
| princeton-vl/DROID-SLAM DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense … | 41 | 2671 | active |
| openai/privacy-filter OpenAI Privacy Filter is a small bidirectional token-classification model (1.5B total, 50M active parameters) for detecting and masking per… | 49 | 2670 | active |
| IceClear/StableSR StableSR is a Python research library that leverages pre-trained Stable Diffusion priors for real-world blind image super-resolution. It pr… | 21 | 2668 | stable |
| aigc3d/LHM LHM is a PyTorch-based large reconstruction model that reconstructs high-fidelity animatable 3D human avatars from a single image in second… | 52 | 2664 | active |
| yxlllc/DDSP-SVC DDSP-SVC is an open-source singing voice conversion system built on Differentiable Digital Signal Processing, designed as a free AI voice c… | 65 | 2656 | active |
| Nutlope/roomGPT RoomGPT is an open-source Next.js web application that lets users upload a photo of a room and generate redesigned variations using the Con… | 31 | 10671 | maintenance |
| google-deepmind/mctx Mctx is a JAX-native Python library implementing Monte Carlo tree search algorithms such as AlphaZero, MuZero, and Gumbel MuZero. It suppor… | 83 | 2654 | active |
| icereed/paperless-gpt A self-hosted companion application for paperless-ngx that uses LLMs and vision models to auto-generate document titles, tags, and dates, a… | 87 | 2651 | active |
| hgmzhn/manga-translator-ui A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,… | 80 | 2651 | active |
| phillipi/pix2pix The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from… | 32 | 10652 | maintenance |
| Tencent/MimicMotion MimicMotion is a diffusion-based framework from Tencent for generating high-quality human motion videos guided by pose sequences, featuring… | 47 | 2647 | active |
| SeldonIO/alibi Alibi is a Python library providing algorithms for explaining and interpreting machine learning models, including black-box, white-box, loc… | 44 | 2644 | active |
| facebookresearch/ParlAI ParlAI is a Python framework from Facebook AI Research for sharing, training, and evaluating dialogue models across many openly available d… | 10 | 10621 | maintenance |
| black-forest-labs/flux2 Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal … | 48 | 2642 | active |
| haoheliu/AudioLDM2 AudioLDM 2 is a Python library and CLI for generating audio, music, and speech from text prompts using latent diffusion models. It includes… | 28 | 2639 | active |
| ultralytics/yolov3 Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation… | 67 | 10596 | maintenance |
| YanjieZe/GMR GMR is a Python library for general motion retargeting that converts human motions into joint commands for diverse humanoid robots in real … | 52 | 2633 | active |
| anliyuan/Ultralight-Digital-Human An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It… | 64 | 2627 | active |
| lucidrains/audiolm-pytorch A PyTorch implementation of AudioLM, Google Research's language modeling approach to audio generation, including a MIT-licensed SoundStream… | 34 | 2627 | active |
| swz30/Restormer Restormer is an efficient Transformer architecture for high-resolution image restoration, published as a CVPR 2022 Oral paper. It provides … | 44 | 2625 | stable |
| sergree/matchering Matchering 2.0 is an open-source Python library, containerized web app, and ComfyUI node for audio matching and mastering. It takes a targe… | 64 | 2615 | active |
| plexe-ai/plexe Plexe is a Python library and CLI that builds machine learning models from natural language descriptions using a multi-agent AI workflow. Y… | 61 | 2614 | active |
| open-gigaai/giga-brain-0 GigaBrain-0/0.7 is an open-source vision-language-action (VLA) model family for generalist embodied agents, powered by world models and a t… | 62 | 2611 | active |
| crowsonkb/k-diffusion A PyTorch library implementing Karras et al. (2022) diffusion models with enhancements like improved sampling algorithms and transformer-ba… | 53 | 2600 | active |
| meta-pytorch/torchrec TorchRec is a PyTorch domain library for building recommendation systems at scale. It provides distributed sharding of large embedding tabl… | 90 | 2599 | active |
| kairos-agi/kairos Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f… | 57 | 2598 | active |
| luca-medeiros/lang-segment-anything A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie… | 42 | 2598 | active |
| kijai/ComfyUI-HunyuanVideoWrapper A set of custom ComfyUI nodes that wrap Tencent's HunyuanVideo text-to-video and image-to-video diffusion model for use inside ComfyUI work… | 38 | 2597 | active |
| Slicer/Slicer 3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi… | 67 | 2595 | stable |
| dreamzero0/dreamzero DreamZero is NVIDIA's World Action Model (WAM) that jointly predicts future video and actions from a pretrained video diffusion backbone, e… | 50 | 2593 | active |
| OmniSVG/OmniSVG OmniSVG is a family of end-to-end multimodal SVG generation models built on pre-trained Vision-Language Models, released with inference cod… | 51 | 2590 | active |
| Mininglamp-AI/Mano-P Mano-P is an open-source GUI-VLA (vision-language-action) agent model and SDK for edge devices, enabling purely vision-driven cross-platfor… | 55 | 2589 | active |
| facebookresearch/demucs Demucs is a state-of-the-art music source separation model from Meta AI that splits songs into stems like drums, bass, and vocals using a h… | 10 | 10359 | maintenance |
| median-research-group/LibMTL LibMTL is an open-source PyTorch library for Multi-Task Learning (MTL). It provides implementations of many MTL architectures and gradient-… | 42 | 2586 | active |
| CodeGeeX CodeGeeX is a family of open multilingual code generation large language models (13B and successors CodeGeeX2/CodeGeeX4) pre-trained on 20+… | 23 | 2585 | active |
| asteroid-team/asteroid Asteroid is a PyTorch-based audio source separation toolkit for researchers, providing modular building blocks (filterbanks, encoders, mask… | 60 | 2584 | active |
| ARISE-Initiative/robosuite robosuite is a modular simulation framework powered by the MuJoCo physics engine for robot learning, offering standardized benchmark enviro… | 72 | 2581 | active |
| googlecolab/colabtools The official Python libraries shipped inside Google Colaboratory, Google's free hosted Jupyter notebook environment for machine learning ed… | 77 | 2580 | active |
| apache/hamilton Apache Hamilton is a lightweight Python library for defining directed acyclic graphs (DAGs) of data transformations as plain, testable Pyth… | 91 | 2575 | active |
| atong01/conditional-flow-matching TorchCFM is a PyTorch library implementing Conditional Flow Matching (CFM), a simulation-free training objective for continuous normalizing… | 74 | 2571 | active |
| Tencent-Hunyuan/HY-World-2.0 HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from … | 58 | 2571 | active |
| ymcui/Chinese-BERT-wwm A collection of Chinese pre-trained language models (BERT-wwm, BERT-wwm-ext, RoBERTa-wwm-ext, RBT3, etc.) built with Whole Word Masking, re… | 67 | 10223 | maintenance |
| advimman/lama LaMa is a PyTorch-based image inpainting model that fills large missing regions in images using fast Fourier convolutions, generalizing wel… | 34 | 10217 | maintenance |
| kayba-ai/agentic-context-engine Agentic Context Engine (ACE) is a Python library that lets LLM-based agents learn from experience by maintaining and evolving an evolving c… | 75 | 2558 | active |
| KomputeProject/kompute Kompute is a general-purpose GPU compute framework built on Vulkan that works across vendor GPUs (AMD, NVIDIA, Qualcomm, etc.) with both C+… | 66 | 2558 | active |
| vita-epfl/Stable-Video-Infinity Stable Video Infinity (SVI) is a research codebase for infinite-length video generation using video diffusion transformers with an error-re… | 55 | 2556 | active |
| graphistry/pygraphistry PyGraphistry is a Python library for loading, shaping, and visually exploring large graphs with GPU-accelerated rendering and analytics via… | 97 | 2551 | active |