function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| huggingface/smollm Hugging Face's repository for the SmolLM and SmolVLM families of compact, fully open language and vision-language models, including trainin… | 59 | 3884 | active |
| Avatarify Avatarify is an open-source application that drives photorealistic avatars in real time for video-conferencing apps like Zoom and Skype, ba… | 23 | 16515 | maintenance |
| IDEA-Research/Grounded-SAM-2 Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X)… | 37 | 3708 | active |
| xinyu1205/recognize-anything Recognize Anything is a collection of open-source image recognition foundation models, including RAM, RAM++, and Tag2Text, that perform ima… | 33 | 3708 | active |
| thu-ml/SageAttention SageAttention is a family of quantized attention kernels (INT8/FP8/FP4) that accelerate transformer inference 2-5x over FlashAttention with… | 40 | 3684 | active |
| zai-org/ChatGLM2-6B ChatGLM2-6B is an open-source bilingual (Chinese-English) 6B-parameter conversational large language model built on the GLM architecture. I… | 29 | 15528 | maintenance |
| ToTheBeginning/PuLID PuLID is the official PyTorch implementation of a NeurIPS 2024 method for inserting a specific person's identity into text-to-image generat… | 40 | 3550 | active |
| ob-f/OpenBot OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo… | 67 | 3442 | active |
| jixiaozhong/Sonic Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation… | 49 | 3273 | active |
| prs-eth/Marigold Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus… | 52 | 3198 | active |
| tekaratzas/RustGPT A transformer-based large language model implemented entirely in pure Rust with no external ML frameworks, using only ndarray for matrix op… | 38 | 3157 | active |
| Djdefrag/QualityScaler QualityScaler is a Windows GUI application that uses AI deep-learning models to upscale, enhance, and de-noise images and videos. It is wri… | 95 | 3138 | active |
| z-x-yang/Segment-and-Track-Anything An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg… | 61 | 3134 | active |
| ridgerchu/matmulfreellm A Python implementation of MatMul-Free LM, a language model architecture that eliminates matrix multiplication operations using ternary wei… | 49 | 3089 | active |
| naver/mast3r MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r… | 37 | 3088 | active |
| Doubiiu/DynamiCrafter DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p… | 27 | 3007 | active |
| williamyang1991/Rerender_A_Video The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.… | 29 | 2999 | stable |
| luminal-ai/luminal Luminal is a high-performance general-purpose ML inference compiler written in Rust that lowers models to a minimal 15-op dataflow IR and c… | 78 | 2956 | active |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2931 | active |
| facebookresearch/omnilingual-asr An open-source multilingual speech recognition library from Meta AI supporting over 1,600 languages, including hundreds never previously co… | 52 | 2898 | active |
| NVlabs/FoundationStereo FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV… | 47 | 2874 | active |
| UX-Decoder/Semantic-SAM Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired … | 33 | 2854 | active |
| TMElyralab/MuseV MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall… | 25 | 2846 | active |
| Camb-ai/MARS5-TTS MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with … | 14 | 2817 | active |
| microsoft/MoGe MoGe is a deep learning model from Microsoft Research that recovers 3D geometry from a single open-domain image, predicting metric point ma… | 66 | 2807 | active |
| autodistill/autodistill Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab… | 29 | 2763 | active |
| kha-white/manga-ocr Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to… | 90 | 2758 | stable |
| ModelTC/LightX2V LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-… | 64 | 2733 | active |
| magic-research/magic-animate MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image … | 44 | 10897 | maintenance |
| yxlllc/DDSP-SVC DDSP-SVC is an open-source singing voice conversion system built on Differentiable Digital Signal Processing, designed as a free AI voice c… | 65 | 2656 | active |
| luca-medeiros/lang-segment-anything A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie… | 42 | 2598 | active |
| facebookresearch/nougat Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py… | 23 | 10063 | maintenance |
| ipazc/mtcnn A Python library implementing the MTCNN (Multitask Cascaded Convolutional Networks) algorithm for face detection and facial landmark alignm… | 23 | 2485 | stable |
| data-infra/cube-studio CubeStudio is an open-source, cloud-native, all-in-one AI platform covering the full machine learning lifecycle (MLOps/MaaS/LLMOps), includ… | 80 | 2448 | active |
| pnnbao97/VieNeu-TTS VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi… | 84 | 2427 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2385 | active |
| jik876/hifi-gan The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k… | 32 | 2367 | stable |
| MouseLand/cellpose Cellpose is a generalist deep learning algorithm for cellular and nucleus segmentation in microscopy images, with human-in-the-loop capabil… | 86 | 2331 | active |
| labmlai/labml A Python library for tracking and monitoring deep learning experiments, with a self-hostable server app for viewing metrics and hardware us… | 30 | 2325 | active |
| XiaomiMiMo/MiMo Xiaomi's MiMo is a 7B-parameter reasoning language model trained from pretraining through posttraining with reinforcement learning, release… | 30 | 2299 | active |
| OlafenwaMoses/ImageAI ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob… | 23 | 8877 | maintenance |
| ermig1979/Simd Simd Library is a free open-source C++ image processing and machine learning library with a C API and Python wrapper. Its algorithms are ha… | 98 | 2265 | active |
| THU-MIG/yoloe YOLOE is the official PyTorch implementation of an open-vocabulary object detection and segmentation model presented at ICCV 2025. It unifi… | 32 | 2256 | active |
| MVIG-SJTU/AlphaPose AlphaPose is an open-source real-time multi-person full-body pose estimation and tracking system built on PyTorch. It detects human keypoin… | 32 | 8596 | maintenance |
| DigitalPhonetics/IMS-Toucan IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive… | 63 | 2207 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| intel/intel-extension-for-transformers Intel's toolkit for accelerating transformer-based GenAI/LLM workloads on Intel platforms, offering state-of-the-art compression (e.g., INT… | 10 | 2174 | active |
| MoonshotAI/MoBA MoBA (Mixture of Block Attention) is a PyTorch implementation of a trainable block-sparse attention mechanism for long-context large langua… | 27 | 2169 | active |
| Tencent-Hunyuan/HunyuanVideo-Avatar HunyuanVideo-Avatar is Tencent's open-source model and inference code for high-fidelity audio-driven human animation, generating talking av… | 45 | 2156 | active |
| espressif/esp-who ESP-WHO is an image processing development platform from Espressif providing face detection, face recognition, pedestrian detection, and QR… | 67 | 2133 | active |
| LiheYoung/Depth-Anything Depth Anything is a monocular depth estimation foundation model trained on 1.5M labeled and 62M+ unlabeled images, released as a Python lib… | 26 | 8195 | maintenance |
| PaddlePaddle/PaddleGAN PaddleGAN is a Python library providing high-performance implementations of classic and state-of-the-art Generative Adversarial Networks bu… | 23 | 8048 | maintenance |
| bigcode-project/starcoder2 StarCoder2 is a family of open code generation language models (3B, 7B, 15B) trained on 600+ programming languages from The Stack v2, with … | 26 | 2087 | active |
| visomaster/VisoMaster VisoMaster is a Python-based desktop application for AI-powered face swapping and face editing in images and videos. It supports multiple s… | 27 | 2052 | active |
| 01-ai/Yi Yi is a family of open-source large language models trained from scratch by 01.AI, including base and chat models in multiple sizes, with b… | 27 | 7836 | maintenance |
| serengil/retinaface RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin… | 61 | 2027 | active |
| TianZerL/Anime4KCPP Anime4KCPP is a high-performance anime image and video upscaler built on CNN-based algorithms, written in C++. It ships as a library plus V… | 78 | 2022 | active |
| juanmc2005/diart Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diariza… | 62 | 2022 | active |
| DEIM DEIMv2 is a real-time object detection framework that extends the DEIM DETR family with DINOv3-pretrained and distilled backbones plus a Sp… | 62 | 1999 | active |
| xingyizhou/CenterNet CenterNet is a PyTorch implementation of the 'Objects as Points' detector, which models objects as single center points detected via keypoi… | 32 | 7573 | maintenance |
| hkchengrex/XMem XMem is a PyTorch model for semi-supervised video object segmentation that tracks objects through long videos using an Atkinson-Shiffrin-in… | 23 | 1983 | stable |
| Netflix/void-model VOID (Video Object and Interaction Deletion) is a research model from Netflix that removes objects from videos along with the physical inte… | 54 | 1965 | active |
| eriklindernoren/PyTorch-YOLOv3 A minimal PyTorch implementation of YOLOv3 supporting training, inference, and evaluation, with compatibility for YOLOv4 and YOLOv7 weights… | 32 | 7440 | maintenance |
| jd-opensource/JoyAI-Echo JoyAI-Echo is a Python framework for long-horizon audio-visual generation, producing coherent multi-shot videos up to ~5 minutes with paire… | 58 | 1943 | active |
| Tencent-Hunyuan/HunyuanOCR HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side … | 59 | 1930 | active |
| sicxu/Deep3DFaceRecon_pytorch A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag… | 32 | 1907 | stable |
| Faceplugin-ltd/Open-Source-Face-Recognition-SDK An open-source face recognition SDK by Faceplugin providing face detection, landmark extraction, feature embedding generation, and face tem… | 64 | 1903 | active |
| vt-vl-lab/3d-photo-inpainting A Python research codebase from a CVPR 2020 paper that converts a single RGB-D image into a 3D photo using layered depth inpainting. It hal… | 32 | 7093 | maintenance |
| we0091234/Chinese_license_plate_detection_recognition A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports … | 71 | 1868 | active |
| Omni-Avatar/OmniAvatar OmniAvatar is an audio-driven full-body avatar video generation model built on Wan2.1 text-to-video diffusion models with LoRA-based audio … | 34 | 1859 | active |
| Tele-AI/Telechat TeleChat is a family of open-source bilingual (Chinese-English) large language models (1B, 7B, 12B) developed by China Telecom's AI team, r… | 67 | 1853 | active |
| descriptinc/descript-audio-codec Descript Audio Codec (.dac) is a high-fidelity neural audio codec based on improved RVQGAN that compresses audio at roughly 90x compression… | 61 | 1846 | stable |
| OpenImagingLab/FlashVSR FlashVSR is a one-step diffusion-based streaming video super-resolution framework that runs at ~17 FPS for 768x1408 video on a single A100 … | 61 | 1799 | active |
| Stability-AI/stable-fast-3d Stable Fast 3D (SF3D) is Stability AI's open-source model that reconstructs a textured, UV-unwrapped 3D mesh from a single input image in a… | 24 | 1794 | active |
| character-ai/Ovi Ovi is a video-plus-audio generation model from Character AI that simultaneously generates synchronized video and audio from text or text+i… | 40 | 1748 | active |
| SHI-Labs/OneFormer OneFormer is a universal image segmentation framework (CVPR 2023) that unifies semantic, instance, and panoptic segmentation in a single tr… | 32 | 1736 | stable |
| AnswerDotAI/ModernBERT ModernBERT is the research repository for a modernized BERT-family bidirectional encoder trained on 2 trillion tokens with an 8192-token co… | 56 | 1713 | active |
| MgArcher/Text_select_captcha A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio… | 69 | 1656 | active |
| Stability-AI/stable-virtual-camera Stable Virtual Camera (SEVA) is a generalist diffusion model for novel view synthesis that generates 3D-consistent views of a scene from an… | 54 | 1652 | active |
| ai-forever/ghost GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i… | 26 | 1582 | active |
| kotaro-kinoshita/yomitoku YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur… | 87 | 1579 | active |
| menyifang/MIMO MIMO is the official PyTorch implementation of a CVPR 2025 paper on controllable character video synthesis using spatially decomposed model… | 34 | 1578 | active |
| Layout-Parser/layout-parser LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout… | 23 | 5774 | maintenance |
| Tencent/AngelSlim AngelSlim is a Python toolkit from Tencent for compressing large language models and related architectures (VLMs, diffusion, audio models) … | 72 | 1547 | active |
| hkchengrex/Tracking-Anything-with-DEVA DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona… | 27 | 1508 | stable |
| Om-Alve/smolGPT A minimal pure-PyTorch implementation for training small GPT-style LLMs from scratch, featuring flash attention, RMSNorm, SwiGLU, RoPE, and… | 23 | 1475 | active |
| microsoft/NeuralSpeech NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro… | 32 | 1462 | active |
| Turing-Project/WriteGPT WriteGPT is a generative text-creation AI framework built on GPT-2 and other models (EAST, CRNN, BERT), fine-tuned to generate Chinese exam… | 23 | 5289 | maintenance |
| neuralchen/SimSwap SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It… | 23 | 5188 | maintenance |
| sql-machine-learning/sqlflow SQLFlow is a compiler that extends SQL with AI-oriented syntax (training, prediction, evaluation, explanation, and mathematical programming… | 23 | 5188 | maintenance |
| facebookresearch/vggsfm VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f… | 29 | 1421 | active |
| CSAILVision/semantic-segmentation-pytorch A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra… | 32 | 5078 | maintenance |
| yfeng95/PRNet PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali… | 32 | 5013 | maintenance |
| om-ai-lab/OmDet OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o… | 57 | 1393 | active |
| QwenAudio/ThinkSound ThinkSound is a PyTorch implementation of a NeurIPS 2025 framework that generates and edits audio from video, text, or audio inputs using C… | 52 | 1378 | active |
| siyuanliii/masa Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros… | 33 | 1377 | active |
| BICLab/SpikingBrain-7B SpikingBrain-7B is a brain-inspired large language model that combines hybrid efficient attention, MoE modules, and spike encoding, with a … | 54 | 1369 | active |
| lxtGH/OMG-Seg Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le… | 47 | 1354 | active |
| yinguobing/head-pose-estimation A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial… | 23 | 1353 | stable |
| owlbarn/owl Owl is an OCaml library for scientific and engineering computing, providing n-dimensional arrays, linear algebra, statistics, optimization,… | 66 | 1351 | active |