function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ByteDance-Seed/SeedVR SeedVR/SeedVR2 are diffusion-transformer based models for generic real-world and AIGC video and image restoration, with SeedVR2 using adver… | 47 | 1334 | active |
| Vahe1994/AQLM Official PyTorch implementation of AQLM, an extreme LLM compression method via additive quantization, extended with PV-Tuning for finetunin… | 57 | 1329 | active |
| seetaface/SeetaFaceEngine SeetaFace Engine is an open-source C++ face recognition engine comprising face detection, face alignment, and face identification modules. … | 32 | 4636 | maintenance |
| Parskatt/RoMa RoMa (romatch) is a Python library for robust dense feature matching between image pairs, estimating pixel-dense warps and reliable certain… | 52 | 1293 | active |
| bytedance/Bernini Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer perf… | 57 | 1287 | active |
| Tianxiaomo/pytorch-YOLOv4 A minimal PyTorch implementation of YOLOv4 (and YOLOv4-tiny) supporting inference and training, with tools to convert Darknet weights to Py… | 32 | 4521 | maintenance |
| zju3dv/MatchAnything MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap… | 64 | 1279 | active |
| city-super/Scaffold-GS Scaffold-GS is a research implementation of a structured 3D Gaussian splatting method that uses anchor points on a sparse voxel grid to dis… | 27 | 1278 | active |
| DreamLM/Dream Dream 7B is an open diffusion large language model (dLLM) with base and instruct checkpoints, plus inference and training code built on Hug… | 44 | 1265 | active |
| lpiccinelli-eth/UniDepth UniDepth is a Python library and research codebase for universal monocular metric depth estimation from single images, based on CVPR 2024 a… | 35 | 1246 | active |
| ZHKKKe/MODNet MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima… | 32 | 4355 | maintenance |
| sh-lee-prml/HierSpeechpp Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera… | 28 | 1238 | active |
| higgsfield-ai/higgsfield Higgsfield is an open-source GPU orchestration and machine learning framework for fault-tolerant, distributed training of very large models… | 23 | 4106 | maintenance |
| Soul-AILab/SoulX-LiveAct SoulX-LiveAct is the official inference code for a real-time human animation framework that generates lifelike, audio/multimodal-controlled… | 54 | 1176 | active |
| NVlabs/imaginaire NVIDIA's PyTorch library containing optimized implementations of image and video synthesis methods, including GAN-based image-to-image tran… | 32 | 4082 | maintenance |
| FoundationVision/GLEE GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o… | 26 | 1170 | active |
| wladradchenko/wunjo.wladradchenko.ru Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a… | 70 | 1169 | active |
| cure-lab/MagicDrive MagicDrive is the official PyTorch implementation of an ICLR 2024 paper for controllable street view generation using diffusion models with… | 35 | 1166 | active |
| facebookresearch/fairseq2 fairseq2 is a PyTorch-based sequence modeling toolkit from Meta FAIR for training custom models for content generation tasks such as langua… | 89 | 1143 | active |
| open-mmlab/mmtracking MMTracking is OpenMMLab's PyTorch-based toolbox for video perception tasks, unifying video object detection, multiple object tracking, sing… | 23 | 3897 | maintenance |
| princeton-vl/RAFT-Stereo RAFT-Stereo is a PyTorch implementation of a deep learning model for stereo matching that estimates disparity maps from stereo image pairs … | 75 | 1119 | stable |
| hkchengrex/Cutie Cutie is a video object segmentation framework with object-level memory reading, a follow-up to XMem offering better consistency, robustnes… | 18 | 1095 | active |
| LujiaJin/One-Pot_Multi-Frame_Denoising Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides … | 60 | 1094 | stable |
| yandex/YaLM-100B YaLM-100B is a GPT-like pretrained language model with 100 billion parameters, trained by Yandex on English and Russian text using DeepSpee… | 32 | 3757 | maintenance |
| lasgroup/SDPO SDPO (Self-Distilled Policy Optimization) is a research library implementing a reinforcement learning framework for post-training large lan… | 56 | 1075 | active |
| microsoft/Biodiversity Microsoft AI for Good Lab's biodiversity research hub providing open-source AI models and tools for wildlife monitoring and conservation, i… | 88 | 1066 | active |
| InternRobotics/InternNav InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation… | 58 | 1061 | active |
| zai-org/GLM-TTS GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur… | 50 | 1055 | active |
| williamyang1991/VToonify Official PyTorch implementation of VToonify, a SIGGRAPH Asia 2022 framework for controllable high-resolution portrait video style transfer … | 32 | 3584 | maintenance |
| qqlu/Entity EntitySeg is an open-source PyTorch toolbox for open-world, high-quality image segmentation, built on Detectron2. It aggregates multiple re… | 32 | 1048 | active |
| Xilinx/finn FINN is an open-source dataflow compiler from AMD/Xilinx that generates highly efficient FPGA accelerators for quantized neural network (QN… | 68 | 1046 | active |
| Tencent-Hunyuan/InstantCharacter InstantCharacter is a tuning-free framework built on diffusion transformers that generates character-consistent images from a single refere… | 29 | 1045 | active |
| ml-tooling/ml-workspace ML Workspace is an all-in-one web-based IDE Docker image specialized for machine learning and data science. It bundles Jupyter, JupyterLab,… | 23 | 3544 | maintenance |
| Jumpat/SegmentAnythingin3D SA3D is a research framework that lifts 2D Segment Anything (SAM) masks into 3D segmentation of objects within a NeRF or 3D Gaussian Splatt… | 40 | 1030 | active |
| kuleshov-group/bd3lms BD3-LMs is a research implementation of Block Discrete Denoising Diffusion Language Models that interpolate between autoregressive and diff… | 34 | 1029 | active |
| zhyever/PatchFusion PatchFusion is a CVPR 2024 end-to-end tile-based framework for high-resolution monocular metric depth estimation from single images. It fus… | 57 | 1026 | active |
| microsoft/tensorwatch TensorWatch is a Python library from Microsoft Research for debugging, monitoring, and visualizing machine learning training in real time, … | 66 | 3471 | maintenance |
| yangxy/PASD PASD (Pixel-Aware Stable Diffusion) is a Python research codebase implementing an ECCV 2024 method for realistic image super-resolution and… | 28 | 1021 | active |
| eragonruan/text-detection-ctpn A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr… | 23 | 3429 | maintenance |
| HumeAI/tada TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,… | 51 | 1009 | active |
| facebookresearch/Mask2Former Mask2Former is the official PyTorch implementation of the CVPR 2022 paper 'Masked-attention Mask Transformer for Universal Image Segmentati… | 10 | 3416 | maintenance |
| minimaxir/gpt-2-simple A Python package that simplifies fine-tuning OpenAI's GPT-2 text-generation model (124M/355M) on custom text and generating text from the r… | 23 | 3400 | maintenance |
| clovaai/CRAFT-pytorch Official PyTorch implementation of CRAFT (Character Region Awareness for Text Detection), a scene text detector that localizes text by pred… | 32 | 3398 | maintenance |
| bytedance/lightseq LightSeq is a high-performance CUDA-based library for training and inference of sequence models like Transformer, BERT, GPT, and BART, with… | 10 | 3295 | maintenance |
| google-research/albert Official TensorFlow implementation and pretrained checkpoints of ALBERT, a lite version of BERT for self-supervised learning of language re… | 10 | 3278 | maintenance |
| anandpawara/Real_Time_Image_Animation A real-time Python application that animates a still image (e.g., a portrait) using facial motion from a live camera or video file, built o… | 32 | 3248 | maintenance |
| mkocabas/VIBE Official PyTorch implementation of VIBE (CVPR 2020), a video-based method for 3D human body pose and shape estimation that predicts SMPL bo… | 23 | 3211 | maintenance |
| google-research/frame-interpolation FILM is the official TensorFlow 2 implementation of a state-of-the-art frame interpolation neural network from Google Research, presented a… | 10 | 3150 | maintenance |
| dbiir/UER-py UER-py is a PyTorch framework for pre-training transformer language models (BERT, GPT-2, T5, ELMo, etc.) and fine-tuning them on downstream… | 32 | 3112 | maintenance |
| biubug6/Pytorch_Retinaface A PyTorch implementation of the RetinaFace single-stage face detection model, supporting mobilenet0.25 and resnet50 backbones with pretrain… | 32 | 2976 | maintenance |
| Tencent/FaceDetection-DSFD DSFD (Dual Shot Face Detector) is Tencent Youtu's high-accuracy face detection network, released with PyTorch inference code and pretrained… | 56 | 2969 | maintenance |
| Alpha-VLLM/LLaMA2-Accessory LLaMA2-Accessory is an open-source Python toolkit for pretraining, finetuning, and deploying large language models and multimodal LLMs, inc… | 29 | 2800 | maintenance |
| YCG09/chinese_ocr An end-to-end Chinese OCR system implemented with TensorFlow and Keras, combining CTPN for text detection with DenseNet + CTC for text reco… | 32 | 2782 | maintenance |
| tensorflow/graphics TensorFlow Graphics is a library of differentiable graphics layers for TensorFlow, including differentiable renderers, spatial transformers… | 64 | 2781 | maintenance |
| HypoX64/DeepMosaics DeepMosaics is a Python application that automatically removes or adds mosaics in images and videos using semantic segmentation and image-t… | 23 | 2633 | maintenance |
| PeterH0323/Smart_Construction A YOLOv5-based object detection application for detecting people, heads, and safety helmets on construction sites, including pretrained wei… | 23 | 2619 | maintenance |
| microsoft/GLIP GLIP is Microsoft's official implementation of Grounded Language-Image Pre-training, a vision-language model that unifies object detection … | 32 | 2607 | maintenance |
| zllrunning/face-parsing.PyTorch A PyTorch implementation of face parsing using a modified BiSeNet architecture, trained on the CelebAMask-HQ dataset. It provides training … | 32 | 2586 | maintenance |
| s3prl/s3prl S3PRL is a PyTorch toolkit for self-supervised speech pre-training and representation learning, bundling many upstream models like wav2vec … | 55 | 2561 | maintenance |
| yfeng95/DECA DECA is the official PyTorch implementation of a SIGGRAPH 2021 method that reconstructs a detailed 3D head model (pose, shape, facial detai… | 32 | 2514 | maintenance |
| leggedrobotics/darknet_ros A ROS package wrapping the YOLO (Darknet) real-time object detector for use in robotic systems. It subscribes to camera image topics and pu… | 23 | 2439 | maintenance |
| Hzzone/pytorch-openpose A PyTorch reimplementation of OpenPose for body and hand pose estimation, with models converted directly from the original OpenPose caffemo… | 32 | 2321 | maintenance |
| MhLiao/DB A PyTorch implementation of DBNet and DBNet++, real-time arbitrary-shape scene text detection models based on differentiable binarization. … | 32 | 2260 | maintenance |
| hustvl/YOLOP YOLOP is a multi-task deep learning network that jointly performs traffic object detection, drivable area segmentation, and lane detection … | 32 | 2234 | maintenance |
| x-flux XLabs AI's training scripts for fine-tuning the FLUX.1 diffusion model with LoRA, ControlNet, and IP-Adapter adapters, using DeepSpeed and … | 23 | 2229 | maintenance |
| mit-han-lab/temporal-shift-module PyTorch implementation of the Temporal Shift Module (TSM), an ICCV 2019 technique that adds temporal modeling to 2D CNNs at zero extra comp… | 32 | 2221 | maintenance |
| alibaba/EasyNLP EasyNLP is a comprehensive PyTorch-based NLP toolkit from Alibaba that provides training, inference, and deployment for pre-trained languag… | 23 | 2184 | maintenance |
| bubbliiiing/yolov4-pytorch A PyTorch implementation of the YOLOv4 object detection model with full training, prediction, and evaluation scripts. It supports training … | 23 | 2160 | maintenance |
| facebookresearch/ELF ELF is an end-to-end, lightweight and flexible C++/Python platform for game research, focused on real-time strategy games. It hosts multipl… | 10 | 2090 | maintenance |
| jiupinjia/SkyAR SkyAR is the official PyTorch implementation of the paper 'Castle in the Sky: Dynamic Sky Replacement and Harmonization in Videos'. It perf… | 32 | 2025 | maintenance |
| facebookresearch/Detic Detic is the official code release for the ECCV 2022 paper 'Detecting Twenty-thousand Classes using Image-level Supervision'. It is an open… | 32 | 2008 | maintenance |
| Music-and-Culture-Technology-Lab/omnizart Omnizart is a Python library and CLI for automatic music transcription, transcribing pitched instruments, vocal melody, chords, drum events… | 88 | 1964 | maintenance |
| yehengchen/Object-Detection-and-Tracking A collection of Python implementations combining YOLO-based object detection with SORT and DeepSORT multi-object tracking. It includes exam… | 32 | 1961 | maintenance |
| mpatacchiola/deepgaze Deepgaze is a Python computer vision library for human-computer interaction built on OpenCV and TensorFlow. It provides CNN-based head pose… | 32 | 1880 | maintenance |
| AlexeyAB/Yolo_mark A Windows and Linux GUI application for drawing bounding boxes around objects in images to create labeled training data for YOLO v2/v3 obje… | 32 | 1840 | maintenance |
| symisc/sod SOD is an embedded, cross-platform computer vision and machine learning library written in C, distributed as a single dependency-free amalg… | 23 | 1798 | maintenance |
| sergiomsilva/alpr-unconstrained An implementation of the ECCV 2018 paper 'License Plate Detection and Recognition in Unconstrained Scenarios', combining a Darknet-based de… | 32 | 1770 | maintenance |
| WXinlong/SOLO Official PyTorch implementation of SOLO and SOLOv2, box-free fully convolutional methods for instance segmentation published at ECCV 2020 a… | 32 | 1758 | maintenance |
| deepset-ai/FARM FARM is a Python framework for fine-tuning and evaluating transformer-based language models for NLP tasks, with a focus on question answeri… | 10 | 1752 | maintenance |
| natethegreate/hent-AI A Python application that automatically detects censor bars and mosaic blurs in illustrated adult content using deep learning (Mask R-CNN) … | 23 | 1730 | maintenance |
| LoSealL/VideoSuperResolution A Python library (pip-installable as VSR) collecting reimplementation of state-of-the-art single-image and video super-resolution neural ne… | 23 | 1687 | maintenance |
| YuliangXiu/ICON ICON is a PyTorch research implementation of a CVPR 2022 method that reconstructs detailed, animatable 3D clothed human avatars from 2D ima… | 23 | 1675 | maintenance |
| lucidrains/PaLM-rlhf-pytorch A PyTorch library implementing Reinforcement Learning from Human Feedback (RLHF) on top of the PaLM transformer architecture, aiming to rep… | 78 | 7868 | experimental |
| chandrikadeb7/Face-Mask-Detection A face mask detection system built with OpenCV and TensorFlow/Keras that uses deep learning (SSD MobileNetV2) to detect whether people are … | 23 | 1606 | maintenance |
| PeterWang512/FALdetector FALdetector is the official PyTorch implementation of the ICCV 2019 paper 'Detecting Photoshopped Faces by Scripting Photoshop'. It provide… | 32 | 1600 | maintenance |
| cmdbug/YOLOv5_NCNN A mobile demo application that deploys the ncnn inference framework on Android and iOS, running a variety of computer vision models includi… | 32 | 1572 | maintenance |
| IDEA-Research/MaskDINO Official PyTorch implementation of Mask DINO, a unified transformer-based framework for object detection and segmentation, built on detectr… | 23 | 1557 | maintenance |
| DevashishPrasad/CascadeTabNet CascadeTabNet is a PyTorch/mmdetection implementation of a CVPR 2020 paper for end-to-end table detection and structure recognition from im… | 32 | 1549 | maintenance |
| FudanNLP/fitlog fitlog (fast + git + log) is a Python tool that helps deep learning practitioners record training logs and manage experiment code, combinin… | 23 | 1512 | maintenance |
| cvg/pixel-perfect-sfm pixsfm is a Python package with a C++ core that improves Structure-from-Motion and visual localization accuracy by refining keypoints, came… | 23 | 1484 | maintenance |
| Sharpiless/Yolov5-deepsort-inference A Python library combining YOLOv5 object detection with DeepSort multi-object tracking to detect, track, and count vehicles and pedestrians… | 66 | 1479 | maintenance |
| peteanderson80/bottom-up-attention A bottom-up attention model based on Faster R-CNN with ResNet-101 trained on Visual Genome, producing features for salient image regions. T… | 32 | 1469 | maintenance |
| microsoft/SpeechT5 Microsoft's unified-modal speech-text pre-training framework implementing SpeechT5 and related models (SpeechLM, SpeechUT, VALL-E X, WavLLM… | 32 | 1449 | maintenance |
| ConnorJL/GPT2 A community Python/TensorFlow implementation of GPT-2 model training and text generation that supports both GPUs and TPUs. It includes scri… | 32 | 1412 | maintenance |
| innnky/emotional-vits Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual … | 32 | 1392 | maintenance |
| zhubenfu/License-Plate-Detect-Recognition-via-Deep-Neural-Networks-accuracy-up-to-99.9 A C++ application that detects and recognizes Chinese license plates in real time using deep neural networks, claiming up to 99.8% accuracy… | 32 | 1384 | maintenance |
| myhub/tr An offline Chinese text detection and recognition OCR SDK with C++ core code and Python bindings, supporting models like CRNN, CTPN, and Pi… | 50 | 1381 | maintenance |
| amazon-science/patchcore-inspection Official implementation of PatchCore, a deep-learning method for industrial image anomaly detection and localization from Roth et al. (2021… | 32 | 1373 | maintenance |
| tensorboy/pytorch_Realtime_Multi-Person_Pose_Estimation A PyTorch implementation of the CVPR'17 Realtime Multi-Person 2D Pose Estimation (OpenPose/rtpose) model. It provides pretrained weights, d… | 32 | 1371 | maintenance |
| dlunion/DBFace DBFace is a real-time, single-stage face detection model implemented in Python, offering small model sizes with high accuracy on the WiderF… | 32 | 1355 | maintenance |