Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: deep-learning

2653 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
huggingface/smollm
Hugging Face's repository for the SmolLM and SmolVLM families of compact, fully open language and vision-language models, including trainin…
593884active
Avatarify
Avatarify is an open-source application that drives photorealistic avatars in real time for video-conferencing apps like Zoom and Skype, ba…
2316515maintenance
IDEA-Research/Grounded-SAM-2
Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X)…
373708active
xinyu1205/recognize-anything
Recognize Anything is a collection of open-source image recognition foundation models, including RAM, RAM++, and Tag2Text, that perform ima…
333708active
thu-ml/SageAttention
SageAttention is a family of quantized attention kernels (INT8/FP8/FP4) that accelerate transformer inference 2-5x over FlashAttention with…
403684active
zai-org/ChatGLM2-6B
ChatGLM2-6B is an open-source bilingual (Chinese-English) 6B-parameter conversational large language model built on the GLM architecture. I…
2915528maintenance
ToTheBeginning/PuLID
PuLID is the official PyTorch implementation of a NeurIPS 2024 method for inserting a specific person's identity into text-to-image generat…
403550active
ob-f/OpenBot
OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo…
673442active
jixiaozhong/Sonic
Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation…
493273active
prs-eth/Marigold
Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus…
523198active
tekaratzas/RustGPT
A transformer-based large language model implemented entirely in pure Rust with no external ML frameworks, using only ndarray for matrix op…
383157active
Djdefrag/QualityScaler
QualityScaler is a Windows GUI application that uses AI deep-learning models to upscale, enhance, and de-noise images and videos. It is wri…
953138active
z-x-yang/Segment-and-Track-Anything
An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg…
613134active
ridgerchu/matmulfreellm
A Python implementation of MatMul-Free LM, a language model architecture that eliminates matrix multiplication operations using ternary wei…
493089active
naver/mast3r
MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r…
373088active
Doubiiu/DynamiCrafter
DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p…
273007active
williamyang1991/Rerender_A_Video
The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.…
292999stable
luminal-ai/luminal
Luminal is a high-performance general-purpose ML inference compiler written in Rust that lowers models to a minimal 15-op dataflow IR and c…
782956active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542931active
facebookresearch/omnilingual-asr
An open-source multilingual speech recognition library from Meta AI supporting over 1,600 languages, including hundreds never previously co…
522898active
NVlabs/FoundationStereo
FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV…
472874active
UX-Decoder/Semantic-SAM
Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired …
332854active
TMElyralab/MuseV
MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall…
252846active
Camb-ai/MARS5-TTS
MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with …
142817active
microsoft/MoGe
MoGe is a deep learning model from Microsoft Research that recovers 3D geometry from a single open-domain image, predicting metric point ma…
662807active
autodistill/autodistill
Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab…
292763active
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
ModelTC/LightX2V
LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-…
642733active
magic-research/magic-animate
MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image …
4410897maintenance
yxlllc/DDSP-SVC
DDSP-SVC is an open-source singing voice conversion system built on Differentiable Digital Signal Processing, designed as a free AI voice c…
652656active
luca-medeiros/lang-segment-anything
A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie…
422598active
facebookresearch/nougat
Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py…
2310063maintenance
ipazc/mtcnn
A Python library implementing the MTCNN (Multitask Cascaded Convolutional Networks) algorithm for face detection and facial landmark alignm…
232485stable
data-infra/cube-studio
CubeStudio is an open-source, cloud-native, all-in-one AI platform covering the full machine learning lifecycle (MLOps/MaaS/LLMOps), includ…
802448active
pnnbao97/VieNeu-TTS
VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi…
842427active
ailia-ai/ailia-models
A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,…
772385active
jik876/hifi-gan
The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k…
322367stable
MouseLand/cellpose
Cellpose is a generalist deep learning algorithm for cellular and nucleus segmentation in microscopy images, with human-in-the-loop capabil…
862331active
labmlai/labml
A Python library for tracking and monitoring deep learning experiments, with a self-hostable server app for viewing metrics and hardware us…
302325active
XiaomiMiMo/MiMo
Xiaomi's MiMo is a 7B-parameter reasoning language model trained from pretraining through posttraining with reinforcement learning, release…
302299active
OlafenwaMoses/ImageAI
ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob…
238877maintenance
ermig1979/Simd
Simd Library is a free open-source C++ image processing and machine learning library with a C API and Python wrapper. Its algorithms are ha…
982265active
THU-MIG/yoloe
YOLOE is the official PyTorch implementation of an open-vocabulary object detection and segmentation model presented at ICCV 2025. It unifi…
322256active
MVIG-SJTU/AlphaPose
AlphaPose is an open-source real-time multi-person full-body pose estimation and tracking system built on PyTorch. It detects human keypoin…
328596maintenance
DigitalPhonetics/IMS-Toucan
IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive…
632207active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992203active
intel/intel-extension-for-transformers
Intel's toolkit for accelerating transformer-based GenAI/LLM workloads on Intel platforms, offering state-of-the-art compression (e.g., INT…
102174active
MoonshotAI/MoBA
MoBA (Mixture of Block Attention) is a PyTorch implementation of a trainable block-sparse attention mechanism for long-context large langua…
272169active
Tencent-Hunyuan/HunyuanVideo-Avatar
HunyuanVideo-Avatar is Tencent's open-source model and inference code for high-fidelity audio-driven human animation, generating talking av…
452156active
espressif/esp-who
ESP-WHO is an image processing development platform from Espressif providing face detection, face recognition, pedestrian detection, and QR…
672133active
LiheYoung/Depth-Anything
Depth Anything is a monocular depth estimation foundation model trained on 1.5M labeled and 62M+ unlabeled images, released as a Python lib…
268195maintenance
PaddlePaddle/PaddleGAN
PaddleGAN is a Python library providing high-performance implementations of classic and state-of-the-art Generative Adversarial Networks bu…
238048maintenance
bigcode-project/starcoder2
StarCoder2 is a family of open code generation language models (3B, 7B, 15B) trained on 600+ programming languages from The Stack v2, with …
262087active
visomaster/VisoMaster
VisoMaster is a Python-based desktop application for AI-powered face swapping and face editing in images and videos. It supports multiple s…
272052active
01-ai/Yi
Yi is a family of open-source large language models trained from scratch by 01.AI, including base and chat models in multiple sizes, with b…
277836maintenance
serengil/retinaface
RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin…
612027active
TianZerL/Anime4KCPP
Anime4KCPP is a high-performance anime image and video upscaler built on CNN-based algorithms, written in C++. It ships as a library plus V…
782022active
juanmc2005/diart
Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diariza…
622022active
DEIM
DEIMv2 is a real-time object detection framework that extends the DEIM DETR family with DINOv3-pretrained and distilled backbones plus a Sp…
621999active
xingyizhou/CenterNet
CenterNet is a PyTorch implementation of the 'Objects as Points' detector, which models objects as single center points detected via keypoi…
327573maintenance
hkchengrex/XMem
XMem is a PyTorch model for semi-supervised video object segmentation that tracks objects through long videos using an Atkinson-Shiffrin-in…
231983stable
Netflix/void-model
VOID (Video Object and Interaction Deletion) is a research model from Netflix that removes objects from videos along with the physical inte…
541965active
eriklindernoren/PyTorch-YOLOv3
A minimal PyTorch implementation of YOLOv3 supporting training, inference, and evaluation, with compatibility for YOLOv4 and YOLOv7 weights…
327440maintenance
jd-opensource/JoyAI-Echo
JoyAI-Echo is a Python framework for long-horizon audio-visual generation, producing coherent multi-shot videos up to ~5 minutes with paire…
581943active
Tencent-Hunyuan/HunyuanOCR
HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side …
591930active
sicxu/Deep3DFaceRecon_pytorch
A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag…
321907stable
Faceplugin-ltd/Open-Source-Face-Recognition-SDK
An open-source face recognition SDK by Faceplugin providing face detection, landmark extraction, feature embedding generation, and face tem…
641903active
vt-vl-lab/3d-photo-inpainting
A Python research codebase from a CVPR 2020 paper that converts a single RGB-D image into a 3D photo using layered depth inpainting. It hal…
327093maintenance
we0091234/Chinese_license_plate_detection_recognition
A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports …
711868active
Omni-Avatar/OmniAvatar
OmniAvatar is an audio-driven full-body avatar video generation model built on Wan2.1 text-to-video diffusion models with LoRA-based audio …
341859active
Tele-AI/Telechat
TeleChat is a family of open-source bilingual (Chinese-English) large language models (1B, 7B, 12B) developed by China Telecom's AI team, r…
671853active
descriptinc/descript-audio-codec
Descript Audio Codec (.dac) is a high-fidelity neural audio codec based on improved RVQGAN that compresses audio at roughly 90x compression…
611846stable
OpenImagingLab/FlashVSR
FlashVSR is a one-step diffusion-based streaming video super-resolution framework that runs at ~17 FPS for 768x1408 video on a single A100 …
611799active
Stability-AI/stable-fast-3d
Stable Fast 3D (SF3D) is Stability AI's open-source model that reconstructs a textured, UV-unwrapped 3D mesh from a single input image in a…
241794active
character-ai/Ovi
Ovi is a video-plus-audio generation model from Character AI that simultaneously generates synchronized video and audio from text or text+i…
401748active
SHI-Labs/OneFormer
OneFormer is a universal image segmentation framework (CVPR 2023) that unifies semantic, instance, and panoptic segmentation in a single tr…
321736stable
AnswerDotAI/ModernBERT
ModernBERT is the research repository for a modernized BERT-family bidirectional encoder trained on 2 trillion tokens with an 8192-token co…
561713active
MgArcher/Text_select_captcha
A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio…
691656active
Stability-AI/stable-virtual-camera
Stable Virtual Camera (SEVA) is a generalist diffusion model for novel view synthesis that generates 3D-consistent views of a scene from an…
541652active
ai-forever/ghost
GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i…
261582active
kotaro-kinoshita/yomitoku
YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur…
871579active
menyifang/MIMO
MIMO is the official PyTorch implementation of a CVPR 2025 paper on controllable character video synthesis using spatially decomposed model…
341578active
Layout-Parser/layout-parser
LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout…
235774maintenance
Tencent/AngelSlim
AngelSlim is a Python toolkit from Tencent for compressing large language models and related architectures (VLMs, diffusion, audio models) …
721547active
hkchengrex/Tracking-Anything-with-DEVA
DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona…
271508stable
Om-Alve/smolGPT
A minimal pure-PyTorch implementation for training small GPT-style LLMs from scratch, featuring flash attention, RMSNorm, SwiGLU, RoPE, and…
231475active
microsoft/NeuralSpeech
NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro…
321462active
Turing-Project/WriteGPT
WriteGPT is a generative text-creation AI framework built on GPT-2 and other models (EAST, CRNN, BERT), fine-tuned to generate Chinese exam…
235289maintenance
neuralchen/SimSwap
SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It…
235188maintenance
sql-machine-learning/sqlflow
SQLFlow is a compiler that extends SQL with AI-oriented syntax (training, prediction, evaluation, explanation, and mathematical programming…
235188maintenance
facebookresearch/vggsfm
VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f…
291421active
CSAILVision/semantic-segmentation-pytorch
A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra…
325078maintenance
yfeng95/PRNet
PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali…
325013maintenance
om-ai-lab/OmDet
OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o…
571393active
QwenAudio/ThinkSound
ThinkSound is a PyTorch implementation of a NeurIPS 2025 framework that generates and edits audio from video, text, or audio inputs using C…
521378active
siyuanliii/masa
Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros…
331377active
BICLab/SpikingBrain-7B
SpikingBrain-7B is a brain-inspired large language model that combines hybrid efficient attention, MoE modules, and spike encoding, with a …
541369active
lxtGH/OMG-Seg
Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le…
471354active
yinguobing/head-pose-estimation
A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial…
231353stable
owlbarn/owl
Owl is an OCaml library for scientific and engineering computing, providing n-dimensional arrays, linear algebra, statistics, optimization,…
661351active

← prev page 25 / 27 next →