Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: deep-learning

2771 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
pq-yang/MatAnyone
MatAnyone is a CVPR 2025 human video matting framework that extracts alpha mattes of target people from video using consistent memory propa…
541605active
meituan/YOLOv6
YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine…
235895maintenance
ml4a/ml4a
ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li…
321602active
autonomousvision/transfuser
Official PyTorch implementation of TransFuser, a transformer-based multi-modal sensor fusion model for end-to-end autonomous driving, publi…
531599stable
X-LANCE/AniTalker
AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v…
241598active
SonyResearch/micro_diffusion
Official implementation of Sony Research's micro-budget approach to training large-scale text-to-image diffusion transformer models from sc…
241594active
facebookresearch/fast3r
Fast3R is the official PyTorch implementation of a CVPR 2025 model from Meta FAIR that reconstructs 3D scenes and estimates camera poses fr…
101593active
google/uncertainty-baselines
A library of high-quality, minimal-dependency implementations of standard and state-of-the-art uncertainty and robustness methods for deep …
771592active
alibaba/Pai-Megatron-Patch
Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V…
561591active
Tencent-Hunyuan/HunyuanWorld-Voyager
HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ…
521590active
intel/auto-round
AutoRound is Intel's advanced quantization toolkit for LLMs and vision-language models, using sign-gradient descent to achieve high accurac…
911588active
databricks/megablocks
MegaBlocks is a lightweight Python library for efficient training of mixture-of-experts (MoE) models, built around its dropless-MoE (dMoE) …
621588active
MoonshotAI/Kimi-Linear
Kimi Linear is a hybrid linear attention architecture (Kimi Delta Attention, based on Gated DeltaNet) released by Moonshot AI with 48B-para…
401588active
FoundationVision/Infinity
Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints…
561587active
Drexubery/ViewCrafter
ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar…
491587active
DeepGraphLearning/torchdrug
TorchDrug is a PyTorch-based machine learning library for drug discovery, covering graph neural networks, deep generative models, and reinf…
231587active
yakhyo/uniface
UniFace is a unified Python library for face analysis that bundles detection, recognition, landmark localization, face parsing, gaze estima…
881586active
Tencent-Hunyuan/HY-WorldPlay
HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f…
551586active
pytorch/FBGEMM
FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere…
931584active
XPixelGroup/HAT
HAT (Hybrid Attention Transformer) is a PyTorch implementation of a state-of-the-art transformer model for image super-resolution and resto…
321583stable
microsoft/MMdnn
MMdnn is a Microsoft toolkit for converting, visualizing, and diagnosing deep learning models across frameworks such as TensorFlow, PyTorch…
395805maintenance
ai-forever/ghost
GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i…
261582active
meta-pytorch/torchtune
Torchtune is a PyTorch-native library for authoring, post-training, and experimenting with large language models. It provides hackable trai…
695801maintenance
AlibabaResearch/DAMO-ConvAI
The official codebase for Alibaba DAMO Academy's Conversational AI research, containing implementations of models from their published pape…
711580active
SalesforceAIResearch/uni2ts
Uni2TS is a PyTorch library for unified pre-training, fine-tuning, inference, and evaluation of universal time series forecasting transform…
601580active
RWKV/rwkv.cpp
A C++ port of the RWKV language model to the ggml tensor library, providing FP32, FP16, and quantized INT4/INT5/INT8 inference focused on C…
371579active
tensorflow/model-optimization
The TensorFlow Model Optimization Toolkit (tfmot) is a Python library providing tools to optimize machine learning models for deployment, i…
821578stable
menyifang/MIMO
MIMO is the official PyTorch implementation of a CVPR 2025 paper on controllable character video synthesis using spatially decomposed model…
341578active
DiffEqML/torchdyn
Torchdyn is a PyTorch library dedicated to numerical deep learning, providing tools for neural differential equations, implicit models, and…
231578active
BloodAxe/pytorch-toolbelt
A Python library of PyTorch extensions providing building blocks for fast R&D prototyping, including encoder-decoder architectures, special…
441574active
FluxML/Zygote.jl
Zygote.jl is a source-to-source automatic differentiation library for Julia that hooks into the Julia compiler to generate gradient (backwa…
971568active
Xilinx/brevitas
Brevitas is a PyTorch library for neural network quantization supporting both post-training quantization (PTQ) and quantization-aware train…
911567active
Chinese-Text-Classification-Pytorch
A PyTorch-based collection of ready-to-run Chinese text classification models including TextCNN, TextRNN, FastText, TextRCNN, BiLSTM with a…
325728maintenance
ZeyueT/AudioX
AudioX is a unified multimodal framework for anything-to-audio generation, producing audio and music conditioned on text, video, image, or …
521552active
kijai/ComfyUI-CogVideoXWrapper
A ComfyUI custom node wrapper for CogVideoX and related video generation models (including Fun variants, CogVideoX 1.5, and Go-with-the-Flo…
391549active
Tencent/AngelSlim
AngelSlim is a Python toolkit from Tencent for compressing large language models and related architectures (VLMs, diffusion, audio models) …
721547active
mosaicml/streaming
StreamingDataset is a Python library from MosaicML for fast, accurate streaming of training data from cloud object storage (S3, GCS, Azure,…
701547active
lucidrains/soundstorm-pytorch
A PyTorch implementation of SoundStorm, Google DeepMind's efficient parallel audio generation model that applies MaskGiT-style masked gener…
391546active
facebookresearch/mmf
MMF is a modular PyTorch framework for vision and language multimodal research from Facebook AI Research. It ships reference implementation…
645633maintenance
google-ai-edge/model-explorer
Model Explorer is a model graph visualizer and debugger from Google AI Edge that renders neural network graphs hierarchically with expandab…
931542active
vibevoice-community/VibeVoice
VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker …
601542active
hustvl/MapTR
MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It …
271542active
tensorflow/gnn
TensorFlow GNN is a Python library for building Graph Neural Networks on TensorFlow, including a GraphTensor type for heterogeneous graphs,…
661541active
lucidrains/DALLE-pytorch
A PyTorch implementation/replication of OpenAI's DALL-E, a text-to-image transformer, including a discrete VAE and optional CLIP for rankin…
235627maintenance
MoonshotAI/Moonlight
Moonlight is a 3B/16B Mixture-of-Experts LLM trained with the Muon optimizer, released by Moonshot AI along with a memory- and communicatio…
361540active
Arthur151/ROMP
ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen…
231538stable
xLLM-AI/xllm
xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su…
781536active
leela-zero/leela-zero
Leela Zero is an open-source Go engine that reimplements AlphaGo Zero, combining Monte Carlo Tree Search with a deep residual convolutional…
235586maintenance
JingyunLiang/SwinIR
Official PyTorch implementation of SwinIR, a Swin Transformer-based model for image restoration tasks including super-resolution, denoising…
235580maintenance
hustvl/LightningDiT
LightningDiT is a research codebase for latent diffusion models implementing VA-VAE and LightningDiT, achieving FID 1.35 on ImageNet-256 wi…
471529active
cchen156/Learning-to-See-in-the-Dark
TensorFlow implementation of 'Learning to See in the Dark' (CVPR 2018), a deep learning model that brightens very dark, short-exposure RAW …
515565maintenance
ByteDance-Seed/Triton-distributed
Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets…
631526active
WenmuZhou/PytorchOCR
A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP…
591523active
keras-rl/keras-rl
A Python library implementing state-of-the-art deep reinforcement learning algorithms (DQN, DDPG, SARSA, and more) that integrates seamless…
235547maintenance
NirAharon/BoT-SORT
BoT-SORT is a state-of-the-art multi-object tracker that combines motion and appearance information with camera motion compensation and an …
321522active
caiyuanhao1998/Retinexformer
Retinexformer is a one-stage Retinex-based Transformer model and toolbox for low-light image enhancement, published at ICCV 2023. It suppor…
661518active
allenzren/open-pi-zero
An open-source re-implementation of the pi0 vision-language-action (VLA) model from Physical Intelligence, built on a pre-trained PaliGemma…
251518active
Phantom-video/Phantom
Phantom is a subject-consistent video generation model from ByteDance that preserves reference subject identity via cross-modal alignment. …
391517active
microsoft/Mage
Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage…
571516active
ZFTurbo/Music-Source-Separation-Training
A Python training framework for music source separation models, supporting many architectures such as MDX23C, Demucs, Band Split RoFormer, …
831512active
decoderesearch/SAELens
SAELens is a Python library for training sparse autoencoders (SAEs) on language model activations and analyzing them for mechanistic interp…
881510active
HiDream-ai/HiDream-O1-Image
HiDream-O1-Image is an open-weights 8B image generation foundation model built on a Pixel-level Unified Transformer (UiT) that natively enc…
531510active
imoneoi/openchat
OpenChat is a library of open-source large language models fine-tuned with C-RLFT, an offline reinforcement learning strategy that learns f…
205490maintenance
wb14123/seq2seq-couplet
A deep learning project that generates Chinese couplets (对联) using a seq2seq model built with TensorFlow. It includes training scripts, a w…
325487maintenance
Flashlight
wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into …
625466maintenance
czczup/ViT-Adapter
Official PyTorch implementation of ViT-Adapter, an ICLR 2023 Spotlight paper introducing a pre-training-free adapter that lets plain Vision…
341503stable
charlesq34/pointnet
Reference implementation of PointNet, a neural network architecture that directly consumes unordered 3D point clouds for classification and…
325459maintenance
TencentARC/MotionCtrl
MotionCtrl is the official implementation of a SIGGRAPH 2024 paper providing a unified and flexible motion controller for video generation …
301500active
mattmireles/gemma-tuner-multimodal
A Python tool for LoRA fine-tuning of Gemma 4 and 3n models on text, images, and audio using Apple Silicon's Metal Performance Shaders. It …
621495active
ibab/tensorflow-wavenet
A TensorFlow implementation of DeepMind's WaveNet generative neural network architecture for raw audio waveform generation. It provides tra…
325428maintenance
open-mmlab/mmengine
MMEngine is the foundational training engine library for OpenMMLab projects, providing a unified training loop, config system, registry, ho…
711492active
bojone/bert4keras
A lightweight, clean reimplementation of BERT and other transformer models (RoBERTa, ALBERT, T5, GPT, ELECTRA, NEZHA) for Keras/tf.keras. I…
235415maintenance
google-deepmind/graph_nets
DeepMind's library for building graph networks (graph neural networks) in TensorFlow and Sonnet, based on the 'Relational inductive biases,…
325406maintenance
graphdeco-inria/diff-gaussian-rasterization
A CUDA-based differentiable rasterization engine for 3D Gaussian Splatting, used in the SIGGRAPH 2023 paper '3D Gaussian Splatting for Real…
291489stable
CUT3R/CUT3R
CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre…
381486active
k2-fsa/icefall
Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi…
641482active
hustvl/DiffusionDrive
DiffusionDrive is a truncated diffusion model for real-time end-to-end autonomous driving, released as the official PyTorch implementation …
441480active
Soul-AILab/SoulX-FlashTalk
SoulX-FlashTalk is a 14B audio-driven talking avatar model that streams infinite real-time video from a reference image and audio, achievin…
581479active
alxndrTL/mamba.py
A simple, readable pure-PyTorch (plus MLX) implementation of the Mamba state-space model architecture with a parallel scan for efficient tr…
531478active
Om-Alve/smolGPT
A minimal pure-PyTorch implementation for training small GPT-style LLMs from scratch, featuring flash attention, RMSNorm, SwiGLU, RoPE, and…
231475active
mit-han-lab/torchsparse
TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a…
261472active
sczhou/Upscale-A-Video
Upscale-A-Video is a diffusion-based model for real-world video super-resolution that takes low-resolution videos and text prompts as input…
261470active
Lightning-AI/lightning-thunder
Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a …
721469active
autonomousvision/mip-splatting
Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim…
271466active
hojonathanho/diffusion
The official reference implementation of Denoising Diffusion Probabilistic Models (DDPM) from the 2020 paper by Jonathan Ho et al., written…
325300maintenance
NVIDIA/tacotron2
NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi…
325296maintenance
yunjey/stargan
Official PyTorch implementation of StarGAN, a unified generative adversarial network for multi-domain image-to-image translation (CVPR 2018…
325296maintenance
FeiYull/TensorRT-Alpha
A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl…
321460active
tensorflow/tpu
A collection of reference models and tools for training machine learning models on Google Cloud TPUs, maintained as a public mirror by the …
725278maintenance
dbolya/yolact
YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++…
515241maintenance
zylo117/Yet-Another-EfficientDet-Pytorch
A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim…
235238maintenance
amdegroot/ssd.pytorch
A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include…
325221maintenance
zsyOAOA/InvSR
InvSR is a Python research library implementing arbitrary-steps image super-resolution via diffusion inversion, leveraging pre-trained diff…
511443active
voice-cloning-app/Voice-Cloning-App
A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom…
231440active
google-deepmind/rlax
RLax is a JAX-based library of building blocks for implementing reinforcement learning agents, providing mathematical operations like value…
831439active
Walter0807/MotionBERT
Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske…
651439active
tensorspace-team/tensorspace
TensorSpace is a neural network 3D visualization framework built on TensorFlow.js, Three.js, and Tween.js. It provides Keras-like APIs to b…
235191maintenance
tianweiy/DMD2
DMD2 is the official PyTorch implementation of Improved Distribution Matching Distillation, a NeurIPS 2024 method that distills diffusion m…
281438active
neuralchen/SimSwap
SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It…
235188maintenance
sql-machine-learning/sqlflow
SQLFlow is a compiler that extends SQL with AI-oriented syntax (training, prediction, evaluation, explanation, and mathematical programming…
235188maintenance

← prev page 10 / 28 next →