Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: deep-learning

2653 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
alibaba/Tora
Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati…
641241active
Roblox/cube
Cube is Roblox's open-source family of foundation models for 3D intelligence, including text-to-3D shape generation and part-controllable m…
581241active
XPixelGroup/HYPIR
Official PyTorch implementation of HYPIR, a SIGGRAPH 2025 method that harnesses diffusion-yielded score priors for image restoration. It pr…
391241active
jcjohnson/fast-neural-style
A Torch (Lua) implementation of feedforward neural style transfer from the ECCV 2016 paper 'Perceptual Losses for Real-Time Style Transfer …
324359maintenance
facebookresearch/deit
Official PyTorch repository for DeiT and related vision transformer architectures (CaiT, ResMLP, PatchConvnet, DeiT III), providing trainin…
104355maintenance
declare-lab/tango
Tango is a family of latent diffusion models for text-to-audio generation, with Tango 2 improving prompt alignment via DPO-based fine-tunin…
451239active
apchenstu/TensoRF
TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie…
441239stable
X-Square-Robot/wall-x
Wall-X is the open-source training and inference stack for X Square Robot's WALL series of embodied foundation models (VLAs) for general-pu…
611236active
meta-pytorch/attention-gym
Attention Gym is a collection of tools, examples, and reference implementations for working with PyTorch's FlexAttention API. It provides a…
911234active
chengzeyi/Comfy-WaveSpeed
A ComfyUI custom node plugin that acts as an all-in-one inference optimization solution for diffusion models, built around First Block Cach…
661231stable
lucidrains/deep-daze
Deep Daze is a simple command line tool for text-to-image generation that combines OpenAI's CLIP with a Siren implicit neural representatio…
234315maintenance
MoonshotAI/FlashKDA
FlashKDA is a set of high-performance CUDA kernels (built on CUTLASS) implementing Kimi Delta Attention, a linear attention mechanism, for …
571229active
CSSLab/maia-chess
Maia is a collection of human-like neural network chess engines trained on millions of human games, targeting skill levels from ELO 1100 to…
611228active
Tencent-Hunyuan/HunyuanCustom
HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video …
401227active
NVIDIA/BigVGAN
BigVGAN is NVIDIA's official PyTorch implementation of a universal neural vocoder (ICLR 2023) that generates high-fidelity raw audio wavefo…
231227stable
alibaba/x-deeplearning
X-DeepLearning (XDL) is an industrial deep learning framework from Alibaba optimized for high-dimension sparse data scenarios such as adver…
234304maintenance
MoonshotAI/Kimi-VL
Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo…
331224active
ElectricAlexis/NotaGen
NotaGen is a symbolic music generation model that produces high-quality classical sheet music using LLM-style training paradigms: pre-train…
321223active
eduardoleao052/js-pytorch
JS-PyTorch is a deep learning library for JavaScript that closely mirrors PyTorch's syntax, providing tensor operations, automatic differen…
161222active
SciML/NeuralPDE.jl
NeuralPDE.jl is a Julia library of physics-informed neural network (PINN) solvers for ordinary, stochastic, and partial differential equati…
991220active
MotrixLab/SMPLer-X
Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an…
591220stable
Project-MONAI/research-contributions
A collection of peer-reviewed research prototype implementations built on the MONAI framework for medical imaging AI. It serves as a fast-t…
431220active
zkonduit/ezkl
EZKL is a Rust-based library and command-line tool that converts deep learning models and arbitrary computational graphs (exported as ONNX)…
751219active
Aratako/Irodori-TTS
Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer…
581219active
lucidrains/perceiver-pytorch
A PyTorch implementation of the Perceiver architecture (General Perception with Iterative Attention) and its follow-up Perceiver IO. It pro…
621217active
fudan-generative-vision/champ
Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,…
254261maintenance
Picsart-AI-Research/Text2Video-Zero
Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta…
304245maintenance
ifzhang/FairMOT
FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification…
324244maintenance
hasktorch/hasktorch
Hasktorch is a Haskell library for tensor math and neural networks, built on bindings to the C++ libtorch libraries that power PyTorch. It …
761211active
datawhalechina/torch-rechub
Torch-RecHub is a lightweight PyTorch framework for building recommendation system models with 30+ out-of-the-box algorithms covering ranki…
941207active
DachunKai/EvTexture
Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e…
541207active
willisma/SiT
Official PyTorch implementation of Scalable Interpolant Transformers (SiT), a family of generative models built on Diffusion Transformers t…
531206active
metavoiceio/metavoice-src
MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in…
264205maintenance
Artelnics/opennn
OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o…
971198active
openvinotoolkit/nncf
NNCF is Intel's Neural Network Compression Framework, a Python library providing post-training and training-time compression algorithms (qu…
941197active
Calamari-OCR/calamari
Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning…
741197active
autonomousvision/stylegan-t
Official training code for StyleGAN-T, an ICML 2023 paper on fast large-scale text-to-image synthesis using GANs. It provides dataset prepa…
311197active
DeepRec-AI/DeepRec
DeepRec is a high-performance deep learning framework for recommendation models, built on TensorFlow 1.15 with Intel and NVIDIA TensorFlow …
241197active
martinpacesa/BindCraft
BindCraft is a Python-based computational pipeline for de novo protein binder design that combines AlphaFold2 backpropagation, ProteinMPNN,…
781196active
haidog-yaqub/MeanFlow
An unofficial PyTorch implementation of MeanFlow and iMF, one-step generative modeling methods based on flow matching. It provides config-d…
591196active
deepseek-ai/DeepSeek-VL
DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf…
254175maintenance
EvolvingLMMs-Lab/LLaVA-OneVision-2
A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis…
721195active
facebookresearch/esm
Meta FAIR's Evolutionary Scale Modeling (ESM) library providing Transformer protein language models with pretrained weights, including ESM-…
104170maintenance
Cerebras/modelzoo
Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and …
771193active
Tencent-Hunyuan/HunyuanWorld-Mirror
HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds…
541191active
shallowdream204/DreamClear
DreamClear is a diffusion-transformer based real-world image restoration model for high-fidelity super-resolution, published at NeurIPS 202…
271191active
a-r-j/graphein
Graphein is a Python library for constructing graph and mesh representations of proteins, RNA, molecules, and biological interaction networ…
761190active
ali-vilab/UniAnimate
UniAnimate is the official code for a research paper on animating a reference human image into a video that follows a driving pose sequence…
311189active
SHI-Labs/Neighborhood-Attention-Transformer
Official PyTorch implementation of the Neighborhood Attention Transformer (NAT/DiNAT), a family of efficient vision transformers with local…
321184stable
mlmed/torchxrayvision
TorchXRayVision is an open-source PyTorch library providing pre-trained deep learning models and a unified interface for publicly available…
911183active
chensjtu/GaussianObject
GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,…
261183active
msracver/Deformable-ConvNets
Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling …
324121maintenance
XPixelGroup/DiffBIR
DiffBIR is a blind image restoration framework that uses generative diffusion priors to restore degraded real-world images. It provides pre…
344119maintenance
mlfoundations/open_flamingo
OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima…
234118maintenance
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371179active
ubicomplab/rPPG-Toolbox
rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and…
511178active
NVlabs/Deep_Object_Pose
NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca…
481178active
balancap/SSD-Tensorflow
A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net…
324101maintenance
AlgRUC/JittorGeometric
JittorGeometric is a graph machine learning library built on the Jittor deep learning framework, providing implementations of 40+ Graph Neu…
621177active
tjiiv-cprg/EPro-PnP
EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc…
411175stable
princeton-nlp/MeZO
MeZO is a memory-efficient zeroth-order optimizer that fine-tunes language models using only forward passes, with the same memory footprint…
291173stable
bytedance/1d-tokenizer
A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i…
291172active
csguoh/MambaIR
MambaIR and MambaIRv2 are PyTorch-based image restoration models built on Mamba state-space models, published at ECCV 2024 and CVPR 2025. T…
541171active
magicleap/SuperGluePretrainedNetwork
SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t…
324072maintenance
baidu-research/warp-ctc
A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter…
324069maintenance
iver56/torch-audiomentations
A PyTorch library for fast audio data augmentation, inspired by audiomentations. It provides GPU-accelerated, differentiable audio transfor…
581167active
SuperBruceJia/EEG-DL
EEG-DL is a deep learning library built on TensorFlow for classifying EEG signals, supporting many architectures including CNNs, RNNs, GCNs…
471167active
nv-tlabs/LLaMA-Mesh
LLaMA-Mesh is a fine-tuned large language model from NVIDIA Research that generates and understands 3D meshes by representing vertex coordi…
281166active
sksq96/pytorch-summary
A PyTorch library providing a Keras-style model.summary() that prints layer types, output shapes, parameter counts, and memory estimates. I…
324053maintenance
facebookresearch/VideoPose3D
A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori…
104052maintenance
yosinski/deep-visualization-toolbox
A GUI toolbox for visualizing and understanding deep neural networks, showing per-unit activations, backprop/deconv, and regularized-optimi…
324051maintenance
Text-to-Audio/AudioLCM
AudioLCM is a PyTorch implementation of an ACM-MM'24 paper for efficient, high-quality text-to-audio generation using latent consistency mo…
371165active
thunlp/OpenKE
OpenKE is an open-source PyTorch-based toolkit for knowledge graph embedding (knowledge representation learning), with C++ accelerated data…
324047maintenance
majianjia/nnom
NNoM is a high-level neural network inference library written in C for microcontrollers. It converts Keras models into optimized on-device …
231164stable
CyberAgentAILab/TANGO
TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion …
391163active
facebookresearch/encodec
EnCodec is a deep learning based neural audio codec from Meta AI that compresses mono 24 kHz and stereo 48 kHz audio to bitrates from 1.5 t…
324041maintenance
3D ResNets for Action Recognition
A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ…
234038maintenance
MCG-NKU/E2FGVI
E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It…
321161stable
minivision-ai/photo2cartoon
A Python deep-learning project from Minivision that converts real portrait photos into cartoon-style avatars using unpaired image translati…
324029maintenance
LibCity/Bigscity-LibCity
LibCity is an open-source PyTorch library for urban spatial-temporal data mining, providing a unified pipeline for traffic prediction resea…
231157active
fundamentalvision/Deformable-DETR
Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl…
324015maintenance
sirius-ai/LPRNet_Pytorch
A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus…
321156stable
gemelo-ai/vocos
Vocos is a fast neural vocoder that synthesizes audio waveforms from acoustic features such as mel-spectrograms or EnCodec tokens. It uses …
651155stable
deepmodeling/Uni-Mol
Uni-Mol is a collection of 3D molecular representation learning frameworks and pretrained models for tasks like molecule property predictio…
341155active
JunMa11/SegLossOdyssey
A curated collection of loss functions for medical image segmentation, accompanying the 'Loss Odyssey in Medical Image Segmentation' survey…
324007maintenance
SystemErrorWang/White-box-Cartoonization
Official TensorFlow implementation of the CVPR 2020 paper 'Learning to Cartoonize Using White-box Cartoon Representations', which converts …
614001maintenance
TJU-Aerial-Robotics/YOPO
YOPO is a learning-based one-stage planner for quadrotor autonomous navigation in obstacle-dense environments, integrating perception, mapp…
791153active
OpenGVLab/VisionLLM
VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc…
331153active
quark0/darts
DARTS is the official PyTorch implementation of the ICLR 2019 paper 'DARTS: Differentiable Architecture Search', which performs neural arch…
323997maintenance
TensorSpeech/TensorFlowTTS
TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me…
233995maintenance
amazon-science/mm-cot
Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw…
313985maintenance
JDAI-CV/fast-reid
FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. …
233981maintenance
open-gigaai/giga-world-1
GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for…
541147active
ShiqiYu/OpenGait
OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models …
671146active
ScorpioLea/AiCE
AiCE is a Python tool that predicts high-fitness protein mutations by sampling sequences from protein inverse folding models such as Protei…
401144active
HengyiWang/spann3r
Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.…
261141active
cvg/glue-factory
Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point…
691140active
THUMNLab/AutoGL
AutoGL is an autoML framework and toolkit for machine learning on graphs, built on PyTorch with PyTorch Geometric and DGL backends. It prov…
471140active
rohitgandikota/sliders
Official implementation of Concept Sliders, LoRA adaptors that enable precise, plug-and-play control of attributes in diffusion models like…
521139active
clovaai/deep-text-recognition-benchmark
Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio…
323942maintenance

← prev page 11 / 27 next →