Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: deep-learning

2653 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ML-GSAI/LLaDA
Official PyTorch implementation of LLaDA, a family of large language diffusion models (8B base/instruct, MoE, and iLLaDA variants) with pre…
623943active
ali-vilab/VACE
VACE is the official implementation of an all-in-one video creation and editing model from Tongyi Lab, built on Wan2.1 diffusion models. It…
413934active
thuml/Transfer-Learning-Library
TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge…
233931active
Tencent-Hunyuan/Hunyuan3D-2.1
Tencent's open-source 3D asset generation model that creates high-fidelity 3D meshes with production-ready PBR materials from single images…
403917active
hustvl/Vim
Vision Mamba (Vim) is a PyTorch implementation of a generic vision backbone built on bidirectional Mamba state space models, published at I…
293899active
Plachtaa/seed-vc
Seed-VC is a Python tool and model for zero-shot voice conversion, real-time voice conversion, and singing voice conversion, cloning a voic…
103888active
safetensors/safetensors
Safetensors is a simple, secure file format and library for storing and distributing tensors, designed as a fast zero-copy alternative to p…
913876stable
NVlabs/VILA
VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d…
573857active
open-mmlab/mmpretrain
MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv…
233850active
Stability-AI/stable-audio-tools
Stability AI's training and inference toolkit for conditional audio generation models, including Stable Audio Open. It supports training cu…
733849active
shenweichen/GraphEmbedding
A Python library providing implementations of classic graph embedding algorithms including DeepWalk, LINE, Node2Vec, SDNE, and Struc2Vec. I…
683845active
neuraloperator/neuraloperator
A PyTorch library for learning neural operators, which map between function spaces rather than finite-dimensional vectors. It provides the …
823836active
google-research/scenic
Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr…
763821active
ddbourgin/numpy-ml
numpy-ml is a collection of machine learning models and algorithms implemented exclusively in NumPy and the Python standard library, coveri…
3216330maintenance
lightly-ai/lightly
LightlySSL is a Python library built on PyTorch for self-supervised learning on images, offering modular implementations of methods like Si…
933797active
google/deepvariant
DeepVariant is a deep learning-based genomic variant caller that converts aligned DNA sequencing reads (BAM/CRAM) into pileup image tensors…
703791stable
SandAI-org/MAGI-1
MAGI-1 is an open-source autoregressive video generation model from Sand.ai, released with Apache-2.0 licensed code and weights. It generat…
593772active
HeartMuLa/heartlib
HeartMuLa is a family of open-source music foundation models that generate music conditioned on lyrics and tags with multilingual support. …
503749active
fudan-generative-vision/hallo2
Hallo2 is a Python research library from Fudan University that animates a single portrait image using audio input, producing long-duration …
263734active
genmoai/mochi
Mochi 1 is Genmo's open-source, state-of-the-art text-to-video generation model released under Apache 2.0, with a Python API, CLI, and Grad…
463713active
microsoft/Bringing-Old-Photos-Back-to-Life
The official PyTorch implementation of 'Bringing Old Photos Back to Life' (CVPR 2020 Oral), a deep learning model that restores old photos …
2315704maintenance
facebookresearch/map-anything
MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end …
773682active
ferdous-alam/GenCAD
GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif…
363669active
HazyResearch/ThunderKittens
ThunderKittens is a C++/CUDA framework of tile-based primitives for writing fast deep learning GPU kernels. It embeds natively into CUDA so…
703659active
shitagaki-lab/see-through
A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant…
583641active
kaldi-asr/kaldi
Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea…
5215469maintenance
opendilab/DI-engine
DI-engine is an open-source reinforcement learning framework from OpenDILab that provides comprehensive implementations of deep RL algorith…
483638active
cmusatyalab/openface
OpenFace is a free and open source Python and Torch implementation of face recognition based on Google's FaceNet deep neural network. It ge…
6515438maintenance
facebookresearch/detr
DETR is Facebook Research's PyTorch implementation of Detection Transformer, an end-to-end object detection model that replaces hand-crafte…
1015354maintenance
facebookresearch/sam-audio
SAM-Audio is Meta's foundation model for isolating any sound in audio using text, visual, or temporal prompts. This repository provides inf…
553612active
albumentations-team/albumentations
Albumentations is a fast, flexible Python image augmentation library for computer vision, supporting images, masks, bounding boxes, keypoin…
1015315maintenance
AI4Finance-Foundation/FinRL-Trading
FinRL-X is an open-source, AI-native modular infrastructure for quantitative trading that unifies data processing, strategy composition, ba…
713592active
ZhaoJ9014/face.evoLVe
A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and…
383589active
AiuniAI/Unique3D
Unique3D is the official implementation of a NeurIPS 2024 paper that generates high-quality textured 3D meshes from a single image in about…
383579active
GVCLab/PersonaLive
PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin…
533552active
sdv-dev/SDV
SDV (Synthetic Data Vault) is a Python library for generating synthetic tabular data using machine learning models ranging from GaussianCop…
993549stable
AliaksandrSiarohin/first-order-model
Official PyTorch/Jupyter implementation of the First Order Motion Model for image animation (NeurIPS 2019). It animates a static source ima…
3215015maintenance
starVLA/starVLA
StarVLA is an open-source, Lego-like modular codebase for developing Vision-Language-Action (VLA) models for generalist robots. It unifies …
693530active
google-research/big_vision
Google Research's official Jax/Flax codebase for training large-scale vision models such as Vision Transformer, SigLIP, MLP-Mixer, and LiT …
423528active
cszn/KAIR
A PyTorch image restoration toolbox providing training and testing code for many restoration models including DnCNN, FFDNet, SRMD, USRNet, …
233523active
neonbjb/tortoise-tts
Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines…
3214870maintenance
pathwaycom/bdh
BDH (Dragon Hatchling) is a biologically inspired large language model architecture that bridges deep learning and neuroscience, implemente…
543519active
NVlabs/FoundationPose
FoundationPose is NVIDIA's unified foundation model for 6D object pose estimation and tracking of novel objects, supporting both model-base…
623516active
MooreThreads/Moore-AnimateAnyone
An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi…
263514active
google-deepmind/alphafold
Open-source implementation of the AlphaFold 2 inference pipeline for predicting protein structures from amino acid sequences, including Alp…
5814811maintenance
NVIDIA/TransformerEngine
Transformer Engine is an NVIDIA library for accelerating Transformer model training and inference on NVIDIA GPUs using low-precision format…
993504active
PKU-YuanGroup/Video-LLaVA
Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into…
273500active
autorope/donkeycar
Donkeycar is an open-source Python library and hardware platform for building small-scale self-driving RC cars with Raspberry Pi or Jetson …
883495active
borisdayma/dalle-mini
DALL·E Mini is a Python library and model that generates images from a text prompt, available via pip and hosted on Hugging Face Model Hub.…
2314740maintenance
facebookresearch/ijepa
Official PyTorch implementation of I-JEPA, a self-supervised learning method that predicts latent representations of image regions from oth…
103489active
NVIDIA/Model-Optimizer
NVIDIA Model Optimizer (ModelOpt) is a Python library of state-of-the-art model optimization techniques including quantization, pruning, di…
913488active
guandeh17/Self-Forcing
Official implementation of Self Forcing, a training method for autoregressive video diffusion models that simulates inference during traini…
373488active
POSTECH-CVLab/PyTorch-StudioGAN
PyTorch-StudioGAN is a PyTorch library providing unified implementations of representative GAN architectures (BigGAN, StyleGAN2/3, etc.) fo…
233487stable
Tencent-Hunyuan/Hunyuan3D-1
Tencent Hunyuan3D-1.0 is an open-source two-stage diffusion-based model for generating 3D assets from text prompts or images. It provides i…
463482active
MiniMax-AI/MiniMax-01
Official repository for MiniMax-Text-01 and MiniMax-VL-01, open-weight large language and vision-language models built on a linear attentio…
343466active
NVlabs/Eagle
Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L…
643462active
shenweichen/DeepCTR-Torch
DeepCTR-Torch is a PyTorch library providing easy-to-use, modular, and extendable implementations of deep-learning-based CTR (click-through…
783450active
XinJingHao/DRL-Pytorch
A unified PyTorch implementation collection of popular deep reinforcement learning algorithms including DQN variants, PPO, DDPG, TD3, SAC, …
443436active
NVlabs/stylegan
The official TensorFlow implementation of StyleGAN, NVIDIA's style-based generator architecture for generative adversarial networks from th…
3214416maintenance
aqlaboratory/openfold
OpenFold is a faithful, trainable PyTorch reproduction of DeepMind's AlphaFold 2 for protein structure prediction. It is memory-efficient a…
483420active
microsoft/nni
NNI (Neural Network Intelligence) is an open-source AutoML toolkit from Microsoft that automates hyperparameter tuning, neural architecture…
1014361maintenance
davidsandberg/facenet
A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC…
3214343maintenance
WongKinYiu/yolov7
Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques…
2314139maintenance
CompVis/latent-diffusion
The official research code and pretrained model zoo for Latent Diffusion Models (LDM), the paper behind Stable Diffusion, enabling high-res…
3214133maintenance
nv-tlabs/kimodo
Kimodo is NVIDIA's official implementation of a kinematic motion diffusion model trained on 700 hours of motion capture data to generate hi…
563365active
OpenTalker/SadTalker
SadTalker is a CVPR 2023 deep learning tool that generates realistic talking head videos from a single portrait image and an audio clip by …
2214040maintenance
VainF/Torch-Pruning
Torch-Pruning is a PyTorch framework for structural neural network pruning based on the DepGraph algorithm from CVPR 2023. It automatically…
503348active
magenta/ddsp
DDSP is a Python library of differentiable digital signal processing components (synthesizers, filters, waveshapers) that can be embedded i…
643344active
opengeos/geoai
GeoAI is a Python package that integrates artificial intelligence with geospatial data analysis, built on PyTorch, Transformers, and segmen…
893327active
Peterande/D-FINE
D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as…
673305active
microsoft/LoRA
loralib is the official PyTorch implementation of LoRA (Low-Rank Adaptation), which fine-tunes large language models by injecting trainable…
2313767maintenance
Tencent-Hunyuan/HunyuanImage-3.0
HunyuanImage-3.0 is Tencent's open-source native multimodal model for text-to-image and image-to-image generation, with inference code and …
573253active
Beckschen/TransUNet
Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat…
633234stable
mit-han-lab/bevfusion
BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re…
103230stable
Jittor/jittor
Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w…
673229active
onnx/onnx-tensorrt
A C++ parser library and backend that converts ONNX models into TensorRT engines for high-performance GPU inference. It is maintained by NV…
923228active
MzeroMiko/VMamba
VMamba is a PyTorch implementation of a visual state space model (SSM) vision backbone based on Mamba, featuring 2D Selective Scan (SS2D) f…
213219active
kerlomz/captcha_trainer
A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren…
553213active
jy0205/Pyramid-Flow
Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m…
223208active
NVIDIA/physicsnemo
NVIDIA PhysicsNeMo is an open-source Python deep-learning framework for building, training, fine-tuning, and inferring physics AI models us…
893198active
Pointcept
Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan…
763196active
facebookresearch/dinov2
PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual…
6813266maintenance
LeelaChessZero/lc0
Lc0 is an open-source, UCI-compliant chess engine that plays chess using neural networks trained via AlphaZero-style self-play reinforcemen…
653193active
stepfun-ai/Step-Video-T2V
Step-Video-T2V is an open-source text-to-video generation model from StepFun, released with inference code and pretrained weights (includin…
253187active
MiniMax-AI/MiniMax-M1
MiniMax-M1 is an open-weight, large-scale hybrid-attention reasoning language model released by MiniMax under Apache-2.0. The repository pr…
323180active
Rudrabha/Wav2Lip
Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In …
4513182maintenance
facebookresearch/tribev2
TRIBE v2 is a multimodal deep learning model from Meta AI that predicts fMRI brain responses to naturalistic video, audio, and text stimuli…
553172active
ali-vilab/VGen
VGen is the official repository for a holistic video generation ecosystem built on diffusion models, including the I2VGen-XL cascaded image…
273155active
megvii-research/NAFNet
NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. …
323148stable
junyanz/CycleGAN
A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I…
3212870maintenance
thuml/Time-Series-Library
TSLib is an open-source Python library providing a unified codebase of advanced deep learning models for general time series analysis. It s…
6612785maintenance
guillaume-be/rust-bert
A Rust-native library providing ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT-2, RoBERTa, BART, etc.), por…
603076active
imbue-bit/AlphaGPT
AlphaGPT is an open-source automated factor factory based on deep reinforcement learning for quantitative finance. It mines and generates a…
553073active
tensorflow/tflite-micro
TensorFlow Lite for Microcontrollers (TFLM) is a C++ port of TensorFlow Lite for running ML models on microcontrollers, DSPs, and other mem…
773059active
RosettaCommons/RFdiffusion
RFdiffusion is an open-source method for de novo protein structure generation using diffusion models, with or without conditional informati…
623026active
deepseek-ai/DreamCraft3D
Official PyTorch implementation of DreamCraft3D, an ICLR 2024 hierarchical 3D content generation method that turns a single 2D image into a…
353021stable
osmr/imgclsmob
A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti…
233016active
ZQPei/deep_sort_pytorch
A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id…
323012active
benedekrozemberczki/pytorch_geometric_temporal
PyTorch Geometric Temporal is a temporal (dynamic) extension library for PyTorch Geometric providing spatiotemporal signal processing with …
682992active
MeiGen-AI/MultiTalk
MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a …
562992active

← prev page 4 / 27 next →