Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: deep-learning

2653 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
AlmondGod/tinyworlds
A minimal Python implementation of DeepMind's Genie autoregressive world model, including a video tokenizer, action tokenizer, and dynamics…
541378active
keyu-tian/SparK
SparK is the official PyTorch implementation of an ICLR 2023 Spotlight paper that applies BERT/MAE-style masked image modeling to convoluti…
221376stable
qubvel/segmentation_models
A Python library providing neural network architectures for image segmentation (Unet, FPN, Linknet, PSPNet) built on Keras and TensorFlow K…
234923maintenance
minimaxir/textgenrnn
A Python 3 library built on Keras/TensorFlow for easily training char-rnn style neural networks that generate text from any dataset in a fe…
234922maintenance
Meituan-AutoML/MobileVLM
MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small…
171370active
pnnl/neuromancer
NeuroMANCER is a PyTorch-based differentiable programming library for solving parametric constrained optimization problems, physics-informe…
721369active
OpenPPL/ppl.nn
PPLNN is a high-performance deep-learning inference engine written in C++ that runs ONNX models on x86 CPUs and NVIDIA GPUs, with a dedicat…
321367active
xandergos/terrain-diffusion
Terrain Diffusion is a Python framework that uses diffusion models as a learned, deterministic replacement for Perlin noise, generating inf…
641365active
ImprintLab/MedSegDiff
MedSegDiff is a diffusion probabilistic model framework for segmenting and reconstructing organs and tissues from medical images, with a tr…
501363active
hustvl/VAD
VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem…
601362active
yuantianyuan01/FastWAM
Official PyTorch codebase for Fast-WAM, a World Action Model for robot manipulation that skips test-time future video imagination and gener…
591362active
bytedance/UNO
UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul…
381362active
MoonInTheRiver/DiffSinger
Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m…
654851maintenance
huggingface/finetrainers
finetrainers is a Hugging Face library for scalable, memory-optimized training (fine-tuning) of diffusion models, including LoRA training o…
621358active
Sense-X/Co-DETR
Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg…
321357stable
blei-lab/edward
Edward is a Python library for probabilistic modeling, inference, and criticism built on TensorFlow. It supports deep generative models, va…
234843maintenance
reiniscimurs/DRL-robot-navigation
A ROS Gazebo simulation project that trains a mobile robot to navigate to random goals while avoiding obstacles using a TD3 deep reinforcem…
581356active
mega-sam/mega-sam
MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua…
481355active
k2-fsa/k2
k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)…
641352active
shivammehta25/Matcha-TTS
Matcha-TTS is a PyTorch-based text-to-speech system that uses conditional flow matching for fast, non-autoregressive speech synthesis. It s…
621349active
wyhuai/DDNM
DDNM is a Python research codebase implementing the Denoising Diffusion Null-Space Model for zero-shot image restoration, published as an I…
321349stable
MegEngine/MegEngine
MegEngine is a fast, scalable deep learning framework with automatic differentiation, developed in C++ with Python bindings. It unifies tra…
234808maintenance
sjvasquez/handwriting-synthesis
A Python implementation of Alex Graves' handwriting synthesis experiments using recurrent neural networks, generating realistic handwritten…
324802maintenance
PKU-VCL-3DV/SLAM3R
SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw…
421344active
muzishen/IMAGDressing
IMAGDressing-v1 is a diffusion-based framework for customizable virtual dressing that generates human images with fixed garments and contro…
441343active
FireRedTeam/FireRed-Image-Edit
FireRed-Image-Edit is an open-source image editing foundation model built on diffusion models, released as PyTorch model weights with infer…
491341active
alibaba/graph-learn
Graph-Learn (formerly AliGraph) is a distributed framework for developing and applying large-scale graph neural networks, with a training l…
361341active
rstudio/tensorflow
An R package that provides full access to the TensorFlow API from R via reticulate, bridging R users to TensorFlow's Python implementation.…
611339active
PKU-YuanGroup/MagicTime
MagicTime is a metamorphic time-lapse video generation pipeline built on diffusion-based text-to-video models, with a MagicAdapter, dynamic…
591338active
IrisRainbowNeko/genshin_auto_fish
A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen…
234758maintenance
liwenxi/SWIFT-AI
SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st…
291334active
mapillary/inplace_abn
A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing…
651333stable
jonathan-laurent/AlphaZero.jl
A generic, simple, and fast Julia implementation of DeepMind's AlphaZero algorithm for training game-playing agents via self-play and MCTS.…
641333active
wenqsun/DimensionX
DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo…
431333active
bytedance/Lance
Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained …
551329active
facebookincubator/AITemplate
AITemplate is a Python framework that compiles deep neural networks into high-performance CUDA (NVIDIA) or HIP (AMD) C++ code for fast fp16…
664724maintenance
LLaVA-VL/LLaVA-NeXT
LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision …
644716maintenance
ACEsuit/mace
MACE is a Python library implementing fast and accurate machine learning interatomic potentials using higher-order equivariant message pass…
891324active
agemagician/ProtTrans
ProtTrans provides state-of-the-art pre-trained Transformer language models for protein sequences, trained on thousands of GPUs and hundred…
331324active
sjtuytc/UnboundedNeRFPytorch
A PyTorch implementation benchmarking state-of-the-art unbounded (large-scale) neural radiance field methods like NeRF++, DVGO, and Block-N…
231324active
tensorflow/lucid
Lucid is a collection of infrastructure and tools for research in neural network interpretability, built on TensorFlow 1.x. It provides fea…
104704maintenance
galilai-group/lejepa
LeJEPA is a Python framework for scalable, theoretically grounded self-supervised representation learning based on Joint-Embedding Predicti…
451322active
ImprintLab/Medical-SAM-Adapter
Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei…
391322active
stared/livelossplot
A Python library that draws live training loss and metric plots inside Jupyter Notebooks for Keras, PyTorch, and other deep learning framew…
911319stable
open-edge-platform/geti
Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru…
981317active
sicara/easy-few-shot-learning
A Python library (easyfsl) with ready-to-use code and tutorial notebooks for few-shot image classification and meta-learning, built on PyTo…
231313stable
yeyupiaoling/VoiceprintRecognition-Pytorch
A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+…
581312active
ali-vilab/MimicBrush
MimicBrush is the official implementation of a zero-shot image editing method that lets users mask a region in a source image and provide a…
241311active
soda-inria/tabicl
TabICLv2 is an open-source tabular foundation model that performs classification and regression via a single forward pass through a pre-tra…
811309active
albermax/innvestigate
iNNvestigate is a Python toolbox providing a common interface and out-of-the-box implementations of many neural network explanation methods…
301309active
huawei-noah/Efficient-Computing
A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu…
321307active
DAMO-NLP-SG/VideoLLaMA2
VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro…
251307active
luosiallen/latent-consistency-model
Official implementation of Latent Consistency Models (LCM), a diffusion-based approach for synthesizing high-resolution images with few-ste…
274615maintenance
STVIR/pysot
PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS…
454600maintenance
donydchen/mvsplat
MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo…
611296active
robodhruv/visualnav-transformer
Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie…
191294active
fundamentalvision/BEVFormer
BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera …
234579maintenance
braindecode/braindecode
Braindecode is an open-source Python toolbox built on PyTorch for decoding raw electrophysiological brain signals such as EEG, ECoG, and ME…
981293active
W2GenAI-Lab/LucidFlux
LucidFlux is a caption-free photo-realistic image restoration model built on a large-scale diffusion transformer, released with inference a…
551293active
plaidml/plaidml
PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m…
104566maintenance
RoyalVane/CLAN
Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation…
321289stable
TencentQQGYLab/ELLA
ELLA is an Efficient Large Language Model Adapter that equips text-to-image diffusion models with LLM-based text understanding via a Timest…
251289active
buoyancy99/diffusion-forcing
Official research code for the NeurIPS paper 'Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion', implementing a metho…
651288active
locuslab/TCN
PyTorch implementation of Temporal Convolutional Networks (TCN) with benchmarks from the paper 'An Empirical Evaluation of Generic Convolut…
324549maintenance
huanngzh/MV-Adapter
MV-Adapter is a plug-and-play adapter that turns pre-trained text-to-image diffusion models (e.g., SDXL, SD2.1) into multi-view consistent …
341285active
Phantom-video/HuMo
HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative …
461283active
e3nn/e3nn
e3nn is a Python/PyTorch library for building E(3)-equivariant neural networks that respect 3D rotation, translation, and mirror symmetries…
701280active
ZhengyiLuo/PHC
Official implementation of the ICCV 2023 paper 'Perpetual Humanoid Control for Real-time Simulated Avatars'. It provides a Python codebase …
441280active
AaronJackson/vrn
Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses…
324517maintenance
plemeri/transparent-background
A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima…
631278active
suragnair/alpha-zero-general
A clean, flexible implementation of the AlphaZero self-play reinforcement learning algorithm that can be adapted to any two-player turn-bas…
324506maintenance
studio-dots-ai/dots.tts
dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain…
791275active
Ma-Lab-Berkeley/CRATE
CRATE is the official PyTorch implementation of the Coding RAte reduction TransformEr, a family of 'white-box' transformer architectures de…
291275active
google-research/simclr
Google Research's official implementation of SimCLR and SimCLRv2, a framework for contrastive learning of visual representations, with 65 p…
104502maintenance
Stable-X/Stable3DGen
Stable3DGen is a modular Python framework for generating 3D assets from images, adapted from Microsoft's TRELLIS with NVIDIA library depend…
331274active
DreamTechAI/Direct3D-S2
Direct3D-S2 is a research framework for high-resolution 3D shape generation from images, built on sparse volumetric representations and a n…
291274active
dcharatan/pixelsplat
pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives…
271274stable
nianticlabs/monodepth2
Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra…
324497maintenance
microsoft/BioGPT
BioGPT is Microsoft's domain-specific generative Transformer language model pre-trained on biomedical text, with implementation code and pr…
324488maintenance
NVlabs/stylegan2-ada-pytorch
Official PyTorch implementation of StyleGAN2-ADA, a generative adversarial network with adaptive discriminator augmentation for training wi…
324487maintenance
leoxiaobin/deep-high-resolution-net.pytorch
Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai…
324480maintenance
luo3300612/Visualizer
A lightweight Python library that extracts attention maps and other local variables from deep inside PyTorch models for visualization. It w…
321269stable
lucidrains/flamingo-pytorch
A PyTorch implementation of DeepMind's Flamingo visual language model architecture, providing the Perceiver Resampler and Gated Cross-Atten…
231269active
amaiya/ktrain
ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul…
251268active
thunlp/OpenNRE
OpenNRE is an open-source Python toolkit for neural relation extraction, extracting relation triples between entities from plain text. It u…
324467maintenance
nv-tlabs/Difix3D
Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau…
321266active
bryandlee/animegan2-pytorch
A PyTorch implementation of AnimeGANv2, a GAN-based image-to-image style transfer model that converts photos into anime-style images. It pr…
324452maintenance
rlawjdghek/StableVITON
StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent…
471261stable
Fictionarry/ER-NeRF
ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr…
241260stable
nv-tlabs/GET3D
GET3D is NVIDIA's PyTorch implementation of a generative model that synthesizes high-quality 3D textured meshes (cars, chairs, animals, bui…
324435maintenance
Francis-Rings/StableAvatar
StableAvatar is an end-to-end video diffusion transformer that generates infinite-length, high-quality talking avatar videos from a referen…
461258active
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2…
581256active
stardist/stardist
StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built …
601255stable
ingra14m/Deformable-3D-Gaussians
Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. …
181255stable
thu-ml/Motus
Motus is the official implementation of a unified latent action world model for robotics, combining a video generation model, a vision-lang…
431246active
LTH14/fractalgen
A PyTorch implementation of Fractal Generative Models (FractalGen), enabling pixel-by-pixel high-resolution image generation. It includes p…
241244active
3DTopia/OpenLRM
OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide…
171244active
HJYao00/Mulberry
Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect…
491243active
cvg/depthsplat
DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim…
561242active
showlab/Tune-A-Video
Tune-A-Video is the official PyTorch implementation of an ICCV 2023 paper that fine-tunes pre-trained text-to-image diffusion models (like …
314364maintenance

← prev page 10 / 27 next →