Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: deep-learning

2771 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
huggingface/swift-coreml-diffusers
A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff…
542757active
CVCUDA/CV-CUDA
CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs…
932723active
magic-research/magic-animate
MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image …
4410898maintenance
KimMeen/Time-LLM
Time-LLM is the official PyTorch implementation of an ICLR 2024 paper that reprograms frozen large language models (Llama, GPT-2, BERT) for…
472689active
princeton-vl/DROID-SLAM
DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense …
412678active
aigc3d/LHM
LHM is a PyTorch-based large reconstruction model that reconstructs high-fidelity animatable 3D human avatars from a single image in second…
522664active
HiDream-ai/HiDream-I1
HiDream-I1 is an open-source 17B-parameter text-to-image generative foundation model based on a Sparse Diffusion Transformer, with full and…
342515active
wgsxm/PartCrafter
PartCrafter is a structured 3D generative model that jointly generates multiple semantically meaningful 3D mesh parts and objects from a si…
532472active
OpenDCAI/DataFlex
DataFlex is a data-centric training framework built on LLaMA-Factory that dynamically schedules LLM training data during optimization. It p…
672432active
HybridRobotics/whole_body_tracking
BeyondMimic's official motion tracking training code, a humanoid control framework built on Isaac Lab that trains sim-to-real-ready whole-b…
692367active
NVlabs/ProtoMotions
ProtoMotions is a GPU-accelerated simulation and reinforcement learning framework for training physically simulated digital humans and huma…
662362active
NVlabs/nvdiffrec
nvdiffrec is NVIDIA's official implementation of a CVPR 2022 oral paper that jointly optimizes triangular 3D meshes, PBR materials, and lig…
762296stable
togethercomputer/OpenChatKit
OpenChatKit is an open-source toolkit from Together, LAION, and Ontocord.ai for building and fine-tuning chat language models, including in…
308983maintenance
nv-tlabs/lyra
Project Lyra is NVIDIA's open series of generative 3D world models, including Lyra 1.0 for feed-forward 3D/4D scene generation from a singl…
592279active
langfengQ/verl-agent
verl-agent is an extension of the veRL framework for training LLM and VLM agents via reinforcement learning, featuring step-independent mul…
562276active
Liuziyu77/Visual-RFT
Official research code for Visual-RFT and Visual-ARFT, applying GRPO-based reinforcement fine-tuning with rule-based verifiable rewards to …
422274active
OlafenwaMoses/ImageAI
ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob…
238882maintenance
jd-opensource/JoyAI-Image
JoyAI-Image is a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editin…
582149active
PKU-YuanGroup/LLaVA-CoT
LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea…
472132active
fastmachinelearning/hls4ml
hls4ml is a Python package that converts machine learning models from Keras, PyTorch, and ONNX into high-level synthesis (C++) code for FPG…
822124active
facebookresearch/theseus
Theseus is a PyTorch-based library for building custom differentiable nonlinear optimization layers, supporting problems in robotics and vi…
232056active
serengil/retinaface
RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin…
612032active
sicxu/Deep3DFaceRecon_pytorch
A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag…
321911stable
diffgram/diffgram
Diffgram is a self-hosted AI datastore for managing schemas, BLOBs, and predictions, with built-in human supervision (data labeling), data …
621909active
PRIME-RL/SimpleVLA-RL
SimpleVLA-RL is an open-source reinforcement learning framework for training Vision-Language-Action (VLA) models for robotic manipulation, …
461845active
YangLing0818/RPG-DiffusionMaster
Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and …
281844active
jd-opensource/JoyAI-VL-Interaction
JoyAI-VL-Interaction is an open 8B-scale vision-language interaction model with a complete deployable real-time streaming system, including…
581840active
Zheng-Chong/CatVTON
CatVTON is a lightweight diffusion model for virtual try-on that swaps clothing onto a person image using a concatenation-based architectur…
401830active
Robbyant/lingbot-vla
LingBot-VLA is a Vision-Language-Action foundation model for robot manipulation, pretrained on 20,000 hours of real-world dual-arm robot da…
541794active
Emu Series
Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid…
571779active
GAP-LAB-CUHK-SZ/gaustudio
GauStudio is a modular PyTorch framework for 3D Gaussian Splatting (3DGS) research and development, supporting novel view synthesis, 3D rec…
491762active
One-2-3-45/One-2-3-45
One-2-3-45 is the official PyTorch implementation of a NeurIPS 2023 paper that converts any single image into a full 360-degree 3D textured…
291718stable
williamyang1991/DualStyleGAN
Official PyTorch implementation of DualStyleGAN, a CVPR 2022 model for exemplar-based high-resolution (1024px) portrait style transfer. It …
321683stable
Stability-AI/stable-virtual-camera
Stable Virtual Camera (SEVA) is a generalist diffusion model for novel view synthesis that generates 3D-consistent views of a scene from an…
541657active
ZJU4HealthCare/HealthGPT
HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled…
631655active
guochengqian/Magic123
Magic123 is the official PyTorch implementation of an ICLR 2024 paper that generates high-quality textured 3D meshes from a single unposed …
901623stable
PKU-Alignment/safe-rlhf
Beaver is a modular open-source framework from Peking University for training large language models with SFT, RLHF, and Safe RLHF (constrai…
531613active
zju3dv/EasyVolcap
EasyVolcap is a PyTorch-based library for accelerating neural volumetric video research, covering volumetric video capture, reconstruction,…
271587active
photosynthesis-team/piq
PyTorch Image Quality (PIQ) is a collection of measures and metrics for image quality assessment in image-to-image tasks, written in pure P…
231574stable
tianrun-chen/SAM-Adapter-PyTorch
A PyTorch library that adapts Meta AI's Segment Anything Model (SAM, SAM2, SAM3) to underperforming downstream segmentation tasks using lig…
671551active
RightNow-AI/autokernel
AutoKernel is an open-source autoresearch pipeline that takes any PyTorch model, profiles it to find GPU kernel bottlenecks, extracts them …
481548active
luciddreamer-cvlab/LucidDreamer
LucidDreamer is the official implementation of a research method that generates 3D Gaussian Splatting scenes from text prompts, published i…
691528active
ATH-MaaS/Ovis
Ovis is an open-source Multimodal Large Language Model (MLLM) architecture that structurally aligns visual and textual embeddings, with rel…
651514active
hkchengrex/Tracking-Anything-with-DEVA
DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona…
271509stable
Vincentqyw/cv-arxiv-daily
An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,…
771495active
CSAILVision/semantic-segmentation-pytorch
A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra…
325080maintenance
alibaba/Logics-Parsing
Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul…
541404active
lxtGH/OMG-Seg
Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le…
471354active
CarperAI/trlx
trlX is a distributed training framework for fine-tuning large language models with reinforcement learning from human feedback (RLHF), supp…
234753maintenance
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
deepseek-ai/DeepSeek-Prover-V2
DeepSeek-Prover-V2 is an open-source large language model for formal theorem proving in Lean 4, trained via reinforcement learning with sub…
341298active
donydchen/mvsplat
MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo…
611296active
BytedTsinghua-SIA/CUDA-Agent
CUDA-Agent is a large-scale agentic reinforcement learning system from ByteDance Seed and Tsinghua that trains LLMs to generate high-perfor…
561290active
zju3dv/MatchAnything
MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap…
641284active
Visual-Agent/DeepEyes
DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning…
441273active
nv-tlabs/Difix3D
Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau…
321268active
rlawjdghek/StableVITON
StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent…
471263stable
ziyc/drivestudio
DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It …
401256active
VainF/pytorch-msssim
A PyTorch library providing fast, differentiable SSIM and MS-SSIM image quality metrics using separable Gaussian filtering for speed. It ca…
231253stable
X-Square-Robot/wall-x
Wall-X is the open-source training and inference stack for X Square Robot's WALL series of embodied foundation models (VLAs) for general-pu…
611252active
Linketic/CityGaussian
Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia…
661251active
unum-cloud/UForm
UForm is a compact multimodal AI library providing tiny image-text embedding models (64-768 dimensions, Matryoshka-style) and small generat…
551248active
3DTopia/OpenLRM
OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide…
171246active
cvg/depthsplat
DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim…
561243active
marian42/mesh_to_sdf
A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte…
321242stable
ZHKKKe/MODNet
MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima…
324359maintenance
apchenstu/TensoRF
TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie…
441237stable
EvolvingLMMs-Lab/LLaVA-OneVision-2
A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis…
721197active
deepseek-ai/DeepSeek-VL
DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf…
254177maintenance
Tencent-Hunyuan/HunyuanWorld-Mirror
HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds…
541195active
trianglesplatting/triangle-splatting
Official implementation of 'Triangle Splatting for Real-Time Radiance Field Rendering' (3DV 2026), which uses 3D triangles as rendering pri…
431193active
ubicomplab/rPPG-Toolbox
rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and…
511185active
chensjtu/GaussianObject
GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,…
261184active
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371178active
majianjia/nnom
NNoM is a high-level neural network inference library written in C for microcontrollers. It converts Keras models into optimized on-device …
231164stable
TJU-Aerial-Robotics/YOPO
YOPO is a learning-based one-stage planner for quadrotor autonomous navigation in obstacle-dense environments, integrating perception, mapp…
791159active
OpenGVLab/VisionLLM
VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc…
331154active
princeton-vl/DPVO
DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-…
321114active
jina-ai/discoart
DiscoArt is a Python library that wraps Disco Diffusion (CLIP-guided diffusion) so generative artists and developers can create AI artworks…
233827maintenance
szymanowiczs/splatter-image
Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net…
261105active
yzslab/gaussian-splatting-lightning
A PyTorch Lightning implementation of 3D Gaussian Splatting with many derived algorithms (Mip-Splatting, LightGaussian, 2DGS, deformable Ga…
581103active
Toni-SM/skrl
skrl is an open-source modular Reinforcement Learning library written in Python, implemented in PyTorch, JAX, and NVIDIA Warp. It supports …
811095active
LujiaJin/One-Pot_Multi-Frame_Denoising
Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides …
601094stable
lizhe00/AnimatableGaussians
Official PyTorch implementation of the CVPR 2024 paper 'Animatable Gaussians', which learns pose-dependent Gaussian maps for high-fidelity …
271094active
mlc-ai/web-stable-diffusion
A project that compiles and runs Stable Diffusion text-to-image models entirely inside web browsers using WebGPU and WebAssembly, with no s…
303723maintenance
Stability-AI/stable-point-aware-3d
SPAR3D is Stability AI's open-source model for fast single-image 3D mesh reconstruction using a two-stage pipeline with point cloud conditi…
301070active
X-LANCE/SLAM-LLM
SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc…
551057active
ShenhanQian/GaussianAvatars
Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri…
561054active
InternRobotics/PointLLM
PointLLM is a multimodal large language model that understands colored 3D point clouds of objects, built on a point cloud encoder fused wit…
651053active
HannesStark/boltzgen
BoltzGen is an open-source all-atom generative diffusion model for designing protein and peptide binders against arbitrary biomolecular tar…
711049active
ml-tooling/ml-workspace
ML Workspace is an all-in-one web-based IDE Docker image specialized for machine learning and data science. It bundles Jupyter, JupyterLab,…
233544maintenance
rl-tools/rl-tools
RLtools is a pure C++ header-only, dependency-free deep reinforcement learning library supporting algorithms like SAC, TD3, and PPO. It com…
651033active
HarborYuan/ovsam
Official PyTorch implementation of Open-Vocabulary SAM (ECCV 2024), a model that unifies SAM's interactive segmentation with CLIP's open-vo…
421033active
JackAILab/ConsistentID
ConsistentID is a diffusion-based portrait generation model and toolkit that preserves facial identity from a single reference image using …
521026active
yangxy/PASD
PASD (Pixel-Aware Stable Diffusion) is a Python research codebase implementing an ECCV 2024 method for realistic image super-resolution and…
281022active
Alpha-VLLM/Lumina-DiMOO
Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu…
551020active
SunOner/sunone_aimbot
An AI-powered aimbot for first-person shooter games that uses YOLO object detection models (YOLOv8/v10/v12) with TensorRT/ONNX acceleration…
651014active
TRI-ML/prismatic-vlms
Prismatic VLMs is a PyTorch-based codebase for training visually-conditioned language models (VLMs) with flexible vision backbones like CLI…
251013active
MeiGen-AI/PosterCraft
PosterCraft is a unified framework for generating high-quality aesthetic posters, published as an ICLR 2026 paper. It provides model weight…
481009active
TinyLLaVA/TinyLLaVA_Factory
TinyLLaVA Factory is an open-source modular PyTorch/HuggingFace codebase for training small-scale large multimodal models (LMMs) that combi…
681004active

← prev page 27 / 28 next →