Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: deep-learning

2771 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Lightning-AI/litgpt
LitGPT is a Python library providing from-scratch, hackable implementations of 20+ open-source large language models with recipes for pretr…
9313629active
Physical-Intelligence/openpi
Open-source repository from Physical Intelligence containing vision-language-action (VLA) models for robotics, including π₀, π₀-FAST, and π…
6613494active
NVIDIA/TensorRT
NVIDIA TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, comprising an inference compiler that converts train…
9413293stable
modelscope/DiffSynth-Studio
DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio …
7913003active
PaddlePaddle/PaddleFormers
PaddleFormers is a Transformers-style library built on PaddlePaddle providing a model zoo of 100+ large language models and vision-language…
9212986active
zai-org/CogVideo
CogVideo/CogVideoX is an open-source family of text-to-video and image-to-video generation models from Zhipu AI (THUDM), with inference and…
4512977active
PaddlePaddle/PaddleNLP
PaddleNLP is an easy-to-use NLP and large language model development kit built on the PaddlePaddle deep learning framework, with a large pr…
6312967active
jacobgil/pytorch-grad-cam
A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A…
7612958active
deepseek-ai/FlashMLA
FlashMLA is DeepSeek's library of optimized CUDA attention kernels implementing Multi-head Latent Attention (MLA), including dense and toke…
6212872active
google-research/vision_transformer
Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp…
7512683stable
PaddlePaddle/PaddleSpeech
PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st…
6612670active
sapientinc/HRM
Official PyTorch implementation of the Hierarchical Reasoning Model (HRM), a 27M-parameter recurrent architecture with high-level and low-l…
5212619active
YaoFANGUK/video-subtitle-remover
An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai…
7112553active
bmaltais/kohya_ss
A Gradio-based GUI and CLI wrapper around Kohya's Stable Diffusion training scripts for fine-tuning diffusion image generation models. It s…
9512548active
Tencent-Hunyuan/HunyuanVideo
HunyuanVideo is Tencent's open-source framework for large-scale video generation, providing PyTorch model definitions, pre-trained weights,…
6212476active
axolotl-ai-cloud/axolotl
Axolotl is a free, open-source, config-driven framework for fine-tuning large language models, supporting SFT, preference learning (DPO/KTO…
9512408active
guoyww/AnimateDiff
Official implementation of AnimateDiff, a plug-and-play motion modeling module that turns personalized text-to-image diffusion models (e.g.…
2912227active
xmu-xiaoma666/External-Attention-pytorch
A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame…
6512183active
PKU-YuanGroup/Open-Sora-Plan
Open-Sora Plan is an open-source effort to reproduce OpenAI's Sora text-to-video model, providing training and inference code for video gen…
5112155active
HKUDS/ViMax
ViMax is an agentic video generation framework that uses coordinated multi-agent collaboration (director, screenwriter, producer, video gen…
7212108active
Tongyi-MAI/Z-Image
Z-Image is a 6B-parameter text-to-image generation foundation model family built on a single-stream diffusion transformer, with a distilled…
4611944active
ostris/ai-toolkit
An all-in-one open-source training toolkit for finetuning diffusion models (image and video) on consumer-grade hardware. It supports many r…
7311838active
speechbrain/speechbrain
SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained…
8311785active
ludwig-ai/ludwig
Ludwig is a declarative, low-code deep learning framework for training, fine-tuning, and deploying AI models — from LLMs to tabular, image,…
9911745active
qubvel-org/segmentation_models.pytorch
A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar…
7011706stable
NVIDIA/cosmos
NVIDIA Cosmos is an open platform of omnimodal world foundation models, datasets, and tools for building Physical AI systems such as robots…
7211641active
milesial/Pytorch-UNet
A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva…
2311613active
CorentinJ/Real-Time-Voice-Cloning
A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from …
6460110maintenance
THU-MIG/yolov10
YOLOv10 is a real-time end-to-end object detection model family that removes NMS post-processing via consistent dual assignments and optimi…
2011336active
kornia/kornia
Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri…
8611327active
salesforce/LAVIS
LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including …
6111262active
facebookresearch/dinov3
Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t…
5911249active
Weights & Biases
Weights & Biases (wandb) is a Python SDK and platform for tracking, visualizing, and managing machine learning experiments, including metri…
9911239active
ultralytics/yolov5
Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif…
6757929maintenance
thu-ml/tianshou
Tianshou is a modular, high-performance deep reinforcement learning library built on pure PyTorch and Gymnasium. It offers both low-level h…
7210943active
triton-inference-server/server
NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo…
9810939stable
Lightricks/LTX-Video
Official repository for LTX-Video, a DiT-based open-weights video generation model from Lightricks that generates high-fidelity video (up t…
4810907active
microsoft/TRELLIS.2
TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f…
5710869active
cumulo-autumn/StreamDiffusion
StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti…
1710806active
OpenVINO
OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su…
9510740stable
lucidrains/denoising-diffusion-pytorch
A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d…
8710679active
Megvii-BaseDetection/YOLOX
YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo…
3410587stable
facebookresearch/xformers
xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient …
8810542active
skypilot-org/skypilot
SkyPilot is an open-source AI compute platform that unifies fragmented infrastructure (Kubernetes, Slurm, VMs, 20+ clouds) into a single po…
9710529active
bigscience-workshop/petals
Petals is a Python library that lets you run and fine-tune large language models (Llama 3.1, Mixtral, Falcon, BLOOM) on a BitTorrent-style …
2310521active
vwxyzjn/cleanrl
CleanRL is a deep reinforcement learning library providing high-quality, single-file implementations of algorithms like PPO, DQN, DDPG, TD3…
5810326active
NVIDIA/cutlass
CUTLASS is NVIDIA's collection of CUDA C++ template abstractions and Python DSLs for implementing high-performance GEMM and related linear …
9910317active
open-mmlab/Amphion
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s…
5110271active
xai-org/grok-1
xAI's open release of the Grok-1 open-weights model (314B-parameter Mixture-of-Experts LLM) with JAX example code for loading and running i…
2552189maintenance
TencentARC/PhotoMaker
PhotoMaker is a personalized text-to-image generation method that encodes multiple reference face photos into a stacked ID embedding to gen…
2610088stable
deepseek-ai/DeepEP
DeepEP is a high-performance GPU communication library for expert parallelism (EP) in MoE training and inference, providing high-throughput…
5710066active
facebookresearch/pytorch3d
PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides…
749954active
espnet/espnet
ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance…
889941active
open-mmlab/mmsegmentation
MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat…
239930stable
huggingface/accelerate
Hugging Face Accelerate is a Python library that lets you run the same PyTorch training and inference code on any device or distributed con…
959838stable
arogozhnikov/einops
einops is a Python library providing readable, framework-agnostic tensor operations via mini-language functions like rearrange, reduce, and…
779581stable
WongKinYiu/yolov9
Official PyTorch implementation of the YOLOv9 object detection paper, featuring Programmable Gradient Information for improved accuracy. It…
169551active
modelscope/facechain
FaceChain is a deep-learning toolchain from ModelScope for generating identity-preserved personal portraits (digital twins) from a single p…
309508active
replicate/cog
Cog is an open-source CLI tool that packages machine learning models into production-ready Docker containers using a simple cog.yaml config…
949463active
Oneflow-Inc/oneflow
OneFlow is an open-source deep learning framework written in C++ with a PyTorch-like Python API, focused on scalable and efficient distribu…
489428active
YaoFANGUK/video-subtitle-extractor
A GUI application that extracts hard-coded (burned-in) subtitles from videos and generates SRT subtitle files using local deep-learning-bas…
709401active
oumi-ai/oumi
Oumi is an open-source Python framework and platform for the end-to-end lifecycle of open-weight LLMs: data synthesis, fine-tuning (SFT, Lo…
869373active
keras-team/autokeras
AutoKeras is an AutoML library for deep learning built on Keras, developed by DATA Lab at Texas A&M University. It automates model architec…
529326active
bytedance/monolith
Monolith is a deep learning framework built on TensorFlow for large-scale recommendation modeling. It provides collisionless embedding tabl…
109298active
LTX-2
Official Python package from Lightricks providing inference pipelines and LoRA training for LTX-2/LTX-2.5, an open-weights DiT-based founda…
839260active
lipku/LiveTalking
LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit…
899238active
Deep Lake
Deep Lake is an open-source database for AI that stores multimodal data (images, video, audio, text, embeddings, annotations) in a format o…
779228active
coqui-ai/TTS
Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini…
2345953maintenance
modelscope/modelscope
ModelScope is a Python library and ecosystem built on the 'Model-as-a-Service' concept, providing unified APIs to download, run inference o…
989111active
studio-dots-ai/dots.ocr
dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi…
519090active
roboflow/rf-detr
RF-DETR is a real-time transformer-based model architecture from Roboflow for object detection, instance segmentation, and keypoint detecti…
879063active
pyro-ppl/pyro
Pyro is a deep universal probabilistic programming library built on Python and PyTorch, supporting Bayesian modeling with variational infer…
669037stable
NVIDIA/apex
NVIDIA-maintained PyTorch extension providing utilities for easy mixed precision and distributed training. It offers up-to-date CUDA and C+…
778993active
dusty-nv/jetson-inference
A C++/Python DNN inference library and tutorial guide ('Hello AI World') for deploying deep learning vision models on NVIDIA Jetson devices…
448969stable
sebastianstarke/AI4Animation
AI4Animation is a deep learning framework for data-driven character animation and control, built around Unity with a Python remake (AI4Anim…
678846active
NVlabs/Sana
SANA is an efficiency-oriented PyTorch codebase for high-resolution text-to-image and text-to-video generation built on Linear Diffusion Tr…
748833active
MIC-DKFZ/nnUNet
nnU-Net is a self-configuring deep learning framework for semantic image segmentation that automatically adapts preprocessing, U-Net archit…
658829stable
bentoml/BentoML
BentoML is a Python framework for building online serving systems for AI apps and model inference, turning model inference scripts into RES…
948808stable
FoundationVision/VAR
Official PyTorch implementation of Visual Autoregressive Modeling (VAR), a NeurIPS 2024 Best Paper-winning method for scalable image genera…
488729active
DepthAnything/Depth-Anything-V2
Depth Anything V2 is a foundation model for monocular depth estimation, trained on 595K synthetic labeled images and 62M+ real unlabeled im…
568709stable
fudan-generative-vision/hallo
Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head …
148664active
MONAI
MONAI is a PyTorch-based open-source framework for deep learning in healthcare imaging, providing domain-specific transforms, 3D architectu…
878634stable
netease-youdao/EmotiVoice
EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn…
178523active
lllyasviel/IC-Light
IC-Light is a Python tool for manipulating the illumination of images using diffusion models, offering text-conditioned and background-cond…
278508active
google-deepmind/alphafold3
Google DeepMind's official implementation of the AlphaFold 3 inference pipeline for predicting biomolecular structures and interactions. It…
838495active
OptimalScale/LMFlow
LMFlow is an extensible Python toolkit for finetuning and inference of large foundation models such as LLaMA, GPT-2, and Galactica. It prov…
648484active
bitsandbytes-foundation/bitsandbytes
bitsandbytes is a Python library providing k-bit quantization primitives for PyTorch, enabling 8-bit (LLM.int8()) and 4-bit (QLoRA) quantiz…
998439active
lucidrains/imagen-pytorch
A PyTorch implementation of Imagen, Google's text-to-image neural network based on cascading DDPMs conditioned on T5 text embeddings. It pr…
238424active
XPixelGroup/BasicSR
BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a…
238367stable
mikel-brostrom/boxmot
BoxMOT is a pluggable Python and C++ library providing state-of-the-art multi-object tracking (MOT) algorithms such as ByteTrack, BoT-SORT,…
958281active
QwenLM/Qwen-Image
Qwen-Image is a 20B MMDiT image generation foundation model from the Qwen team, with strong complex text rendering (especially Chinese) and…
488265active
lltcggie/waifu2x-caffe
A Windows GUI/CLI application that reimplements the waifu2x image upscaling and noise-reduction tool using the Caffe deep learning framewor…
568221stable
brycedrennan/imaginAIry
A Python library and CLI tool (imaginairy/aimg) for generating images and videos with Stable Diffusion and Stable Video Diffusion models. I…
638179active
google-research/bert
Google Research's official TensorFlow implementation of BERT, the Bidirectional Encoder Representations from Transformers language model, a…
1040046maintenance
jessevig/bertviz
BertViz is an interactive Python library for visualizing attention mechanisms in Transformer language models such as BERT, GPT-2, and RoBER…
508160stable
shenweichen/DeepCTR
DeepCTR is a Python library of easy-to-use, modular, and extendible deep-learning based CTR (click-through rate) prediction models built on…
778050stable
suno-ai/bark
Bark is Suno's open-source transformer-based text-to-audio model that generates highly realistic multilingual speech, music, background noi…
3039249maintenance
InternLM/lmdeploy
LMDeploy is a toolkit for compressing, quantizing, deploying, and serving large language models, built around its high-performance TurboMin…
958024active
lanpa/tensorboardX
A Python library for writing TensorBoard event files from PyTorch, Chainer, MXNet, NumPy, and other frameworks without needing TensorFlow. …
787998active
open-mmlab/mmpose
MMPose is an open-source pose estimation toolbox and benchmark built on PyTorch as part of the OpenMMLab ecosystem. It provides implementat…
397855active

← prev page 2 / 28 next →