Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: deep-learning

2771 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
wang-xinyu/tensorrtx
A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f…
767827active
TencentARC/GFPGAN
GFPGAN is a Python library built on PyTorch that restores and enhances real-world degraded face photos using GAN-based priors. It provides …
2337657maintenance
deepseek-ai/DeepGEMM
DeepGEMM is a high-performance CUDA BLAS kernel library for NVIDIA tensor cores, providing FP8, FP4, and BF16 GEMMs plus fused MoE and othe…
847738active
PaddlePaddle/ERNIE
Official repository for Baidu's ERNIE 4.5 family of large multimodal models and ERNIEKit, an industrial-grade training toolkit built on Pad…
657738active
MeiGen-AI/InfiniteTalk
InfiniteTalk is an open-source model and framework for unlimited-length audio-driven talking video generation, supporting both image-to-vid…
557700active
meta-llama/llama-models
Meta's official repository of Llama large language models with Python utilities for working with them, including model cards, licenses, and…
557685active
meituan-longcat/LongCat-Video
LongCat-Video is a 13.6B-parameter foundational video generation model from Meituan that unifies text-to-video, image-to-video, and video-c…
547665active
babysor/MockingBird
MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real…
5436909maintenance
1adrianb/face-alignment
A Python library built on PyTorch that detects 2D and 3D facial landmarks in images using the FAN deep learning face alignment network. It …
707538active
SkyworkAI/SkyReels-V2
SkyReels-V2 is an open-source infinite-length film/video generative model using an AutoRegressive Diffusion-Forcing architecture, released …
487462active
EleutherAI/gpt-neox
GPT-NeoX is EleutherAI's library for training large-scale autoregressive transformer language models on GPUs, built on NVIDIA's Megatron an…
627459active
LargeWorldModel/LWM
Large World Model (LWM) is a family of open-source 7B-parameter multimodal autoregressive transformer models trained on long videos and boo…
257425active
apple/ml-fastvlm
Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an…
287411active
facebookresearch/SlowFast
PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta…
657410active
facebookresearch/sam-3d-objects
SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code…
557322active
vladmandic/sdnext
SD.Next is an open-source, self-hosted WebUI server application for AI generative image and video creation built on Stable Diffusion and Di…
767320active
arcee-ai/mergekit
mergekit is a Python toolkit for merging pre-trained large language models directly in weight space, supporting many merge methods (SLERP, …
637310active
google/flax
Flax is a neural network library and ecosystem for JAX designed for flexibility, featuring the newer NNX API with first-class Python refere…
997303active
PaddlePaddle/Paddle-Lite
Paddle Lite is a high-performance, lightweight deep learning inference engine from Baidu's PaddlePaddle ecosystem, designed for mobile, emb…
587273active
mit-han-lab/streaming-llm
StreamingLLM is a research framework from MIT Han Lab implementing the Attention Sinks method (ICLR 2024) for efficient streaming language …
277268stable
Zyphra/Zonos
Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I…
247244active
kohya-ss/sd-scripts
A collection of Python training, generation, and utility scripts for Stable Diffusion and other image generation models, most widely used f…
897210active
BVLC/caffe
Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp…
2334556maintenance
tensorflow/tensorboard
TensorBoard is TensorFlow's visualization toolkit, a suite of web applications for inspecting and understanding machine learning training r…
887206active
CMU-Perceptual-Computing-Lab/openpose
OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota…
2334413maintenance
ControlNet
ControlNet is a neural network architecture that adds conditional control (edges, poses, depth, etc.) to pretrained text-to-image diffusion…
3134091maintenance
yangchris11/samurai
SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio…
277112active
threestudio-project/threestudio
threestudio is a unified open-source framework for 3D content generation from text prompts, single images, and few-shot images by lifting 2…
207059active
deepseek-ai/DeepSpec
DeepSpec is a full-stack Python codebase from DeepSeek for training and evaluating draft models used in speculative decoding of large langu…
547041active
apple/corenet
CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation …
457007active
TheTom/turboquant_plus
TurboQuant+ is a Python reference implementation of the TurboQuant KV cache compression method (ICLR 2026), using PolarQuant codebooks and …
567006active
deeppavlov/DeepPavlov
DeepPavlov is an open-source Python NLP library built on PyTorch and Hugging Face transformers for developing, training, and deploying stat…
396989active
PaddlePaddle/models
PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda…
236932active
eriklindernoren/ML-From-Scratch
A Python library providing bare-bones NumPy implementations of fundamental machine learning models and algorithms, from linear regression t…
3232528maintenance
clearml/clearml
ClearML is an open-source Python SDK and platform for MLOps/LLMOps that provides auto-logged experiment tracking, data versioning, pipeline…
996840stable
facebookresearch/fairseq
Fairseq is a PyTorch-based sequence modeling toolkit from Facebook AI Research for training custom models for translation, summarization, l…
1032231maintenance
kijai/ComfyUI-WanVideoWrapper
ComfyUI custom nodes wrapping WanVideo (Wan2.1) and related video generation models. It provides a standalone sandbox for quickly implement…
586678active
tencent-ailab/IP-Adapter
IP-Adapter is a lightweight (22M parameter) adapter that adds image prompt capability to pretrained text-to-image diffusion models like Sta…
286677stable
FoundationVision/ByteTrack
ByteTrack is a PyTorch-based multi-object tracking (MOT) library implementing the ECCV 2022 paper 'Multi-Object Tracking by Associating Eve…
326654stable
yangjianxin1/Firefly
Firefly is an open-source one-stop training tool for large language models, supporting pretraining, instruction fine-tuning (SFT), and DPO …
216653active
AILab-CVC/YOLO-World
YOLO-World is a real-time open-vocabulary object detection model and Python toolkit from Tencent AI Lab and HUST, published at CVPR 2024. I…
296529active
FareedKhan-dev/kimi-k3-in-c
A dependency-free C99 inference engine that runs the 2.78-trillion-parameter Kimi K3 model on a single CPU with as little as 8 GB of RAM by…
796524active
open-mmlab/mmdetection3d
MMDetection3D is OpenMMLab's next-generation platform for general 3D object detection, built on PyTorch. It provides a modular toolbox with…
236518active
rtqichen/torchdiffeq
torchdiffeq is a PyTorch library of differentiable ordinary differential equation (ODE) solvers, best known as the canonical implementation…
396477stable
open-mmlab/mmcv
MMCV is the foundational computer vision library for the OpenMMLab ecosystem, providing image/video I/O, data transformations, and CUDA ope…
526470stable
TMElyralab/MuseTalk
MuseTalk is a real-time, high-fidelity lip-sync model that modifies a face region in video according to input audio via latent space inpain…
456459active
HVision-NKU/StoryDiffusion
StoryDiffusion is the official implementation of a NeurIPS 2024 Spotlight paper introducing Consistent Self-Attention for character-consist…
246452active
haifengl/smile
SMILE is a comprehensive, high-performance machine learning framework for the JVM with idiomatic APIs for Java, Scala, and Kotlin. It cover…
996413active
multimodal-art-projection/YuE
YuE is a family of open-source foundation models based on the LLaMA2 architecture that generate full songs (up to five minutes) with vocals…
326403active
tensorflow/serving
TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. It mana…
866360stable
KevinMusgrave/pytorch-metric-learning
A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes…
486339active
yl4579/StyleTTS2
StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa…
296336active
mindee/doctr
docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe…
906315active
canopyai/Orpheus-TTS
Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a…
456314active
meta-llama/llama3
The official Meta repository for Llama 3, providing model weights download scripts, tokenizer, and minimal example code for running inferen…
1029247maintenance
flashinfer-ai/flashinfer
FlashInfer is a GPU kernel library and kernel generator for LLM inference, providing unified APIs for attention, GEMM, and MoE operations w…
906252active
RangiLyu/nanodet
NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN…
236252stable
PaddlePaddle/PaddleX
PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti…
926251active
meta-pytorch/gpt-fast
A minimal (<1000 lines) PyTorch-native implementation of fast transformer text generation, demonstrating low-latency LLM inference with int…
446249active
ByteDance-Seed/Depth-Anything-3
Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,…
596213active
skorch-dev/skorch
skorch is a Python library that wraps PyTorch neural networks in a scikit-learn compatible API, providing estimators like NeuralNetClassifi…
856173active
timeseriesAI/tsai
tsai is an open-source deep learning library built on PyTorch and fastai for time series and sequential data tasks such as classification, …
846111active
Akegarasu/lora-scripts
SD-Trainer is a GUI application and set of scripts for training LoRA and Dreambooth fine-tunes of Stable Diffusion diffusion models, wrappi…
666110active
bytedance/MegaTTS3
MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro…
596091active
open-edge-platform/anomalib
Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local…
986088active
svc-develop-team/so-vits-svc
A deep learning framework based on SoftVC VITS for singing voice conversion (SVC), letting users train models that convert one singing voic…
1028125maintenance
bytedance/LatentSync
LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to…
336026active
om-ai-lab/VLM-R1
VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi…
636015active
OFA-Sys/Chinese-CLIP
Chinese-CLIP is a Chinese version of the CLIP model trained on ~200 million Chinese image-text pairs, built on open_clip. It provides APIs,…
665998active
z-lab/dflash
DFlash is a lightweight block diffusion model used as a draft model for speculative decoding of large language models, drafting entire toke…
715967active
lucidrains/x-transformers
A concise PyTorch library implementing full-attention transformer architectures (encoder, decoder, encoder-decoder, and vision transformers…
855942active
pjreddie/darknet
Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete…
3226492maintenance
DeepLabCut/DeepLabCut
DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks …
895745stable
NVIDIA/DALI
NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea…
925734active
google/gemma_pytorch
The official PyTorch implementation of Google's Gemma family of open large language models, including text-only and multimodal variants. It…
105719active
google-deepmind/gemma
The official JAX-based Python library from Google DeepMind for running, sampling from, and fine-tuning the Gemma family of open-weight larg…
875695active
meta-pytorch/captum
Captum is a model interpretability and understanding library for PyTorch, providing implementations of algorithms like Integrated Gradients…
815693active
open-mmlab/OpenPCDet
OpenPCDet is a PyTorch-based open-source toolbox for LiDAR-based 3D object detection. It provides official implementations of models like P…
535692active
alexlenail/NN-SVG
A web-based tool for creating publication-ready neural network architecture diagrams parametrically, supporting FCNN, LeNet-style CNN, and …
615686stable
apple/ml-depth-pro
Depth Pro is Apple's reference implementation of a foundation model for zero-shot metric monocular depth estimation, producing sharp high-r…
305683active
huggingface/alignment-handbook
A collection of robust training recipes and scripts from Hugging Face for aligning large language models with human and AI preferences, cov…
665671active
OpenSenseNova/SenseNova-U1
SenseNova-U is a series of open-weight unified multimodal models (e.g., SenseNova-U1.5-8B-MoT) built on the NEO-unify architecture that com…
595668active
pytorch/torchtitan
torchtitan is a PyTorch-native platform for large-scale training of generative AI models, offering a clean-room implementation of PyTorch's…
795667active
Fanghua-Yu/SUPIR
SUPIR is a Python-based photo-realistic image restoration system built on SDXL diffusion priors and LLaVA captioning, presented at CVPR 202…
365649active
fla-org/flash-linear-attention
A PyTorch library providing hardware-efficient implementations of emerging sequence model architectures, including linear attention, sparse…
885627active
huggingface/parler-tts
Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via …
255586active
matterport/Mask_RCNN
A Python implementation of the Mask R-CNN model for object detection and instance segmentation, built on Keras and TensorFlow with a ResNet…
2325567maintenance
mosaicml/composer
Composer is an open-source PyTorch-based deep learning training library by MosaicML (now Databricks) for training neural networks faster an…
655495active
LaurentMazare/tch-rs
tch-rs is a Rust crate providing thin bindings to the C++ API of PyTorch (libtorch), staying close to the original API. It enables tensor o…
675479active
lyuwenyu/RT-DETR
Official implementation of RT-DETR and RT-DETRv2, real-time object detection transformers that outperform YOLO models, in PyTorch and Paddl…
745476active
tensorflow/rust
TensorFlow Rust provides idiomatic Rust language bindings for TensorFlow via its C API. It lets Rust programs build and run TensorFlow comp…
105476active
obss/sahi
SAHI (Slicing Aided Hyper Inference) is a Python vision library for detecting small objects in large images via sliced/tiled inference, wor…
995473active
HarisIqbal88/PlotNeuralNet
A LaTeX/TikZ-based library for drawing neural network architecture diagrams, with a Python interface for generating the TikZ code. It is co…
2324955maintenance
karpathy/minGPT
A minimal, clean PyTorch re-implementation of OpenAI's GPT covering both training and inference in roughly 300 lines of code. It is designe…
3224840maintenance
isl-org/MiDaS
MiDaS is a Python library with pretrained models for robust monocular depth estimation from a single image, based on the TPAMI 2022 paper a…
105420stable
facebookresearch/sapiens
Sapiens is a family of foundation models from Meta Reality Labs for human-centric vision tasks including 2D pose estimation, body-part segm…
615418active
apple/coremltools
Apple's official Python package for converting machine learning models from TensorFlow, PyTorch, scikit-learn, XGBoost, and LibSVM into the…
805399active
InternLM/xtuner
XTuner is an open-source LLM training engine from InternLM designed for fine-tuning ultra-large-scale Mixture-of-Experts (MoE) models, with…
675183active
transformerlab/transformerlab-app
Transformer Lab is an open-source desktop application (built with Electron and Python) that provides a unified GUI for training, fine-tunin…
845179active
h2oai/h2o-llmstudio
H2O LLM Studio is a framework and no-code GUI for fine-tuning state-of-the-art large language models, built by H2O.ai. It supports LoRA and…
965172active

← prev page 3 / 28 next →