domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| isaac-sim/IsaacSim NVIDIA Isaac Sim is an open-source robotics simulation application built on NVIDIA Omniverse for developing, simulating, and testing AI-dri… | 74 | 3960 | active |
| huggingface/smollm Hugging Face's repository for the SmolLM and SmolVLM families of compact, fully open language and vision-language models, including trainin… | 59 | 3884 | active |
| HumanAIGC-Engineering/OpenAvatarChat OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c… | 74 | 3723 | active |
| NExT-GPT/NExT-GPT NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide… | 37 | 3638 | active |
| ob-f/OpenBot OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo… | 67 | 3442 | active |
| Kedreamix/Linly-Talker Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, … | 48 | 3436 | active |
| NVlabs/stylegan The official TensorFlow implementation of StyleGAN, NVIDIA's style-based generator architecture for generative adversarial networks from th… | 32 | 14416 | maintenance |
| CompVis/latent-diffusion The official research code and pretrained model zoo for Latent Diffusion Models (LDM), the paper behind Stable Diffusion, enabling high-res… | 32 | 14133 | maintenance |
| amov-lab/Prometheus Prometheus is an open-source autonomous drone software system platform built on PX4 flight controller firmware and ROS. It provides onboard… | 57 | 3239 | active |
| sonos/tract Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op… | 99 | 3045 | active |
| deepseek-ai/DreamCraft3D Official PyTorch implementation of DreamCraft3D, an ICLR 2024 hierarchical 3D content generation method that turns a single 2D image into a… | 35 | 3021 | stable |
| leigest519/ScreenCoder ScreenCoder is a UI-to-code generation system that converts screenshots or design mockups into clean, editable HTML/CSS using a modular mul… | 62 | 2951 | active |
| TMElyralab/MuseV MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall… | 25 | 2846 | active |
| kairos-agi/kairos Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f… | 57 | 2598 | active |
| apple/axlearn AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf… | 71 | 2372 | active |
| Xilinx/PYNQ PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA… | 78 | 2339 | active |
| facebookresearch/ImageBind A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud… | 54 | 9064 | maintenance |
| microsoft/LLaVA-Med LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su… | 40 | 2231 | active |
| ZiYang-xie/WorldGen WorldGen is a Python library that generates full 3D scenes in seconds from text prompts or images, supporting 360-degree consistent explora… | 53 | 2076 | active |
| PRIS-CV/DemoFusion DemoFusion is a CVPR 2024 framework that extends open-source latent diffusion models like SDXL to generate high-resolution images without a… | 48 | 2041 | stable |
| A9T9/RPA Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible … | 96 | 1985 | active |
| OpenMotionLab/MotionGPT MotionGPT is a unified motion-language model that treats 3D human motion as a foreign language by converting motion into discrete motion to… | 32 | 1961 | active |
| 2U1/Qwen-VL-Series-Finetune An open-source Python repository providing training scripts for fine-tuning Alibaba's Qwen-VL series of vision-language models (Qwen2-VL, Q… | 67 | 1960 | active |
| tensorlayer/TensorLayer TensorLayer is a TensorFlow-based deep learning and reinforcement learning library offering customizable neural layers for researchers and … | 23 | 7381 | maintenance |
| probcomp/Gen.jl Gen.jl is a general-purpose probabilistic programming system embedded in Julia that lets users write generative models as probabilistic pro… | 62 | 1850 | active |
| ROCm/FastFlowLM FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama… | 85 | 1809 | active |
| TencentARC/BrushNet BrushNet is the official PyTorch implementation of an ECCV 2024 plug-and-play image inpainting model that embeds pixel-level masked image f… | 25 | 1745 | active |
| facebookresearch/multimodal TorchMultimodal is a PyTorch library from Meta for training state-of-the-art multimodal multi-task models at scale, covering both content u… | 77 | 1732 | active |
| vdaas/vald Vald is a highly scalable, distributed approximate nearest neighbor (ANN) dense vector search engine built on cloud-native architecture and… | 91 | 1718 | active |
| robin-shaun/XTDrone XTDrone is a customizable UAV simulation platform built on PX4, ROS, and Gazebo, supporting multi-rotors, fixed-wing, VTOL vehicles, and ot… | 48 | 1711 | active |
| software-mansion/react-native-executorch React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi… | 88 | 1702 | active |
| ZJU4HealthCare/HealthGPT HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled… | 63 | 1654 | active |
| OminousIndustries/PhoneDriver PhoneDriver is a Python-based mobile automation agent that uses Qwen3-VL vision-language models to visually understand and control Android … | 38 | 1614 | active |
| pixeltable/pixeltable Pixeltable is a Python library providing declarative, incremental data infrastructure for multimodal AI applications, unifying storage of i… | 92 | 1613 | active |
| pypose/pypose PyPose is a PyTorch-based Python library for differentiable robotics on manifolds, combining deep perceptual models with physics-based opti… | 91 | 1605 | active |
| semperai/amica Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. … | 32 | 1593 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1590 | active |
| FoundationVision/Infinity Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints… | 56 | 1587 | active |
| byjlw/video-analyzer A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs … | 60 | 1559 | active |
| hustvl/LightningDiT LightningDiT is a research codebase for latent diffusion models implementing VA-VAE and LightningDiT, achieving FID 1.35 on ImageNet-256 wi… | 47 | 1529 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1516 | active |
| qupath/qupath QuPath is an open-source desktop application for bioimage analysis, aimed especially at digital pathology and whole-slide imaging. It provi… | 79 | 1430 | active |
| jrzaurin/pytorch-widedeep A PyTorch library for multimodal deep learning that combines tabular data with text and images using Wide and Deep model architectures. It … | 62 | 1416 | active |
| affinelayer/pix2pix-tensorflow A TensorFlow implementation of pix2pix, a conditional GAN that learns a mapping from input images to output images. It is a faithful port o… | 32 | 5081 | maintenance |
| ARahim3/mlx-tune A Python library for fine-tuning LLMs, vision-language, audio (TTS/STT), embedding, OCR, and JEPA models natively on Apple Silicon Macs usi… | 75 | 1389 | active |
| bytedance/UNO UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul… | 38 | 1362 | active |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| robodhruv/visualnav-transformer Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie… | 19 | 1294 | active |
| PrunaAI/pruna Pruna is an open-source Python model optimization framework that makes AI models faster, smaller, cheaper, and greener via caching, quantiz… | 84 | 1275 | active |
| Stable-X/Stable3DGen Stable3DGen is a modular Python framework for generating 3D assets from images, adapted from Microsoft's TRELLIS with NVIDIA library depend… | 33 | 1274 | active |
| Renumics/spotlight Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe… | 93 | 1272 | active |
| lucidrains/flamingo-pytorch A PyTorch implementation of DeepMind's Flamingo visual language model architecture, providing the Perceiver Resampler and Gated Cross-Atten… | 23 | 1269 | active |
| GML-MMGroup/GMTalker GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l… | 46 | 1217 | active |
| Artelnics/opennn OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o… | 97 | 1198 | active |
| NVlabs/alpasim AlpaSim is an open-source, Python-based autonomous vehicle simulation platform for developing and testing end-to-end AV policies in closed … | 75 | 1196 | active |
| qualcomm/ai-hub-models Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-… | 89 | 1195 | active |
| Cerebras/modelzoo Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and … | 77 | 1193 | active |
| xemle/home-gallery HomeGallery is a self-hosted, open-source web gallery for browsing personal photos and videos with a mobile-friendly interface. It offers A… | 72 | 1174 | active |
| bytedance/1d-tokenizer A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i… | 29 | 1172 | active |
| simpler-env/SimplerEnv SIMPLER (SimplerEnv) is a collection of simulated environments built on SAPIEN/ManiSkill for evaluating real-world robot manipulation polic… | 51 | 1147 | active |
| TowhidKashem/snapchat-clone A Snapchat clone web application built with React, Redux Toolkit, and TypeScript, featuring camera-based face filters with Three.js augment… | 64 | 1136 | active |
| FlagOpen/RoboBrain2.5 RoboBrain 2.5 is an open-source embodied AI foundation model from BAAI that combines multimodal large language model capabilities with 3D s… | 50 | 1132 | active |
| BeingBeyond/Being-H Being-H is a family of human-centric embodied foundation models, including VLA models (Being-H0.5, Being-H0) and latent world-action models… | 63 | 1126 | active |
| FlagAI-Open/FlagAI FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int… | 64 | 3869 | maintenance |
| alibaba-damo-academy/RynnVLA-002 RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation… | 43 | 1119 | active |
| gabber-dev/gabber Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr… | 44 | 1111 | active |
| TencentARC/T2I-Adapter Official implementation of T2I-Adapter, lightweight adapter models that add controllable conditioning (sketch, canny, lineart, depth, pose)… | 31 | 3801 | maintenance |
| kerberos-io/agent Kerberos Agent is an open-source, scalable video surveillance application written in Go with a React frontend, designed to connect to IP ca… | 95 | 1103 | active |
| rhymes-ai/Aria Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda… | 23 | 1087 | active |
| DSE-MSU/DeepRobust DeepRobust is a PyTorch library for adversarial robustness research, providing implementations of attack and defense methods for both image… | 45 | 1085 | active |
| SimpleITK/SimpleITK SimpleITK is a simplified C++ interface to the Insight Toolkit (ITK) for multi-dimensional image analysis, including filtering, segmentatio… | 98 | 1084 | stable |
| NVlabs/Fast-dLLM NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based … | 57 | 1082 | active |
| autonomousvision/navsim NAVSIM is a data-driven pseudo-simulation framework and benchmark for autonomous vehicle planning, evaluating driving agents non-reactively… | 50 | 1076 | active |
| NVIDIA/DreamDojo NVIDIA's official PyTorch codebase for DreamDojo, a generalist robot world model pretrained on 44k hours of human egocentric video and post… | 48 | 1059 | active |
| tensorflow/hub TensorFlow Hub is a Python library for reusing parts of trained TensorFlow models (SavedModels) for transfer learning, wrapping them as Ker… | 24 | 3523 | maintenance |
| manycoretech/aholo-viewer Aholo Viewer is a high-performance TypeScript renderer for 3D Gaussian Splatting (3DGS) scenes and meshes, using a chunked streaming LOD sc… | 79 | 1021 | active |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3454 | maintenance |
| Alpha-VLLM/Lumina-DiMOO Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu… | 55 | 1015 | active |
| MeshAnything MeshAnything is an autoregressive transformer model that generates artist-created 3D meshes (up to 1600 faces in V2) aligned with a given s… | 31 | 1015 | active |
| EvolvingLMMs-Lab/Otter Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa… | 21 | 3436 | maintenance |
| huawei-noah/noah-research A collection of research code subprojects released by Huawei Noah's Ark Lab, each in its own directory. It is not an official Huawei produc… | 76 | 1004 | active |
| ZJUI-AI4H/Hulu-Med Hulu-Med is a family of open-source transparent generalist medical vision-language models ranging from 4B to 235B parameters, covering text… | 62 | 1001 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| Alpha-VLLM/LLaMA2-Accessory LLaMA2-Accessory is an open-source Python toolkit for pretraining, finetuning, and deploying large language models and multimodal LLMs, inc… | 29 | 2800 | maintenance |
| Tencent/MedicalNet MedicalNet provides a series of 3D-ResNet pre-trained models trained on 23 diverse medical imaging datasets, with PyTorch transfer-learning… | 57 | 2255 | maintenance |
| atriumlts/subpixel A TensorFlow reimplementation of the efficient sub-pixel convolutional neural network (ESPCN) for single-image super-resolution, based on S… | 32 | 2123 | maintenance |
| tianqiraf/DouZero_For_HappyDouDiZhu A Python desktop application that applies the DouZero reinforcement-learning Dou Dizhu (Chinese card game) AI to the popular Happy DouDiZhu… | 23 | 2065 | maintenance |
| NVlabs/alpamayo NVIDIA Alpamayo 1 is an open 10B-parameter reasoning vision-language-action (VLA) model for autonomous vehicles that pairs driving trajecto… | 59 | 2005 | maintenance |
| SummitKwan/transparent_latent_gan TL-GAN is a Python/TensorFlow project that makes a GAN's latent space transparent by discovering feature axes, enabling controlled image sy… | 32 | 1973 | maintenance |
| openai/Video-Pre-Training OpenAI's Video PreTraining (VPT) codebase for learning Minecraft agents by watching unlabeled online videos, including behavioral cloning a… | 41 | 1737 | maintenance |
| ispysoftware/iSpy iSpy is an open source video surveillance application for Windows that connects to webcams and IP cameras, providing live viewing, motion d… | 62 | 1613 | maintenance |
| NVIDIAGameWorks/kaolin-wisp NVIDIA Kaolin Wisp is a PyTorch library and engine for neural fields research, built on NVIDIA Kaolin Core. It provides differentiable rend… | 23 | 1498 | maintenance |
| google-research/disentanglement_lib disentanglement_lib is an open-source Python library for research on learning disentangled representations, supporting models like BetaVAE,… | 10 | 1425 | maintenance |
| lfz/DSB2017 The winning solution of team 'grt123' for the 2017 Data Science Bowl (DSB2017), a deep learning pipeline for detecting lung cancer from CT … | 32 | 1241 | maintenance |
| mit-han-lab/tinyml MIT Han Lab's TinyML research repository containing projects like TinyTL and NetAug for memory-efficient deep learning on microcontrollers … | 32 | 1211 | maintenance |
| HyperGAN/HyperGAN HyperGAN is a composable GAN (generative adversarial network) framework built on PyTorch, offering both a Python API and a CLI with a user … | 23 | 1183 | maintenance |
| 18601949127/DiDiCallCar An Android ride-hailing demo app modeled on Didi, built end-to-end by one developer including the Apache+PHP+MySQL backend. It adds RFID/NF… | 32 | 1164 | maintenance |
| alexandre01/deepsvg Official PyTorch code for the NeurIPS 2020 paper 'DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation'. It provides a … | 32 | 1164 | maintenance |
| Timthony/self_drive A self-driving RC car project based on Raspberry Pi and TensorFlow/Keras. It collects camera images while a human drives the car on a taped… | 32 | 1132 | maintenance |
| yu-takagi/StableDiffusionReconstruction Research codebase reproducing Takagi and Nishimoto's CVPR 2023 method for reconstructing images a person viewed from fMRI brain activity us… | 30 | 1127 | maintenance |