Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
isaac-sim/IsaacSim
NVIDIA Isaac Sim is an open-source robotics simulation application built on NVIDIA Omniverse for developing, simulating, and testing AI-dri…
743960active
huggingface/smollm
Hugging Face's repository for the SmolLM and SmolVLM families of compact, fully open language and vision-language models, including trainin…
593884active
HumanAIGC-Engineering/OpenAvatarChat
OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c…
743723active
NExT-GPT/NExT-GPT
NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide…
373638active
ob-f/OpenBot
OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo…
673442active
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
NVlabs/stylegan
The official TensorFlow implementation of StyleGAN, NVIDIA's style-based generator architecture for generative adversarial networks from th…
3214416maintenance
CompVis/latent-diffusion
The official research code and pretrained model zoo for Latent Diffusion Models (LDM), the paper behind Stable Diffusion, enabling high-res…
3214133maintenance
amov-lab/Prometheus
Prometheus is an open-source autonomous drone software system platform built on PX4 flight controller firmware and ROS. It provides onboard…
573239active
sonos/tract
Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op…
993045active
deepseek-ai/DreamCraft3D
Official PyTorch implementation of DreamCraft3D, an ICLR 2024 hierarchical 3D content generation method that turns a single 2D image into a…
353021stable
leigest519/ScreenCoder
ScreenCoder is a UI-to-code generation system that converts screenshots or design mockups into clean, editable HTML/CSS using a modular mul…
622951active
TMElyralab/MuseV
MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall…
252846active
kairos-agi/kairos
Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f…
572598active
apple/axlearn
AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf…
712372active
Xilinx/PYNQ
PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA…
782339active
facebookresearch/ImageBind
A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud…
549064maintenance
microsoft/LLaVA-Med
LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su…
402231active
ZiYang-xie/WorldGen
WorldGen is a Python library that generates full 3D scenes in seconds from text prompts or images, supporting 360-degree consistent explora…
532076active
PRIS-CV/DemoFusion
DemoFusion is a CVPR 2024 framework that extends open-source latent diffusion models like SDXL to generate high-resolution images without a…
482041stable
A9T9/RPA
Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible …
961985active
OpenMotionLab/MotionGPT
MotionGPT is a unified motion-language model that treats 3D human motion as a foreign language by converting motion into discrete motion to…
321961active
2U1/Qwen-VL-Series-Finetune
An open-source Python repository providing training scripts for fine-tuning Alibaba's Qwen-VL series of vision-language models (Qwen2-VL, Q…
671960active
tensorlayer/TensorLayer
TensorLayer is a TensorFlow-based deep learning and reinforcement learning library offering customizable neural layers for researchers and …
237381maintenance
probcomp/Gen.jl
Gen.jl is a general-purpose probabilistic programming system embedded in Julia that lets users write generative models as probabilistic pro…
621850active
ROCm/FastFlowLM
FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama…
851809active
TencentARC/BrushNet
BrushNet is the official PyTorch implementation of an ECCV 2024 plug-and-play image inpainting model that embeds pixel-level masked image f…
251745active
facebookresearch/multimodal
TorchMultimodal is a PyTorch library from Meta for training state-of-the-art multimodal multi-task models at scale, covering both content u…
771732active
vdaas/vald
Vald is a highly scalable, distributed approximate nearest neighbor (ANN) dense vector search engine built on cloud-native architecture and…
911718active
robin-shaun/XTDrone
XTDrone is a customizable UAV simulation platform built on PX4, ROS, and Gazebo, supporting multi-rotors, fixed-wing, VTOL vehicles, and ot…
481711active
software-mansion/react-native-executorch
React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi…
881702active
ZJU4HealthCare/HealthGPT
HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled…
631654active
OminousIndustries/PhoneDriver
PhoneDriver is a Python-based mobile automation agent that uses Qwen3-VL vision-language models to visually understand and control Android …
381614active
pixeltable/pixeltable
Pixeltable is a Python library providing declarative, incremental data infrastructure for multimodal AI applications, unifying storage of i…
921613active
pypose/pypose
PyPose is a PyTorch-based Python library for differentiable robotics on manifolds, combining deep perceptual models with physics-based opti…
911605active
semperai/amica
Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. …
321593active
Tencent-Hunyuan/HunyuanWorld-Voyager
HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ…
521590active
FoundationVision/Infinity
Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints…
561587active
byjlw/video-analyzer
A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs …
601559active
hustvl/LightningDiT
LightningDiT is a research codebase for latent diffusion models implementing VA-VAE and LightningDiT, achieving FID 1.35 on ImageNet-256 wi…
471529active
microsoft/Mage
Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage…
571516active
qupath/qupath
QuPath is an open-source desktop application for bioimage analysis, aimed especially at digital pathology and whole-slide imaging. It provi…
791430active
jrzaurin/pytorch-widedeep
A PyTorch library for multimodal deep learning that combines tabular data with text and images using Wide and Deep model architectures. It …
621416active
affinelayer/pix2pix-tensorflow
A TensorFlow implementation of pix2pix, a conditional GAN that learns a mapping from input images to output images. It is a faithful port o…
325081maintenance
ARahim3/mlx-tune
A Python library for fine-tuning LLMs, vision-language, audio (TTS/STT), embedding, OCR, and JEPA models natively on Apple Silicon Macs usi…
751389active
bytedance/UNO
UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul…
381362active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
robodhruv/visualnav-transformer
Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie…
191294active
PrunaAI/pruna
Pruna is an open-source Python model optimization framework that makes AI models faster, smaller, cheaper, and greener via caching, quantiz…
841275active
Stable-X/Stable3DGen
Stable3DGen is a modular Python framework for generating 3D assets from images, adapted from Microsoft's TRELLIS with NVIDIA library depend…
331274active
Renumics/spotlight
Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe…
931272active
lucidrains/flamingo-pytorch
A PyTorch implementation of DeepMind's Flamingo visual language model architecture, providing the Perceiver Resampler and Gated Cross-Atten…
231269active
GML-MMGroup/GMTalker
GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l…
461217active
Artelnics/opennn
OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o…
971198active
NVlabs/alpasim
AlpaSim is an open-source, Python-based autonomous vehicle simulation platform for developing and testing end-to-end AV policies in closed …
751196active
qualcomm/ai-hub-models
Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-…
891195active
Cerebras/modelzoo
Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and …
771193active
xemle/home-gallery
HomeGallery is a self-hosted, open-source web gallery for browsing personal photos and videos with a mobile-friendly interface. It offers A…
721174active
bytedance/1d-tokenizer
A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i…
291172active
simpler-env/SimplerEnv
SIMPLER (SimplerEnv) is a collection of simulated environments built on SAPIEN/ManiSkill for evaluating real-world robot manipulation polic…
511147active
TowhidKashem/snapchat-clone
A Snapchat clone web application built with React, Redux Toolkit, and TypeScript, featuring camera-based face filters with Three.js augment…
641136active
FlagOpen/RoboBrain2.5
RoboBrain 2.5 is an open-source embodied AI foundation model from BAAI that combines multimodal large language model capabilities with 3D s…
501132active
BeingBeyond/Being-H
Being-H is a family of human-centric embodied foundation models, including VLA models (Being-H0.5, Being-H0) and latent world-action models…
631126active
FlagAI-Open/FlagAI
FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int…
643869maintenance
alibaba-damo-academy/RynnVLA-002
RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation…
431119active
gabber-dev/gabber
Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr…
441111active
TencentARC/T2I-Adapter
Official implementation of T2I-Adapter, lightweight adapter models that add controllable conditioning (sketch, canny, lineart, depth, pose)…
313801maintenance
kerberos-io/agent
Kerberos Agent is an open-source, scalable video surveillance application written in Go with a React frontend, designed to connect to IP ca…
951103active
rhymes-ai/Aria
Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda…
231087active
DSE-MSU/DeepRobust
DeepRobust is a PyTorch library for adversarial robustness research, providing implementations of attack and defense methods for both image…
451085active
SimpleITK/SimpleITK
SimpleITK is a simplified C++ interface to the Insight Toolkit (ITK) for multi-dimensional image analysis, including filtering, segmentatio…
981084stable
NVlabs/Fast-dLLM
NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based …
571082active
autonomousvision/navsim
NAVSIM is a data-driven pseudo-simulation framework and benchmark for autonomous vehicle planning, evaluating driving agents non-reactively…
501076active
NVIDIA/DreamDojo
NVIDIA's official PyTorch codebase for DreamDojo, a generalist robot world model pretrained on 44k hours of human egocentric video and post…
481059active
tensorflow/hub
TensorFlow Hub is a Python library for reusing parts of trained TensorFlow models (SavedModels) for transfer learning, wrapping them as Ker…
243523maintenance
manycoretech/aholo-viewer
Aholo Viewer is a high-performance TypeScript renderer for 3D Gaussian Splatting (3DGS) scenes and meshes, using a chunked streaming LOD sc…
791021active
towhee-io/towhee
Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s…
233454maintenance
Alpha-VLLM/Lumina-DiMOO
Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu…
551015active
MeshAnything
MeshAnything is an autoregressive transformer model that generates artist-created 3D meshes (up to 1600 faces in V2) aligned with a given s…
311015active
EvolvingLMMs-Lab/Otter
Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa…
213436maintenance
huawei-noah/noah-research
A collection of research code subprojects released by Huawei Noah's Ark Lab, each in its own directory. It is not an official Huawei produc…
761004active
ZJUI-AI4H/Hulu-Med
Hulu-Med is a family of open-source transparent generalist medical vision-language models ranging from 4B to 235B parameters, covering text…
621001active
siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System
A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro…
391001active
Alpha-VLLM/LLaMA2-Accessory
LLaMA2-Accessory is an open-source Python toolkit for pretraining, finetuning, and deploying large language models and multimodal LLMs, inc…
292800maintenance
Tencent/MedicalNet
MedicalNet provides a series of 3D-ResNet pre-trained models trained on 23 diverse medical imaging datasets, with PyTorch transfer-learning…
572255maintenance
atriumlts/subpixel
A TensorFlow reimplementation of the efficient sub-pixel convolutional neural network (ESPCN) for single-image super-resolution, based on S…
322123maintenance
tianqiraf/DouZero_For_HappyDouDiZhu
A Python desktop application that applies the DouZero reinforcement-learning Dou Dizhu (Chinese card game) AI to the popular Happy DouDiZhu…
232065maintenance
NVlabs/alpamayo
NVIDIA Alpamayo 1 is an open 10B-parameter reasoning vision-language-action (VLA) model for autonomous vehicles that pairs driving trajecto…
592005maintenance
SummitKwan/transparent_latent_gan
TL-GAN is a Python/TensorFlow project that makes a GAN's latent space transparent by discovering feature axes, enabling controlled image sy…
321973maintenance
openai/Video-Pre-Training
OpenAI's Video PreTraining (VPT) codebase for learning Minecraft agents by watching unlabeled online videos, including behavioral cloning a…
411737maintenance
ispysoftware/iSpy
iSpy is an open source video surveillance application for Windows that connects to webcams and IP cameras, providing live viewing, motion d…
621613maintenance
NVIDIAGameWorks/kaolin-wisp
NVIDIA Kaolin Wisp is a PyTorch library and engine for neural fields research, built on NVIDIA Kaolin Core. It provides differentiable rend…
231498maintenance
google-research/disentanglement_lib
disentanglement_lib is an open-source Python library for research on learning disentangled representations, supporting models like BetaVAE,…
101425maintenance
lfz/DSB2017
The winning solution of team 'grt123' for the 2017 Data Science Bowl (DSB2017), a deep learning pipeline for detecting lung cancer from CT …
321241maintenance
mit-han-lab/tinyml
MIT Han Lab's TinyML research repository containing projects like TinyTL and NetAug for memory-efficient deep learning on microcontrollers …
321211maintenance
HyperGAN/HyperGAN
HyperGAN is a composable GAN (generative adversarial network) framework built on PyTorch, offering both a Python API and a CLI with a user …
231183maintenance
18601949127/DiDiCallCar
An Android ride-hailing demo app modeled on Didi, built end-to-end by one developer including the Apache+PHP+MySQL backend. It adds RFID/NF…
321164maintenance
alexandre01/deepsvg
Official PyTorch code for the NeurIPS 2020 paper 'DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation'. It provides a …
321164maintenance
Timthony/self_drive
A self-driving RC car project based on Raspberry Pi and TensorFlow/Keras. It collects camera images while a human drives the car on a taped…
321132maintenance
yu-takagi/StableDiffusionReconstruction
Research codebase reproducing Takagi and Nishimoto's CVPR 2023 method for reconstructing images a person viewed from fMRI brain activity us…
301127maintenance

← prev page 23 / 24 next →