Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: machine-learning

5378 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Mega4alik/ollm
oLLM is a lightweight Python library for large-context LLM inference built on Hugging Face Transformers and PyTorch. It offloads weights an…
602788active
numz/ComfyUI-SeedVR2_VideoUpscaler
The official ComfyUI integration of ByteDance's SeedVR2 model for high-quality video and image upscaling, provided as custom nodes. It can …
542786active
huggingface/setfit
SetFit is a Python library for efficient, prompt-free few-shot fine-tuning of Sentence Transformers for text classification. It achieves hi…
642784active
rhasspy/piper
Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr…
1011281maintenance
Geniusay/ChopperBot
ChopperBot is a fully automated AI bot that monitors popular live streams across platforms like Douyu, Huya, Bilibili, Douyin, and Twitch, …
452779active
FasterDecoding/Medusa
Medusa is a framework that accelerates LLM text generation by adding multiple decoding heads to an existing model, avoiding the need for a …
182770active
ideogram-oss/ideogram4
Ideogram 4 is an open-weight text-to-image foundation model trained from scratch, with inference code and weights released in Python. It fe…
542766active
autodistill/autodistill
Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab…
292763active
NVlabs/stylegan2
The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit…
3211184maintenance
AutoArk/GPA
GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an…
542762active
NVIDIA/FastPhotoStyle
FastPhotoStyle is NVIDIA's official PyTorch implementation of the ECCV 2018 paper 'A Closed-form Solution to Photorealistic Image Stylizati…
2311177maintenance
Stan
Stan is a C++ probabilistic programming library for full Bayesian inference via NUTS/Hamiltonian Monte Carlo, approximate inference via ADV…
862761stable
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
apple/turicreate
Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj…
1011159maintenance
echohive42/AI-reads-books-page-by-page
A Python script that reads PDF books page by page, using AI to extract knowledge points and generate progressive summaries at configurable …
612756active
huggingface/swift-coreml-diffusers
A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff…
542756active
voxelmorph/voxelmorph
VoxelMorph is a Python library for learning-based image registration and alignment, using unsupervised deep learning to model deformations …
762748active
pyro-ppl/numpyro
NumPyro is a lightweight probabilistic programming library providing a NumPy backend for Pyro, built on JAX for automatic differentiation a…
892746active
artidoro/qlora
QLoRA is the official implementation of the QLoRA paper, an efficient finetuning approach that backpropagates through a frozen 4-bit quanti…
2910998maintenance
Bytez-com/docs
Bytez is a serverless model inference platform offering a unified API with a single key to run 175k+ open and closed-source AI models (LLMs…
562718active
prophesier/diff-svc
Diff-SVC is a deep learning project that performs singing voice conversion using diffusion models, transforming input singing audio into a …
622717active
lengstrom/fast-style-transfer
A TensorFlow implementation of fast neural style transfer that applies the style of famous paintings to photos and videos in real time. It …
3210962maintenance
mljs/ml
ml.js is an umbrella library bundling the mljs organization's machine learning tools for JavaScript, including clustering, classification, …
232715active
vllm-project/vllm-ascend
vllm-ascend is a community-maintained hardware plugin that enables vLLM to run large language model inference on Huawei Ascend NPUs. It imp…
832711active
bmild/nerf
The official TensorFlow implementation of NeRF (Neural Radiance Fields), the ECCV 2020 paper representing scenes as neural radiance fields …
3910927maintenance
intel/neural-compressor
Intel Neural Compressor is an open-source Python library providing state-of-the-art low-bit quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4),…
922704active
yuweihao/MambaOut
MambaOut is a PyTorch implementation of Gated CNN models from the CVPR 2025 paper 'MambaOut: Do We Really Need Mamba for Vision?', which qu…
192704stable
magic-research/magic-animate
MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image …
4410897maintenance
dscripka/openWakeWord
openWakeWord is an open-source Python library for detecting wake words (or phrases) in streaming audio, with pre-trained models for common …
492702active
TMElyralab/MusePose
MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re…
282701active
IDEA-Research/T-Rex
T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete…
482699active
facebookresearch/HyperAgents
A research framework from Meta AI implementing self-referential, self-improving LLM agents that can optimize their own code for arbitrary c…
572697active
SkyworkAI/SkyReels-V1
SkyReels V1 is an open-source human-centric video foundation model with Text-to-Video and Image-to-Video variants, fine-tuned from HunyuanV…
252696active
secretflow/secretflow
SecretFlow is a unified Python framework for privacy-preserving data analysis and machine learning. It layers cryptographic devices (MPC, H…
682695active
roboflow/maestro
maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2…
622694active
naiveHobo/InvoiceNet
InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN…
322694active
intro-skipper/intro-skipper
A Jellyfin server plugin that analyzes the audio of TV episodes to automatically detect intro and credit sequences. It provides media segme…
942692active
baaivision/EVA
EVA is a family of large-scale vision foundation models from BAAI, including masked image models (EVA-01/02) and scaled CLIP models (EVA-CL…
232691active
openai/DALL-E
The official PyTorch package for the discrete VAE (dVAE) component of OpenAI's DALL·E model. It does not include the transformer that gener…
1010834maintenance
qualcomm/aimet
AIMET (AI Model Efficiency Toolkit) is a Python library from Qualcomm providing advanced quantization and compression techniques for traine…
992688active
torinmb/mediapipe-touchdesigner
A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac…
862686active
KimMeen/Time-LLM
Time-LLM is the official PyTorch implementation of an ICLR 2024 paper that reprograms frozen large language models (Llama, GPT-2, BERT) for…
472685active
bytedance/InfiniteYou
InfiniteYou (InfU) is a research framework from ByteDance for identity-preserved text-to-image generation built on Diffusion Transformers l…
372685active
yuqinie98/PatchTST
Official PyTorch implementation of PatchTST, an ICLR 2023 Transformer model for long-term time series forecasting based on patching and cha…
322685stable
thewh1teagle/kokoro-onnx
A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T…
672680active
MrGiovanni/UNetPlusPlus
Official implementation of UNet++, a nested U-Net architecture for medical image segmentation, in both Keras and PyTorch. It redesigns skip…
772679stable
HiLab-git/SSL4MIS
A benchmark and code collection of semi-supervised learning methods for medical image segmentation, re-implementing approaches like Mean Te…
442676active
tryolabs/norfair
Norfair is a lightweight, customizable Python library for real-time multi-object tracking that works with any detector outputting (x, y) co…
312676stable
PowerHouseMan/ComfyUI-AdvancedLivePortrait
A ComfyUI custom node implementing LivePortrait for fast facial expression editing and animation with real-time preview. It can edit expres…
232675active
stochasticai/xTuring
xTuring is a Python library for fine-tuning, evaluating, and running open-source large language models such as LLaMA, GPT-J, GPT-2, Qwen, a…
522674active
JIA-Lab-research/LISA
LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati…
312674active
georgia-tech-db/evadb
EvaDB is a Python database system for building AI-powered applications, exposing a SQL API that queries structured data, videos, and unstru…
102673active
ZHZisZZ/dllm
dLLM is a Python library that unifies training, inference, and evaluation of diffusion language models such as LLaDA and Dream. It builds o…
592672active
princeton-vl/DROID-SLAM
DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense …
412671active
openai/privacy-filter
OpenAI Privacy Filter is a small bidirectional token-classification model (1.5B total, 50M active parameters) for detecting and masking per…
492670active
IceClear/StableSR
StableSR is a Python research library that leverages pre-trained Stable Diffusion priors for real-world blind image super-resolution. It pr…
212668stable
aigc3d/LHM
LHM is a PyTorch-based large reconstruction model that reconstructs high-fidelity animatable 3D human avatars from a single image in second…
522664active
yxlllc/DDSP-SVC
DDSP-SVC is an open-source singing voice conversion system built on Differentiable Digital Signal Processing, designed as a free AI voice c…
652656active
Nutlope/roomGPT
RoomGPT is an open-source Next.js web application that lets users upload a photo of a room and generate redesigned variations using the Con…
3110671maintenance
google-deepmind/mctx
Mctx is a JAX-native Python library implementing Monte Carlo tree search algorithms such as AlphaZero, MuZero, and Gumbel MuZero. It suppor…
832654active
icereed/paperless-gpt
A self-hosted companion application for paperless-ngx that uses LLMs and vision models to auto-generate document titles, tags, and dates, a…
872651active
hgmzhn/manga-translator-ui
A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,…
802651active
phillipi/pix2pix
The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from…
3210652maintenance
Tencent/MimicMotion
MimicMotion is a diffusion-based framework from Tencent for generating high-quality human motion videos guided by pose sequences, featuring…
472647active
SeldonIO/alibi
Alibi is a Python library providing algorithms for explaining and interpreting machine learning models, including black-box, white-box, loc…
442644active
facebookresearch/ParlAI
ParlAI is a Python framework from Facebook AI Research for sharing, training, and evaluating dialogue models across many openly available d…
1010621maintenance
black-forest-labs/flux2
Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal …
482642active
haoheliu/AudioLDM2
AudioLDM 2 is a Python library and CLI for generating audio, music, and speech from text prompts using latent diffusion models. It includes…
282639active
ultralytics/yolov3
Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation…
6710596maintenance
YanjieZe/GMR
GMR is a Python library for general motion retargeting that converts human motions into joint commands for diverse humanoid robots in real …
522633active
anliyuan/Ultralight-Digital-Human
An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It…
642627active
lucidrains/audiolm-pytorch
A PyTorch implementation of AudioLM, Google Research's language modeling approach to audio generation, including a MIT-licensed SoundStream…
342627active
swz30/Restormer
Restormer is an efficient Transformer architecture for high-resolution image restoration, published as a CVPR 2022 Oral paper. It provides …
442625stable
sergree/matchering
Matchering 2.0 is an open-source Python library, containerized web app, and ComfyUI node for audio matching and mastering. It takes a targe…
642615active
plexe-ai/plexe
Plexe is a Python library and CLI that builds machine learning models from natural language descriptions using a multi-agent AI workflow. Y…
612614active
open-gigaai/giga-brain-0
GigaBrain-0/0.7 is an open-source vision-language-action (VLA) model family for generalist embodied agents, powered by world models and a t…
622611active
crowsonkb/k-diffusion
A PyTorch library implementing Karras et al. (2022) diffusion models with enhancements like improved sampling algorithms and transformer-ba…
532600active
meta-pytorch/torchrec
TorchRec is a PyTorch domain library for building recommendation systems at scale. It provides distributed sharding of large embedding tabl…
902599active
kairos-agi/kairos
Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f…
572598active
luca-medeiros/lang-segment-anything
A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie…
422598active
kijai/ComfyUI-HunyuanVideoWrapper
A set of custom ComfyUI nodes that wrap Tencent's HunyuanVideo text-to-video and image-to-video diffusion model for use inside ComfyUI work…
382597active
Slicer/Slicer
3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi…
672595stable
dreamzero0/dreamzero
DreamZero is NVIDIA's World Action Model (WAM) that jointly predicts future video and actions from a pretrained video diffusion backbone, e…
502593active
OmniSVG/OmniSVG
OmniSVG is a family of end-to-end multimodal SVG generation models built on pre-trained Vision-Language Models, released with inference cod…
512590active
Mininglamp-AI/Mano-P
Mano-P is an open-source GUI-VLA (vision-language-action) agent model and SDK for edge devices, enabling purely vision-driven cross-platfor…
552589active
facebookresearch/demucs
Demucs is a state-of-the-art music source separation model from Meta AI that splits songs into stems like drums, bass, and vocals using a h…
1010359maintenance
median-research-group/LibMTL
LibMTL is an open-source PyTorch library for Multi-Task Learning (MTL). It provides implementations of many MTL architectures and gradient-…
422586active
CodeGeeX
CodeGeeX is a family of open multilingual code generation large language models (13B and successors CodeGeeX2/CodeGeeX4) pre-trained on 20+…
232585active
asteroid-team/asteroid
Asteroid is a PyTorch-based audio source separation toolkit for researchers, providing modular building blocks (filterbanks, encoders, mask…
602584active
ARISE-Initiative/robosuite
robosuite is a modular simulation framework powered by the MuJoCo physics engine for robot learning, offering standardized benchmark enviro…
722581active
googlecolab/colabtools
The official Python libraries shipped inside Google Colaboratory, Google's free hosted Jupyter notebook environment for machine learning ed…
772580active
apache/hamilton
Apache Hamilton is a lightweight Python library for defining directed acyclic graphs (DAGs) of data transformations as plain, testable Pyth…
912575active
atong01/conditional-flow-matching
TorchCFM is a PyTorch library implementing Conditional Flow Matching (CFM), a simulation-free training objective for continuous normalizing…
742571active
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from …
582571active
ymcui/Chinese-BERT-wwm
A collection of Chinese pre-trained language models (BERT-wwm, BERT-wwm-ext, RoBERTa-wwm-ext, RBT3, etc.) built with Whole Word Masking, re…
6710223maintenance
advimman/lama
LaMa is a PyTorch-based image inpainting model that fills large missing regions in images using fast Fourier convolutions, generalizing wel…
3410217maintenance
kayba-ai/agentic-context-engine
Agentic Context Engine (ACE) is a Python library that lets LLM-based agents learn from experience by maintaining and evolving an evolving c…
752558active
KomputeProject/kompute
Kompute is a general-purpose GPU compute framework built on Vulkan that works across vendor GPUs (AMD, NVIDIA, Qualcomm, etc.) with both C+…
662558active
vita-epfl/Stable-Video-Infinity
Stable Video Infinity (SVI) is a research codebase for infinite-length video generation using video diffusion transformers with an error-re…
552556active
graphistry/pygraphistry
PyGraphistry is a Python library for loading, shaping, and visually exploring large graphs with GPU-accelerated rendering and analytics via…
972551active

← prev page 12 / 54 next →