function: machine-learning
5378 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| GML-MMGroup/GMTalker GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l… | 45 | 1222 | active |
| jx-sec/jxwaf JXWAF is an open-source web application firewall powered by an AI large language model, combining an AI security model, a semantic analysis… | 67 | 1221 | active |
| InternScience/GraphGen GraphGen is a Python framework for knowledge-graph-guided synthetic data generation for LLM training. It builds fine-grained knowledge grap… | 58 | 1216 | active |
| TIGER-AI-Lab/OpenResearcher OpenResearcher is a fully open-source pipeline for synthesizing long-horizon deep research trajectories using LLM agents with retrieval and… | 54 | 1215 | active |
| kijai/ComfyUI-segment-anything-2 A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm… | 42 | 1215 | active |
| taranis-ai/taranis-ai Taranis AI is a self-hosted open-source OSINT platform that collects news articles from web sources and uses NLP/AI to enrich, cluster, and… | 94 | 1208 | active |
| bytedance/Fastbot_Android Fastbot is a model-based GUI testing tool for Android apps that models GUI transitions using machine learning and reinforcement learning to… | 44 | 1203 | active |
| cleanlab/cleanvision CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate … | 58 | 1199 | active |
| microsoft/RecAI RecAI is a Microsoft research project exploring ways to integrate large language models into recommender systems (LLM4Rec). It includes a c… | 56 | 1194 | active |
| diodiogod/TTS-Audio-Suite A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt… | 84 | 1190 | active |
| ali-vilab/UniAnimate UniAnimate is the official code for a research paper on animating a reference human image into a video that follows a driving pose sequence… | 31 | 1189 | active |
| VladimirYugay/Gaussian-SLAM A research implementation of a dense RGBD SLAM system that uses 3D Gaussian Splatting as its scene representation to photorealistically rec… | 27 | 1182 | active |
| Facico/Chinese-Vicuna Chinese-Vicuna is a low-resource LLaMA+LoRA solution for building Chinese instruction-following language models, structured after Alpaca. I… | 37 | 4111 | maintenance |
| MediaBrain-SJTU/MING MING (明医) is a Chinese medical consultation large language model fine-tuned on medical instruction data, with variants built on bloomz-7b a… | 40 | 1177 | active |
| wladradchenko/wunjo.wladradchenko.ru Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a… | 70 | 1170 | active |
| FoundationVision/GLEE GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o… | 26 | 1170 | active |
| soniqo/speech-swift An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v… | 83 | 1164 | active |
| CyberAgentAILab/TANGO TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion … | 38 | 1163 | active |
| 3D ResNets for Action Recognition A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ… | 23 | 4038 | maintenance |
| meta-pytorch/torchcodec TorchCodec is a PyTorch-native library for decoding and encoding videos, audio, and images into PyTorch tensors on CPU and CUDA GPU, built … | 86 | 1162 | active |
| TensorSpeech/TensorFlowTTS TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me… | 23 | 3996 | maintenance |
| ttttccxxui/DataInfra-RedactionEverything A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla… | 60 | 1147 | active |
| CharlesWiltgen/Axiom Axiom is a collection of 273 battle-tested skills, 42 agents, 17 commands, and bundled developer tools (xclog, xcsym, xcui, xcprof) that gi… | 61 | 1146 | active |
| agent-topia/evolving_personality JPAF is a Python framework that gives LLM agents structured, evolving MBTI personalities based on Carl Jung's psychological types. It uses … | 48 | 1146 | active |
| Aaronontheweb/dotnet-skills A plugin of 30 skills and 5 specialized sub-agents that give AI coding assistants (Claude Code, Codex, GitHub Copilot, OpenCode) expert kno… | 81 | 1138 | active |
| andabi/deep-voice-conversion A TensorFlow implementation of deep neural networks for voice conversion (voice style transfer) that converts a source speaker's voice into… | 32 | 3938 | maintenance |
| K-Dense-AI/k-dense-byok K-Dense BYOK is a free, open-source desktop application providing 'Kady', an AI research assistant (co-scientist) that runs locally and use… | 78 | 1131 | active |
| AutoArk/open-audio-opd An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models … | 52 | 1129 | active |
| jerry-ai-dev/MODULAR-RAG-MCP-SERVER A modular, pluggable RAG (Retrieval-Augmented Generation) framework exposed as an MCP (Model Context Protocol) server, so AI assistants lik… | 47 | 1127 | active |
| ZiwenZhuang/parkour Official code for 'Robot Parkour Learning' (CoRL 2023), a reinforcement learning system that trains quadrupedal robots to perform vision-ba… | 50 | 1121 | active |
| dominostars/playtranslate PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for … | 81 | 1118 | active |
| FutureUniant/Tailor Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea… | 37 | 1117 | active |
| BytedTsinghua-SIA/MemAgent MemAgent is a reinforcement-learning framework for training LLM agents that process arbitrarily long contexts via a memory mechanism within… | 55 | 1102 | active |
| google/fully-homomorphic-encryption Google's repository of demos for fully homomorphic encryption (FHE), originally a C++ transpiler and now showcasing the HEIR MLIR-based FHE… | 86 | 3754 | maintenance |
| Bogdanovich77/DeekSeek-OCR---Dockerized-API A Dockerized REST API and batch processing scripts that convert PDF documents to Markdown using the DeepSeek-OCR model behind a FastAPI bac… | 38 | 1092 | active |
| hinthornw/trustcall A Python library that improves LLM tool calling reliability by having models generate JSON patch operations instead of full JSON blobs. It … | 38 | 1091 | active |
| nobodywho-ooo/nobodywho NobodyWho is an open-source (EUPL-1.2) on-device LLM inference engine written in Rust, built on llama.cpp, with SDKs for Kotlin, Swift, Pyt… | 86 | 1088 | active |
| InternRobotics/InternNav InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation… | 57 | 1088 | active |
| dtsola/xiaoyaosearch XiaoyaoSearch is a cross-platform desktop application (Electron + Python/FastAPI) that lets users find local files using AI-powered semanti… | 70 | 1085 | active |
| HZJQF/help_tool A PyQt5-based Windows GUI tool that uses inference models to identify the encryption or hashing algorithm behind a given ciphertext and att… | 24 | 1084 | active |
| olivia-ai/olivia Olivia is an open-source chatbot written in Go that uses a neural network for natural language understanding, aiming to be a free alternati… | 10 | 3717 | maintenance |
| Pokee-AI/PokeeResearchOSS An open-source repository for Pokee's 7B-parameter DeepResearch agent that performs multi-turn web search and content reading to answer com… | 38 | 1077 | active |
| trailofbits/anamorpher Anamorpher is a tool for crafting and visualizing image scaling attacks that hide multi-modal prompt injections in images, revealed only wh… | 55 | 1076 | active |
| IAHispano/Applio Applio is an open-source, MIT-licensed voice conversion suite built on RVC that lets users convert audio into other voices, train custom vo… | 90 | 3677 | maintenance |
| GanymedeNil/document.ai A universal local knowledge base solution that stores documents as vectors in a vector database and uses GPT3.5 to generate answers from re… | 30 | 3669 | maintenance |
| rafael-fuente/diffractsim A flexible Python library for simulating and visualizing diffraction and physical optics phenomena using scalar diffraction techniques like… | 65 | 1066 | active |
| bytedance/SandboxFusion A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via … | 63 | 1062 | active |
| Makememo/MemoAI MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and… | 95 | 1061 | active |
| zai-org/GLM-TTS GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur… | 50 | 1061 | active |
| YeeZTech/YeeZ-Privacy-Computing Fidelius is a privacy computing middleware built on Intel SGX trusted execution environments, enabling secure data collaboration where orig… | 77 | 1057 | active |
| Dicklesworthstone/swiss_army_llama A FastAPI-based REST service that exposes local LLM capabilities including text embeddings, completions, semantic similarity, and semantic … | 32 | 1055 | active |
| DrizzleTime/Foxel Foxel is a self-hosted private cloud storage application that unifies file management across pluggable storage backends (S3, WebDAV, cloud … | 80 | 1054 | active |
| Artrajz/vits-simple-api A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.… | 69 | 1051 | active |
| smxiazi/NEW_xp_CAPTCHA xp_CAPTCHA is a Burp Suite extension (Java plugin) that automatically recognizes CAPTCHAs during brute-force attacks, using a companion Pyt… | 23 | 1050 | active |
| stay-leave/weibo-public-opinion-analysis A Python project for Weibo public opinion analysis that combines a web crawler, LDA topic modeling, sentiment analysis, and spatiotemporal … | 32 | 1046 | active |
| airbnb/chronon Chronon is an open-source data platform from Airbnb for computing, backfilling, and serving ML features. It handles batch and streaming fea… | 75 | 1043 | active |
| foolwood/SiamMask Official PyTorch implementation of SiamMask, a deep learning framework for fast online visual object tracking and video object segmentation… | 35 | 3550 | maintenance |
| datamllab/rlcard RLCard is a Python toolkit for reinforcement learning research in card games, providing environments for Blackjack, Leduc Hold'em, Texas Ho… | 23 | 3541 | maintenance |
| TIGER-AI-Lab/verl-tool VerlTool is a unified, extensible framework built on verl for training LLM agents with tool use via reinforcement learning. It decouples ac… | 66 | 1037 | active |
| centerforaisafety/HarmBench HarmBench is a standardized, open-source evaluation framework for automated red teaming of large language models, comparing attack methods … | 26 | 1037 | active |
| BabitMF/bmf BMF (Babit Multimedia Framework) is a cross-platform, multi-language multimedia and video processing framework developed by ByteDance, offe… | 75 | 1034 | active |
| videoflow/videoflow Videoflow is a Python framework for building distributed video and stream processing pipelines as directed acyclic graphs of producers, pro… | 65 | 1034 | active |
| XYZ-AI-Lab/axrl AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout … | 56 | 1034 | active |
| soCzech/TransNetV2 TransNet V2 is a deep neural network for shot boundary detection in videos, achieving state-of-the-art results on benchmarks like ClipShots… | 32 | 1033 | stable |
| abizovnuralem/go2_ros2_sdk An unofficial ROS2 SDK for the Unitree Go2 quadruped robot (AIR/PRO/EDU), connecting over WebRTC (Wi-Fi) or CycloneDDS (Ethernet). It provi… | 57 | 1026 | active |
| microsoft/tensorwatch TensorWatch is a Python library from Microsoft Research for debugging, monitoring, and visualizing machine learning training in real time, … | 66 | 3472 | maintenance |
| silvery107/rl-mpc-locomotion A Python framework combining deep reinforcement learning with model predictive control (MPC) for quadruped robot locomotion, where a policy… | 67 | 1022 | active |
| autonomousvision/gaussian-opacity-fields Gaussian Opacity Fields (GOF) is a Python/CUDA research implementation for efficient, adaptive surface reconstruction in unbounded scenes u… | 25 | 1018 | active |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3453 | maintenance |
| Sxela/WarpFusion WarpFusion is a Stable Diffusion-based video-to-video style transfer tool distributed as a Jupyter/Colab notebook. It applies AI animation … | 31 | 1016 | active |
| eragonruan/text-detection-ctpn A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr… | 23 | 3429 | maintenance |
| HumeAI/tada TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,… | 51 | 1010 | active |
| sematic-ai/sematic Sematic is an open-source ML pipeline development platform that lets engineers write end-to-end pipelines in pure Python. Pipelines can run… | 30 | 1002 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| argilla-io/distilabel Distilabel is a Python framework for building scalable pipelines that generate synthetic data and AI feedback, based on verified research p… | 74 | 3385 | maintenance |
| pytorch-yolo-v3 A minimal PyTorch implementation of the YOLO v3 object detection algorithm, supporting detection on images and video with configurable reso… | 32 | 3312 | maintenance |
| isaac-sim/IsaacGymEnvs A collection of example reinforcement learning environments for NVIDIA Isaac Gym, a GPU-accelerated physics simulator. It provides a Gym-st… | 10 | 2953 | maintenance |
| jbesomi/texthero Texthero is a Python toolkit for text preprocessing, representation, and visualization, designed to work on top of Pandas Series and DataFr… | 23 | 2906 | maintenance |
| YCG09/chinese_ocr An end-to-end Chinese OCR system implemented with TensorFlow and Keras, combining CTPN for text detection with DenseNet + CTC for text reco… | 32 | 2780 | maintenance |
| zllrunning/video-object-removal A PyTorch application that removes objects from videos by drawing a bounding box around them. It combines SiamMask for object tracking and … | 32 | 2710 | maintenance |
| knazeri/edge-connect EdgeConnect is a PyTorch implementation of a two-stage generative adversarial model for image inpainting, published at ICCV 2019. It first … | 32 | 2620 | maintenance |
| MaybeShewill-CV/lanenet-lane-detection An unofficial TensorFlow implementation of the LaneNet deep neural network for real-time lane detection, based on the IEEE IV paper 'Toward… | 32 | 2563 | maintenance |
| young-geng/EasyLM EasyLM is a JAX/Flax-based framework for pre-training, finetuning, evaluating, and serving large language models like LLaMA. It scales trai… | 32 | 2515 | maintenance |
| TIBCOSoftware/flogo Project Flogo is an ultra-light, Go-based open source ecosystem for building event-driven applications using triggers and actions. It suppo… | 23 | 2491 | maintenance |
| CjangCjengh/MoeGoe MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V… | 23 | 2421 | maintenance |
| Roujack/mathAI mathAI is a photo-based math problem solver written in Python: it takes an image containing a handwritten or printed arithmetic expression,… | 32 | 2369 | maintenance |
| NVIDIA/waveglow WaveGlow is a PyTorch implementation of a flow-based generative network that synthesizes high-quality speech audio from mel-spectrograms, c… | 32 | 2339 | maintenance |
| Hzzone/pytorch-openpose A PyTorch reimplementation of OpenPose for body and hand pose estimation, with models converted directly from the original OpenPose caffemo… | 32 | 2320 | maintenance |
| facebookresearch/pyrobot PyRobot is a lightweight, high-level Python library providing hardware-independent APIs for robotic manipulation and navigation, built on t… | 10 | 2307 | maintenance |
| rsennrich/subword-nmt A Python library and CLI toolset for unsupervised word segmentation into subword units, best known for byte pair encoding (BPE) used in neu… | 23 | 2274 | maintenance |
| MetaGLM/FinGLM FinGLM is an open, community-driven financial LLM project centered on a dialog-based question-answering system that analyzes Chinese listed… | 27 | 2261 | maintenance |
| TigerResearch/TigerBot TigerBot is a multi-language, multi-task large language model project from TigerResearch, providing pretrained and chat-tuned model weights… | 29 | 2260 | maintenance |
| ankush-me/SynthText SynthText is a Python tool for generating synthetic scene-text images with ground-truth bounding boxes, as described in the CVPR 2016 paper… | 32 | 2146 | maintenance |
| Mukosame/Anime2Sketch Anime2Sketch is a PyTorch-based sketch extractor that converts anime art, illustrations, and manga into line drawings using pretrained GAN … | 32 | 2129 | maintenance |
| dotnet/spark .NET for Apache Spark provides high-performance C# and F# bindings for Apache Spark, exposing DataFrames, SparkSQL, and Structured Streamin… | 78 | 2096 | maintenance |
| facebookresearch/ELF ELF is an end-to-end, lightweight and flexible C++/Python platform for game research, focused on real-time strategy games. It hosts multipl… | 10 | 2090 | maintenance |
| jiupinjia/SkyAR SkyAR is the official PyTorch implementation of the paper 'Castle in the Sky: Dynamic Sky Replacement and Harmonization in Videos'. It perf… | 32 | 2025 | maintenance |
| guanshuicheng/invoice A Flask-based OCR microservice that recognizes Chinese VAT invoices (electronic, regular, and special) using a YOLOv3 + CRNN + CTC deep lea… | 32 | 1984 | maintenance |
| QwenLM/Qwen-Audio Official repository for Qwen-Audio, Alibaba Cloud's large audio-language model with pretrained and chat variants. It provides model weights… | 27 | 1948 | maintenance |
| carefree0910/carefree-creator carefree-creator is a Python library and CLI that serves AI image generation endpoints built on Stable Diffusion and related models, poweri… | 32 | 1934 | maintenance |