function: llm-inference
3362 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| kijai/ComfyUI-WanVideoWrapper ComfyUI custom nodes wrapping WanVideo (Wan2.1) and related video generation models. It provides a standalone sandbox for quickly implement… | 58 | 6678 | active |
| simplescaling/s1 s1 is an open-source research project implementing simple test-time scaling for large language models, including the s1K dataset of 1,000 c… | 33 | 6668 | active |
| microsoft/LLMLingua LLMLingua is a Microsoft library for prompt compression that removes non-essential tokens from prompts and KV caches to speed up LLM infere… | 53 | 6609 | active |
| julep-ai/julep Julep is an open-source framework and platform for building durable, composable AI agents as dataflows that can crash, resume, retry, and e… | 68 | 6591 | active |
| BlockRunAI/ClawRouter ClawRouter is an open-source TypeScript LLM router built for autonomous AI agents, offering 71 models from OpenAI, Anthropic, Google, xAI, … | 78 | 6571 | active |
| microsoft/call-center-ai An AI-powered call center service that lets you initiate or receive phone calls handled by an LLM-driven agent via a simple API call. Built… | 64 | 6563 | active |
| FareedKhan-dev/kimi-k3-in-c A dependency-free C99 inference engine that runs the 2.78-trillion-parameter Kimi K3 model on a single CPU with as little as 8 GB of RAM by… | 79 | 6524 | active |
| souzatharsis/podcastfy Podcastfy is an open-source Python package and CLI that transforms multimodal content (websites, PDFs, images, YouTube videos, topics) into… | 60 | 6521 | active |
| shibing624/pycorrector pycorrector is a Python toolkit for Chinese text error correction, supporting phonetic, visual, and grammar error types with models includi… | 84 | 6515 | active |
| Andyyyy64/whichllm whichllm is a Python CLI tool that auto-detects your GPU, CPU, and RAM, then ranks local LLMs from HuggingFace that will actually run and p… | 81 | 6488 | active |
| josStorer/RWKV-Runner RWKV-Runner is a lightweight desktop application that automates downloading, managing, and running RWKV large language models with a GUI. I… | 86 | 6462 | active |
| haifengl/smile SMILE is a comprehensive, high-performance machine learning framework for the JVM with idiomatic APIs for Java, Scala, and Kotlin. It cover… | 99 | 6413 | active |
| multimodal-art-projection/YuE YuE is a family of open-source foundation models based on the LLaMA2 architecture that generate full songs (up to five minutes) with vocals… | 32 | 6403 | active |
| kvcache-ai/Mooncake Mooncake is a KVCache-centric disaggregated serving platform for LLM inference, originally built to serve Kimi by Moonshot AI. It separates… | 91 | 6398 | active |
| drumih/turbo-fieldfare TurboFieldfare is a custom Swift + Metal runtime that runs Gemma 4 26B-A4B inference in roughly 2 GB of RAM on Apple Silicon Macs by stream… | 80 | 6397 | active |
| lavague-ai/LaVague LaVague is an open-source Python framework for building AI Web Agents powered by a Large Action Model. It combines a World Model that inter… | 26 | 6389 | active |
| antvis/Infographic AntV Infographic is a declarative infographic generation and rendering engine for the web, with a custom syntax, ~200 built-in templates, t… | 78 | 6385 | active |
| genkit-ai/genkit Genkit is an open-source framework by Google for building full-stack, AI-powered agentic applications in JavaScript/TypeScript, Go, Python,… | 87 | 6378 | stable |
| e2b-dev/fragments An open-source Next.js template for building AI-generated apps, similar to Claude Artifacts or v0. It uses the E2B SDK to securely execute … | 68 | 6370 | active |
| tensorflow/serving TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. It mana… | 86 | 6360 | stable |
| Arthur-Ficial/apfel apfel is a Swift CLI tool that exposes Apple's built-in on-device Foundation Models LLM on Apple Silicon Macs as a UNIX-friendly command, a… | 75 | 6325 | active |
| canopyai/Orpheus-TTS Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a… | 45 | 6314 | active |
| google-ai-edge/LiteRT-LM LiteRT-LM is Google's production-ready, high-performance open-source framework for running large language models on edge devices, built as … | 86 | 6298 | active |
| AIOS AIOS is an 'AI Agent Operating System' that embeds LLMs into an OS-like kernel providing scheduling, context switching, memory, storage, an… | 75 | 6291 | active |
| github/CopilotForXcode GitHub Copilot for Xcode is an AI coding assistant plugin for Xcode providing code completions, chat, agent mode, and code review for Swift… | 88 | 6285 | active |
| neuphonic/neutts NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit… | 60 | 6256 | active |
| meta-llama/llama3 The official Meta repository for Llama 3, providing model weights download scripts, tokenizer, and minimal example code for running inferen… | 10 | 29247 | maintenance |
| flashinfer-ai/flashinfer FlashInfer is a GPU kernel library and kernel generator for LLM inference, providing unified APIs for attention, GEMM, and MoE operations w… | 90 | 6252 | active |
| Sylinko/Everywhere Everywhere is a cross-platform desktop AI assistant with on-screen awareness that reads the context of whatever app you're using via access… | 83 | 6249 | active |
| meta-pytorch/gpt-fast A minimal (<1000 lines) PyTorch-native implementation of fast transformer text generation, demonstrating low-latency LLM inference with int… | 44 | 6249 | active |
| Shaunwei/RealChar RealChar is an open-source application for creating, customizing, and talking to AI characters/companions in realtime via voice or text. It… | 48 | 6213 | active |
| Eigenwise/atomic-agents Atomic Agents is a lightweight, modular Python framework for building agentic AI pipelines and applications from small, composable componen… | 91 | 6202 | active |
| Forget-C/Jellyfish Jellyfish is a self-hosted, end-to-end production workspace for AI-generated short dramas, covering script input, storyboard breakdown, cha… | 73 | 6195 | active |
| PrefectHQ/marvin Marvin is a Python framework for building AI applications with LLMs, offering structured-output utilities (extract, cast, classify, generat… | 88 | 6192 | active |
| modelscope/FunClip FunClip is an open-source, locally deployed video clipping tool that uses FunASR Paraformer models for speech recognition and subtitle gene… | 95 | 6190 | active |
| NEKOparapa/AiNiee AiNiee is an AI-powered translation tool focused on automatically translating complex long-form content such as RPG/SLG game text, Epub/TXT… | 94 | 6174 | active |
| microsoft/TaskWeaver TaskWeaver is a code-first agent framework from Microsoft that plans and executes data analytics tasks by generating and running code snipp… | 10 | 6172 | active |
| ByteDance-Seed/Bagel BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I… | 55 | 6159 | active |
| wenda-LLM/wenda Wenda is a self-hosted LLM invocation platform designed for efficient content generation in constrained environments, integrating local/off… | 31 | 6159 | active |
| microsoft/fara Fara1.5 is a family of open-weight computer use agent (CUA) models from Microsoft Research, released at 4B, 9B, and 27B scales and built on… | 58 | 6153 | active |
| microsoft/semantic-kernel Semantic Kernel is a model-agnostic SDK and orchestration framework from Microsoft for building AI agents and multi-agent systems in .NET, … | 93 | 28504 | maintenance |
| Beingpax/VoiceInk VoiceInk is a native macOS voice-to-text dictation app that transcribes speech to text almost instantly using local AI models (Parakeet, Wh… | 84 | 6115 | active |
| princeton-nlp/tree-of-thought-llm Official implementation of the Tree of Thoughts (ToT) framework for deliberate LLM problem solving, generalizing chain-of-thought prompting… | 20 | 6057 | stable |
| basketikun/chatgpt2api A self-hosted reverse-engineered implementation of ChatGPT's official web interfaces, exposing OpenAI-compatible API endpoints for text gen… | 77 | 6012 | active |
| PawanOsman/OpenCursor OpenCursor is an open-source, Cursor-like AI coding agent distributed as a VS Code extension. It provides agentic chat that reads workspace… | 96 | 6004 | active |
| gluonfield/enchanted Enchanted is an open-source iOS/macOS/visionOS chat app that connects to a self-hosted Ollama server to chat with private local language mo… | 75 | 5998 | active |
| z-lab/dflash DFlash is a lightweight block diffusion model used as a draft model for speculative decoding of large language models, drafting entire toke… | 71 | 5967 | active |
| kuafuai/DevOpsGPT DevOpsGPT is a multi-agent system that combines LLMs with DevOps tools to convert natural language requirements into working software. It a… | 19 | 5966 | active |
| YILING0013/AI_NovelGenerator A Python GUI application that uses large language models to generate multi-chapter long-form novels with automatic context continuity, fore… | 65 | 5960 | active |
| Conway-Research/automaton Automaton is a TypeScript runtime for continuously running, self-improving AI agents that pay for their own compute via an Ethereum wallet … | 73 | 5957 | active |
| microsoft/Webwright Webwright is a Python framework from Microsoft that turns coding LLMs into state-of-the-art browser agents by giving them a terminal to lau… | 57 | 5947 | active |
| cactus-compute/cactus Cactus is a hybrid edge-cloud AI inference engine for mobile devices, wearables, smart home devices, and robots, built in C++ with custom q… | 86 | 5934 | active |
| samchon/typia Typia is a TypeScript transformer library that compiles pure TypeScript types into super-fast runtime validators, JSON serializers/parsers,… | 95 | 5877 | active |
| TideDra/zotero-arxiv-daily A Python tool deployed as a GitHub Action that recommends new arXiv (and bioRxiv/medRxiv/chemRxiv) papers daily based on the contents of yo… | 81 | 5872 | active |
| MadcowD/ell ell is a lightweight Python library for language model programming that treats prompts as functions (language model programs) rather than s… | 34 | 5864 | active |
| sqlchat/sqlchat SQL Chat is a chat-based SQL client and editor that lets users query, modify, add, and delete data using natural language, powered by OpenA… | 65 | 5843 | active |
| iyaja/llama-fs LlamaFS is a self-organizing file manager that uses the Llama 3 model to automatically rename and organize files based on their content and… | 40 | 5836 | active |
| kserve/kserve KServe is a CNCF incubating, Kubernetes-native platform for serving both generative and predictive AI models at scale. It provides a standa… | 94 | 5834 | stable |
| OpenAI PHP OpenAI PHP is a community-maintained PHP SDK for interacting with the OpenAI API, covering chat, responses, embeddings, audio, images, fine… | 97 | 5829 | active |
| Mai-with-u/MaiBot MaiBot (MaiSaka) is an open-source, LLM-powered conversational agent designed to behave like a warm, human-like 'digital life form' in QQ g… | 87 | 5821 | active |
| SamuelSchmidgall/AgentLaboratory Agent Laboratory is an end-to-end autonomous research workflow framework that uses specialized LLM-driven agents to assist human researcher… | 38 | 5809 | active |
| Michael-A-Kuykendall/shimmy Shimmy is a single-binary, OpenAI-compatible LLM inference server written in pure Rust, running GGUF models on a WebGPU-based engine (Airfr… | 84 | 5808 | active |
| Giskard-AI/giskard-oss Giskard is an open-source Python library for testing, evaluating, and red-teaming LLM-based agents and RAG systems. Its v3 rewrite provides… | 95 | 5773 | active |
| amazon-science/chronos-forecasting Chronos is a Python library providing pretrained foundation models for time series forecasting, including Chronos-2 which handles univariat… | 88 | 5759 | active |
| 21st-dev/magic-mcp Magic MCP is a legacy compatibility proxy package that forwards MCP requests to the unified 21st MCP server, which lets AI coding agents li… | 63 | 5735 | active |
| memovai/mimiclaw MimiClaw is a bare-metal AI agent framework written in pure C that runs an OpenClaw-style personal AI assistant on a $5 ESP32-S3 microcontr… | 69 | 5731 | active |
| google/agents-cli A Python CLI plus agent skills that equip coding assistants (Claude Code, Codex, Antigravity, etc.) to build, evaluate, and deploy AI agent… | 81 | 5729 | active |
| google/gemma_pytorch The official PyTorch implementation of Google's Gemma family of open large language models, including text-only and multimodal variants. It… | 10 | 5719 | active |
| serge-chat/serge Serge is a self-hosted web chat interface for running LLMs locally via llama.cpp, built with a SvelteKit frontend and FastAPI backend. It s… | 10 | 5714 | active |
| HKUDS/AI-Researcher AI-Researcher is an autonomous research system that orchestrates the full scientific research pipeline—from literature review and hypothesi… | 41 | 5701 | active |
| google-deepmind/gemma The official JAX-based Python library from Google DeepMind for running, sampling from, and fine-tuning the Gemma family of open-weight larg… | 87 | 5695 | active |
| dnhkng/GLaDOS A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,… | 64 | 5689 | active |
| Q00/ouroboros Ouroboros is an 'Agent OS' that manages AI coding agent workflows through interview-gated requirement capture, staged evaluation, and a bud… | 78 | 5687 | active |
| mindcraft-bots/mindcraft Mindcraft is a JavaScript application that connects LLM-powered AI agents to Minecraft via the Mineflayer library, letting language models … | 81 | 5680 | active |
| OpenSenseNova/SenseNova-U1 SenseNova-U is a series of open-weight unified multimodal models (e.g., SenseNova-U1.5-8B-MoT) built on the NEO-unify architecture that com… | 59 | 5668 | active |
| nextai-translator/bob-plugin-openai-translator A Bob plugin for macOS that uses OpenAI, Google Gemini, MiniMax, or OpenAI-compatible APIs to translate text, polish wording, and fix gramm… | 96 | 5655 | active |
| ConnectAI-E/feishu-openai A self-hosted Go application that integrates OpenAI models (GPT-4, GPT-4V, DALL·E-3, Whisper) into Feishu/Lark as a chatbot. It supports vo… | 35 | 5638 | active |
| fla-org/flash-linear-attention A PyTorch library providing hardware-efficient implementations of emerging sequence model architectures, including linear attention, sparse… | 88 | 5627 | active |
| Vision-CAIR/MiniGPT-4 Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a sing… | 30 | 25627 | maintenance |
| sohzm/cheating-daddy Cheating Daddy is a free, open-source Electron desktop app that acts as a real-time AI assistant during video calls, interviews, and meetin… | 78 | 5580 | active |
| ngxson/smolvlm-realtime-webcam A browser-based demo that streams webcam frames to a llama.cpp server running SmolVLM 500M for real-time object detection and scene descrip… | 29 | 5571 | active |
| alexzhang13/rlm A plug-and-play Python inference library for Recursive Language Models (RLMs), a paradigm where an LLM programmatically examines and decomp… | 80 | 5539 | active |
| father-bot/chatgpt_telegram_bot A self-hostable Telegram bot that brings ChatGPT and Claude models into Telegram using your own OpenAI/Anthropic/OpenRouter API keys. It su… | 78 | 5531 | active |
| dograh-hq/dograh Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re… | 84 | 5505 | active |
| ATH-MaaS/ComfyUI-Copilot ComfyUI-Copilot is an AI-powered custom node for ComfyUI that acts as an intelligent assistant for building, debugging, and optimizing imag… | 53 | 5489 | active |
| HisMax/RedInk RedInk is a self-hostable web application that generates complete Xiaohongshu (Little Red Book) image-and-text posts from a single topic se… | 57 | 5481 | active |
| mostlygeek/llama-swap A Go-based proxy service that lets you run multiple local AI inference servers (llama.cpp, vllm, ComfyUI, etc.) and hot-swap between models… | 85 | 5480 | active |
| zarazhangrui/codebase-to-course A Claude Code skill that converts any codebase into a self-contained, interactive single-page HTML course for non-technical 'vibe coders'. … | 48 | 5464 | active |
| lm-sys/RouteLLM RouteLLM is a Python framework for serving and evaluating LLM routers that route queries between strong and weak models to cut costs. It pr… | 24 | 5408 | active |
| lingodotdev/lingo.dev Lingo.dev is an open-source localization engineering toolkit that turns LLMs into stateful translation APIs, persisting glossaries, brand v… | 87 | 5407 | active |
| TaskingAI/TaskingAI TaskingAI is an open-source Backend-as-a-Service (BaaS) platform for developing and deploying LLM-based AI agents. It unifies access to hun… | 17 | 5401 | active |
| docker/genai-stack A Docker Compose-based starter stack for building GenAI applications, combining LangChain, Ollama for local LLM inference, and Neo4j as a d… | 71 | 5397 | active |
| deepseek-ai/DeepSeek-VL2 DeepSeek-VL2 is a series of Mixture-of-Experts vision-language models (Tiny, Small, and 4.5B activated parameters) with inference code and … | 25 | 5374 | active |
| winfunc/deepreasoning DeepReasoning is a high-performance Rust-based LLM inference API and chat UI that combines DeepSeek R1's chain-of-thought reasoning traces … | 31 | 5363 | active |
| karpathy/llm-council A local web app that sends a user's question to multiple LLMs via OpenRouter, has them anonymously review and rank each other's answers, an… | 40 | 24369 | maintenance |
| PurpleAILAB/Decepticon Decepticon is an autonomous AI red-team hacking agent that uses LLMs (built on LangChain/LangGraph) to plan and execute context-aware offen… | 80 | 5335 | active |
| cloudflare/vibesdk An open-source SDK for building your own vibe-coding platform where users describe full-stack apps and an AI coding agent plans, edits, dep… | 73 | 5334 | active |
| vllm-project/semantic-router vLLM Semantic Router is an open-source programmable routing and control layer for building Mixture-of-Models systems across heterogeneous L… | 76 | 5315 | active |
| ysharma3501/LuxTTS LuxTTS is a lightweight zipvoice-based text-to-speech model for high-quality zero-shot voice cloning, generating clear 48kHz speech at up t… | 54 | 5306 | active |
| BuilderIO/ai-shell AI Shell is an open-source CLI tool that converts natural language prompts into shell commands using OpenAI, with options to run, revise, o… | 57 | 5283 | active |