function: llm-inference
3362 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Vahe1994/AQLM Official PyTorch implementation of AQLM, an extreme LLM compression method via additive quantization, extended with PV-Tuning for finetunin… | 57 | 1329 | active |
| Tongjilibo/bert4torch bert4torch is a PyTorch library providing an elegant reimplementation of transformer models (BERT, RoBERTa, T5, GPT, ChatGLM, LLaMA, etc.) … | 82 | 1328 | active |
| ChaokunHong/MetaScreener MetaScreener is an open-source AI tool that automates title/abstract and full-text PDF screening for systematic reviews using an ensemble o… | 71 | 1328 | active |
| facebookincubator/AITemplate AITemplate is a Python framework that compiles deep neural networks into high-performance CUDA (NVIDIA) or HIP (AMD) C++ code for fast fp16… | 66 | 4724 | maintenance |
| inference-labs-inc/dsperse DSperse is a proving-system-agnostic tool for verifiable AI that decomposes ONNX neural network models into circuit-compatible segments and… | 82 | 1327 | active |
| bugbasesecurity/pentest-copilot Pentest Copilot is an open-source, AI-driven penetration testing agent that connects to a Kali attack box, autonomously runs security tools… | 64 | 1327 | active |
| kodustech/kodus-ai Kodus (Kody) is an open-source AI code review agent that reviews pull requests on GitHub, GitLab, Bitbucket, and Azure Repos. It is model-a… | 82 | 1326 | active |
| Duxiaoman-DI/XuanYuan XuanYuan is a family of open-source Chinese financial-domain large language models from Duxiaoman, including base, chat, and quantized vari… | 29 | 1326 | active |
| AutumnWhj/ChatGPT-wechat-bot A WeChat bot that connects to ChatGPT/OpenAI's API so your WeChat account can reply to messages with AI-generated answers. It supports priv… | 53 | 4713 | maintenance |
| theroyallab/tabbyAPI TabbyAPI is a FastAPI-based OpenAI-compatible API server for running large language models locally using the ExllamaV3 backend on NVIDIA GP… | 71 | 1324 | active |
| ConardLi/easy-llm-cli Easy LLM CLI is an open-source command-line AI agent forked from Gemini CLI that works with multiple LLM providers, including Gemini, OpenA… | 48 | 1323 | active |
| neurogen-dev/NeuroAPI NeuroAPI is a Russian AI API gateway providing unified OpenAI/Anthropic/Gemini-compatible access to models like GPT, Claude, Gemini, and De… | 62 | 1321 | active |
| yukkcat/gemini-business2api A self-hosted gateway service that exposes Gemini Business through an OpenAI-compatible API, with multi-account load balancing and an admin… | 58 | 1321 | active |
| meta-pytorch/segment-anything-fast A fast, batched offline inference-oriented fork of Meta's Segment Anything (SAM) image segmentation model. It applies optimizations like bf… | 45 | 1321 | active |
| Robitx/gp.nvim Gp.nvim is a Neovim plugin that brings GPT-powered chat sessions, instructable text/code operations, speech-to-text, and image generation d… | 36 | 1320 | active |
| sanchit-gandhi/whisper-jax An optimized JAX implementation of OpenAI's Whisper speech recognition model, built on Hugging Face Transformers, offering up to 70x faster… | 30 | 4682 | maintenance |
| lean-dojo/LeanCopilot Lean Copilot integrates large language models natively into the Lean 4 theorem prover for proof automation. It provides tactic suggestion, … | 90 | 1316 | active |
| alibaba/rtp-llm RTP-LLM is Alibaba's high-performance LLM inference engine written in C++/CUDA, optimized with kernels like PagedAttention and FlashAttenti… | 66 | 1316 | active |
| evilsocket/nerve Nerve is a simple Agent Development Kit (ADK) for building, running, evaluating, and orchestrating LLM-based agents using YAML definitions … | 10 | 1316 | active |
| Capsize-Games/airunner AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat… | 79 | 1314 | active |
| aws-samples/claude-prompt-generator A Python web application that generates, translates, and iteratively evaluates prompts for Anthropic Claude 3 via AWS Bedrock, including co… | 10 | 1314 | active |
| huggingface/llm-vscode A VSCode extension providing LLM-powered features like Copilot-style ghost-text code completion, with configurable backends including the H… | 59 | 1313 | active |
| run-bigpig/jcp JCP (韭菜盘) is a cross-platform desktop application built with Wails, Go, and React that provides AI-driven analysis of Chinese A-share stock… | 64 | 1310 | active |
| JailbrokenAI/wallbreaker Wallbreaker is a Claude-Code-style terminal harness for red-teaming LLMs, driving an autonomous agent loop that runs jailbreak attacks (PAI… | 57 | 1310 | active |
| PentesterFlow/agent PentesterFlow is a terminal-based agentic AI CLI assistant for penetration testers and bug bounty hunters. It orchestrates LLM-driven recon… | 71 | 1308 | active |
| HugeCatLab/ChatTutor ChatTutor is a visual and interactive AI tutoring application that gives LLMs the ability to teach on an electronic whiteboard. It combines… | 54 | 1308 | active |
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| SSShooter/ebook-to-mindmap A browser-based AI application that parses EPUB and PDF ebooks and converts them into chapter summaries, per-chapter mind maps, or a single… | 72 | 1306 | active |
| meridianlabs-ai/inspect_petri Inspect Petri is an alignment auditing agent that automatically probes language models for concerning behaviors like sycophancy, reward hac… | 80 | 1303 | active |
| tjake/Jlama Jlama is a modern LLM inference engine written in Java, supporting popular transformer models like Llama, Mistral, Gemma, and Qwen with qua… | 51 | 1303 | active |
| openai/simple-evals A lightweight Python library from OpenAI for evaluating language models against benchmarks like MMLU, GPQA, MATH, HumanEval, SimpleQA, Heal… | 60 | 4612 | maintenance |
| turboderp-org/exllamav2 ExLlamaV2 is a fast Python inference library for running large language models locally on modern consumer GPUs, with support for EXL2/GPTQ … | 61 | 4611 | maintenance |
| SakanaAI/text-to-lora Text-to-LoRA (T2L) is a hypernetwork that generates LoRA adapters for large language models in a single forward pass, using only a natural … | 30 | 1300 | active |
| stenolabs/stenoai Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire… | 84 | 1299 | active |
| dabit3/react-native-ai React Native AI is a full-stack framework for building cross-platform mobile AI apps with React Native and an Express server proxy. It prov… | 69 | 1298 | active |
| DimiMikadze/orca Orca is an AI agent application for deep LinkedIn profile analysis that scrapes posts, comments, reactions, and interaction networks, then … | 75 | 1297 | active |
| foxhui/WebAI2API WebAI2API is a self-hosted Node.js service that exposes web-based AI services (LMArena, Gemini, ChatGPT, DeepSeek, etc.) as OpenAI-compatib… | 57 | 1292 | active |
| fastclaw-ai/fastclaw FastClaw is a lightweight AI agent runtime written in Go, distributed as a single binary. It acts as an agent factory that creates, manages… | 75 | 1291 | active |
| Exafunction/windsurf.nvim A native Neovim plugin for Windsurf (formerly Codeium) that provides AI-powered code completion and chat inside the editor. It integrates w… | 65 | 1291 | active |
| panyanyany/Twocast Twocast is a self-hostable AI podcast generator that creates two-person conversational podcast episodes from topics, links, documents, or w… | 32 | 1290 | active |
| TencentQQGYLab/ELLA ELLA is an Efficient Large Language Model Adapter that equips text-to-image diffusion models with LLM-based text understanding via a Timest… | 25 | 1289 | active |
| xunbu/docutranslate DocuTranslate is a lightweight local document translation tool powered by large language models, supporting formats such as pdf, docx, xlsx… | 80 | 1285 | active |
| DropbaseHQ/dropbase Dropbase is a local-first, self-hosted platform that uses AI to generate Python web apps, combining a drag-and-drop UI builder with editabl… | 27 | 1285 | active |
| EvanZhouDev/openai-oauth A TypeScript toolkit that turns a ChatGPT account into an OpenAI-compatible API via a local dev proxy, TypeScript SDK, and client adapters.… | 65 | 1283 | active |
| Xiangyue-Zhang/auto-deep-researcher-24x7 An open-source Python framework where LLM agents autonomously run deep learning experiments 24/7, covering hypothesis formation, code imple… | 52 | 1283 | active |
| coleam00/claude-memory-compiler A Python tool that gives Claude Code an evolving memory by capturing conversation transcripts via hooks, extracting decisions and lessons w… | 48 | 1283 | active |
| poetiq-ai/poetiq-arc-agi-solver A Python research codebase that reproduces Poetiq's record-breaking, leaderboard-topping submissions to the ARC-AGI-1 and ARC-AGI-2 abstrac… | 42 | 1282 | active |
| hyperonym/basaran Basaran is an open-source alternative to the OpenAI text completion API that serves Hugging Face Transformers-based text generation models … | 10 | 1282 | active |
| charmbracelet/mods Mods is a Go CLI tool that brings LLM-based AI to the command line, designed to work in shell pipelines by ingesting command output and for… | 10 | 4524 | maintenance |
| lasantosr/intelli-shell IntelliShell is a command template and snippet manager for shells, offering smart search, dynamic variable completions, bookmarks, and AI-p… | 93 | 1279 | active |
| perminder-klair/subwave SUB/WAVE is a self-hostable personal internet radio station that broadcasts a single shared Icecast stream to all listeners simultaneously.… | 77 | 1278 | active |
| SanMuzZzZz/LuaN1aoAgent LuaN1aoAgent is an autonomous AI-driven penetration testing agent built in TypeScript on the Pi SDK, using graph-based cognitive reasoning … | 81 | 1276 | active |
| rinadelph/Agent-MCP Agent-MCP is a TypeScript framework for building coordinated multi-agent AI systems over the Model Context Protocol (MCP). It provides a pe… | 56 | 1276 | active |
| neo4j/neo4j-graphrag-python The official Neo4j first-party Python library for building graph retrieval-augmented generation (GraphRAG) applications. It provides retrie… | 93 | 1275 | active |
| PrunaAI/pruna Pruna is an open-source Python model optimization framework that makes AI models faster, smaller, cheaper, and greener via caching, quantiz… | 84 | 1275 | active |
| giuseppe99barchetta/SuggestArr SuggestArr is a self-hosted web application that automatically recommends and requests movies, TV shows, and anime to Jellyseerr/Overseerr … | 89 | 1274 | active |
| turboderp-org/exllamav3 ExLlamaV3 is a Python library for fast quantization and inference of large language models on consumer-class GPUs, featuring the EXL3 quant… | 87 | 1274 | active |
| ai-genie/chatgpt-vscode Genie AI is a Visual Studio Code extension that integrates OpenAI's ChatGPT models (GPT-4o, GPT-4, GPT-3.5, o1) as an AI pair programmer. I… | 21 | 1274 | active |
| ABexit/ASR-LLM-TTS An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) … | 60 | 1271 | active |
| datalayer/jupyter-mcp-server A Model Context Protocol (MCP) server that lets AI agents connect to, edit, and execute Jupyter Notebooks in real time. It supports STDIO a… | 88 | 1270 | active |
| okisdev/ChatChat Chat Chat is a self-hostable web platform that unifies chat and search interfaces to multiple AI providers such as OpenAI, Anthropic, Coher… | 65 | 1270 | active |
| Storia-AI/sage Sage is an open-source tool that lets developers chat with any codebase to quickly understand how it works, similar to a self-hosted GitHub… | 10 | 1268 | active |
| vipshop/cache-dit Cache-DiT is a PyTorch-native inference engine that accelerates Diffusion Transformer (DiT) models with hybrid caching, parallelism, quanti… | 82 | 1267 | active |
| DreamLM/Dream Dream 7B is an open diffusion large language model (dLLM) with base and instruct checkpoints, plus inference and training code built on Hug… | 44 | 1265 | active |
| jofizcd/Soul-of-Waifu Soul of Waifu is an open-source desktop application for creating AI companion characters with Live2D/VRM avatars, voice chat, and persisten… | 94 | 1263 | active |
| muscriptor/muscriptor MuScriptor is a multi-instrument music transcription model by Kyutai and Mirelo that converts audio recordings into MIDI and sheet music. I… | 79 | 1263 | active |
| stair-lab/kg-gen kg-gen is a Python library that extracts knowledge graphs from arbitrary plain text or conversation messages using LLMs, with model routing… | 63 | 1260 | active |
| razzant/ouroboros Ouroboros is an open-source, general-purpose AI agent with persistent identity, durable memory, and the ability to rewrite its own implemen… | 79 | 1258 | active |
| yym68686/uni-api uni-api is a self-hosted Python service that unifies management of multiple LLM provider APIs behind a single OpenAI-compatible endpoint. I… | 86 | 1257 | active |
| kubeai-project/kubeai KubeAI is a Kubernetes operator for serving machine learning models in production, supporting LLMs via vLLM and Ollama, vector embeddings, … | 89 | 1256 | active |
| Xerxes-2/clewdr ClewdR is a high-performance Rust proxy for Claude.ai and Claude Code that exposes native Claude and OpenAI-compatible API endpoints from a… | 83 | 1256 | active |
| tigicion/dao-code Dao Code is an open-source terminal-native AI coding agent written in TypeScript, targeting DeepSeek V4 with 1M context and Claude Code-com… | 76 | 1254 | active |
| facebookresearch/llm-transparency-tool An interactive toolkit from Meta Research for analyzing the internal workings of Transformer-based language models. It visualizes contribut… | 10 | 1254 | active |
| HelgeSverre/ollama-gui Ollama GUI is a web-based chat interface for interacting with local LLMs served through the Ollama API. It offers local chat history via In… | 70 | 1253 | active |
| TencentCloudADP/youtu-graphrag Youtu-GraphRAG is a Python framework for graph-based retrieval-augmented generation that unifies graph schema construction, community detec… | 57 | 1253 | active |
| Neroued/ninfer NInfer is a from-scratch C++/CUDA inference engine optimized for maximum single-GPU performance on a narrow, explicitly registered set of Q… | 58 | 1252 | active |
| FoloUp/FoloUp FoloUp is an open-source, self-hostable web application that conducts AI-powered voice interviews with job candidates. It generates tailore… | 58 | 1250 | active |
| ModelCloud/GPTQModel GPTQModel is a production-ready Python toolkit for quantizing (compressing) large language models using GPTQ, AWQ, and related methods, wit… | 91 | 1248 | active |
| Blueturboguy07/cue cue is an open-source desktop AI copilot that floats a glass panel over your screen, capturing your screen, microphone, and meeting audio t… | 78 | 1247 | active |
| HJYao00/Mulberry Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect… | 49 | 1243 | active |
| RUC-NLPIR/Search-o1 Search-o1 is a research framework that enhances large reasoning models (like QwQ and R1) with agentic search and retrieval-augmented genera… | 44 | 1241 | active |
| turnstonelabs/turnstone Turnstone is a self-hosted, local-first orchestration platform for tool-using AI agents, letting LLMs use real tools like shell, files, sea… | 78 | 1239 | active |
| xllm-go/bypass A Go server that reverse-engineers the chat interfaces of multiple AI providers (Coze, DeepSeek, Cursor, Windsurf, Grok, Bing Copilot, You,… | 62 | 1239 | active |
| 0xacx/chatGPT-shell-cli A lightweight shell script that lets you chat with OpenAI's ChatGPT models and generate DALL-E images directly from the terminal, requiring… | 32 | 1238 | active |
| pytorch/serve TorchServe is a flexible, production-ready model server for serving, optimizing, and scaling PyTorch models over HTTP with support for CPU,… | 10 | 4346 | maintenance |
| mrwadams/attackgen AttackGen is a Streamlit-based cybersecurity tool that uses large language models to generate tailored incident response testing scenarios … | 95 | 1237 | active |
| astaxie/TokenHub TokenHub is a self-hosted enterprise AI gateway written in Go that provides a governance layer over AI model access, including model routin… | 81 | 1237 | active |
| tsinghua-fib-lab/AgentSociety AgentSociety is an LLM-native agent simulation framework for building large-scale multi-agent simulations of social and urban environments.… | 85 | 1236 | active |
| Ksuriuri/index-tts-vllm A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and… | 58 | 1235 | active |
| MoonshotAI/FlashKDA FlashKDA is a set of high-performance CUDA kernels (built on CUTLASS) implementing Kimi Delta Attention, a linear attention mechanism, for … | 57 | 1229 | active |
| inclusionAI/AWorld AWorld is an open-source agent framework and runtime that orchestrates tools, memory, context, and execution for building autonomous AI age… | 78 | 1227 | active |
| microsoft/MInference MInference is a Microsoft library that accelerates long-context LLM inference using dynamic sparse attention, reducing pre-fill latency by … | 49 | 1226 | active |
| apple/python-apple-fm-sdk Python bindings for Apple's Foundation Models framework, giving access to the on-device foundation model behind Apple Intelligence on macOS… | 75 | 1225 | active |
| run-house/kubetorch Kubetorch is a Python library that lets you distribute and run ML workloads (training, inference, data processing) on Kubernetes directly f… | 83 | 1224 | active |
| MoonshotAI/Kimi-VL Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo… | 33 | 1224 | active |
| s2b-dev/smart-second-brain An open-source Obsidian plugin that adds a privacy-focused AI assistant to your note vault. It embeds your notes and uses a RAG pipeline wi… | 81 | 1223 | active |
| ipa-lab/hackingBuddyGPT HackingBuddyGPT is a Python framework that helps ethical hackers and security researchers use LLMs and LLM-based autonomous agents for pene… | 67 | 1223 | active |
| mrousavy/react-native-fast-tflite A high-performance TensorFlow Lite library for React Native built on Nitro Modules, using the low-level C/C++ TFLite core API with zero-cop… | 84 | 1222 | active |
| allwefantasy/auto-coder Auto-Coder is an open-source AI coding agent and CLI tool (powered by Byzer-LLM) that provides chat, one-shot command, server, and RAG mode… | 70 | 1222 | active |
| jieyefriic/rp-engine A YAML-native agent workflow execution engine written in Rust that parses declarative workflow files describing nodes, edges, prompts, data… | 51 | 1221 | active |