function: llm-inference
3362 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ConnorJL/GPT2 A community Python/TensorFlow implementation of GPT-2 model training and text generation that supports both GPUs and TPUs. It includes scri… | 32 | 1412 | maintenance |
| a16z-infra/llama2-chatbot A Streamlit-based chatbot web app for chatting with LLaMA 2 models (7B, 13B, 70B) hosted on Replicate API endpoints. It includes per-sessio… | 28 | 1410 | maintenance |
| vas3k/TaxHacker TaxHacker is a self-hosted AI-powered accounting app for freelancers, indie-hackers, and small businesses. It uses LLMs to OCR and analyze … | 85 | 6658 | experimental |
| nraiden/cofounder Cofounder is an AI-powered framework that generates full-stack web applications, including backend, database, and stateful frontend with ge… | 22 | 6657 | experimental |
| Yifan-Song793/RestGPT RestGPT is an LLM-based autonomous agent that controls real-world applications like TMDB and Spotify by planning and executing RESTful API … | 28 | 1404 | maintenance |
| bupticybee/ChineseAiDungeonChatGPT A Chinese-language clone of AI Dungeon that uses OpenAI's ChatGPT API as the storytelling model. It offers both a CLI and a simple GUI app … | 22 | 1404 | maintenance |
| neo4j/NaLLM NaLLM is a demo application from Neo4j exploring synergies between Neo4j graph databases and Large Language Models. It provides a backend A… | 10 | 1404 | maintenance |
| gotzmann/llama.go A pure-Go reimplementation of llama.cpp-style LLM inference, running LLaMA-family models on CPU without C++ dependencies. It includes tenso… | 21 | 1397 | maintenance |
| safety-research/bloom Bloom is a Python tool that automatically generates behavioral evaluation suites for LLMs, probing target models for behaviors like sycopha… | 55 | 1392 | maintenance |
| anus-dev/ANUS ANUS is a Grok-powered AI agent for the terminal, distributed as an npm CLI, whose stated goal is to evolve into an agent that contributes … | 38 | 6554 | experimental |
| 198808xc/Pangu-Weather Official implementation of Pangu-Weather, a deep learning model for fast, accurate medium-range global weather forecasting using 3D neural … | 22 | 1388 | maintenance |
| doonny/PipeCNN PipeCNN is an OpenCL-based FPGA accelerator for large-scale convolutional neural network inference, written in C with pipelined kernels. It… | 32 | 1386 | maintenance |
| nishiwen1214/ChatReviewer ChatReviewer is a Python application built on the ChatGPT API that analyzes academic papers, summarizing their strengths and weaknesses and… | 30 | 1377 | maintenance |
| keirp/automatic_prompt_engineer Automatic Prompt Engineer (APE) is a Python library implementing the research method from 'Large Language Models Are Human-Level Prompt Eng… | 32 | 1362 | maintenance |
| Jackywine/Bella Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper… | 48 | 6379 | experimental |
| strnad/CrewAI-Studio CrewAI Studio is a Streamlit-based graphical application for creating, managing, and running CrewAI multi-agent crews and tasks without wri… | 88 | 1346 | maintenance |
| yuvalsuede/ai-component-generator A Next.js web application that generates UI components (HTML, Tailwind CSS, React, Material UI) from free-form text prompts using OpenAI's … | 30 | 1346 | maintenance |
| chengzeyi/stable-fast Stable Fast is an ultra lightweight performance optimization framework for HuggingFace Diffusers pipelines on NVIDIA GPUs. It achieves SOTA… | 24 | 1302 | maintenance |
| ItsPi3141/alpaca-electron Alpaca Electron is an Electron desktop app that provides a simple ChatGPT-style UI for chatting with Alpaca and other LLaMA-based local LLM… | 21 | 1302 | maintenance |
| goldfishh/chatgpt-tool-hub An open-source Python tool engine that lets ChatGPT (and other LLMs) invoke tools like web search, math computation, code execution, and co… | 21 | 1267 | maintenance |
| UKPLab/EasyNMT EasyNMT is a Python library providing easy access to state-of-the-art neural machine translation for 100+ languages using models like Opus-… | 23 | 1260 | maintenance |
| andrewyng/translation-agent A Python demonstration library by Andrew Ng implementing an agentic machine translation workflow: an LLM translates text, reflects on its o… | 24 | 5804 | experimental |
| ShannonAI/service-streamer Service Streamer is a Python middleware that queues discrete web service requests into mini-batches for deep learning model inference, impr… | 32 | 1241 | maintenance |
| google-gemini/deprecated-generative-ai-js The deprecated Google AI JavaScript/TypeScript SDK for the Gemini API, now in limited maintenance with support ending November 30, 2025. Us… | 10 | 1237 | maintenance |
| srush/MiniChain MiniChain is a tiny Python library for building applications with large language models by annotating Python functions that call LLMs and c… | 21 | 1232 | maintenance |
| m1guelpf/auto-commit A Rust CLI tool that uses OpenAI's GPT-3.5 to automatically generate commit messages from your staged git changes. It supports dry-run outp… | 23 | 1221 | maintenance |
| nihui/realsr-ncnn-vulkan A command-line tool implementing the RealSR real-world super-resolution model using the ncnn inference framework with Vulkan GPU accelerati… | 23 | 1213 | maintenance |
| UMass-Embodied-AGI/3D-LLM 3D-LLM is the research code for a large language model that takes 3D representations (objects and scenes) as input, built on BLIP-2/LAVIS. … | 28 | 1212 | maintenance |
| dotnet/modernize-dotnet An AI-powered GitHub Copilot agent plugin that analyzes .NET applications, generates upgrade plans, and applies code changes to modernize t… | 77 | 1209 | maintenance |
| KwaiKEG/KwaiAgents KwaiAgents is an open-source suite from Kuaishou for building generalized information-seeking agents powered by LLMs. It includes KAgentSys… | 27 | 1200 | maintenance |
| mozilla-ai/any-agent any-agent is a Python library providing a single unified interface for building, serving, and evaluating AI agents across multiple agent fr… | 71 | 1197 | maintenance |
| LiuHC0428/LAW-GPT A Chinese-language legal conversational model (LawGPT_zh / 獬豸) fine-tuned from ChatGLM-6B with LoRA on legal Q&A datasets, including answer… | 30 | 1194 | maintenance |
| rockchip-linux/rknn-toolkit2 RKNN-Toolkit2 is Rockchip's software development kit for converting trained AI models to RKNN format and deploying them on Rockchip NPU chi… | 23 | 1193 | maintenance |
| visual-openllm/visual-openllm An open-source tool that interactively connects different visual models with an LLM, built on ChatGLM, Visual ChatGPT, and Stable Diffusion… | 30 | 1186 | maintenance |
| VoltaML/voltaML VoltaML is a lightweight Python library that compiles and optimizes machine learning and deep learning models for high-performance inferenc… | 32 | 1176 | maintenance |
| punica-ai/punica Punica is a Python system for serving many LoRA-finetuned LLMs from a single copy of the base model on one GPU, using a custom CUDA kernel … | 19 | 1175 | maintenance |
| IndieKKY/bilibili-subtitle A browser extension (哔哔君) that displays a clickable subtitle list panel for Bilibili videos, enabling jump-to-timestamp navigation, subtitl… | 52 | 1174 | maintenance |
| abhagsain/ai-cli A GPT-3/ChatGPT-powered command-line tool that answers CLI command questions directly from the terminal. Built with oclif in TypeScript and… | 23 | 1169 | maintenance |
| bestony/ChatGPT-Feishu A ChatGPT bot for Feishu (Lark) that lets users chat with OpenAI's models directly in Feishu messages. It is deployed as a single Node.js e… | 22 | 1169 | maintenance |
| uber-research/PPLM PPLM (Plug and Play Language Model) is a research implementation for controlled text generation that steers the topic and attributes of GPT… | 32 | 1153 | maintenance |
| sentient-agi/ROMA ROMA (Recursive Open Meta-Agent) is a Python meta-agent framework for building hierarchical, high-performance multi-agent systems. It enabl… | 51 | 5174 | experimental |
| peterw/Chat-with-Github-Repo A Python application that lets you chat with any GitHub repository. It clones a repo, embeds its documents into Activeloop Deep Lake, and s… | 30 | 1142 | maintenance |
| openinterpreter/01 An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python … | 16 | 5156 | experimental |
| lupantech/chameleon-llm Chameleon is a research framework for plug-and-play compositional reasoning with large language models, where an LLM planner synthesizes pr… | 20 | 1139 | maintenance |
| fixie-ai/ai-jsx AI.JSX is a JavaScript/TypeScript framework for building AI applications using JSX components, supporting prompt engineering, tools, docume… | 29 | 1133 | maintenance |
| ray-project/llmperf LLMPerf is a Python library for benchmarking and validating the performance of LLM APIs. It runs load tests measuring inter-token latency a… | 10 | 1127 | maintenance |
| microsoft/Windows-Machine-Learning Microsoft's Windows Machine Learning samples and tools repository, providing a high-performance ONNX inference API powered by ONNX Runtime … | 39 | 1123 | maintenance |
| google-deepmind/dramatron Dramatron is a research tool from DeepMind that uses pre-trained large language models to hierarchically co-write theatre scripts and scree… | 32 | 1112 | maintenance |
| PaddlePaddle/Paddle.js Paddle.js is a browser-based deep learning inference engine for Baidu PaddlePaddle models, running via WebGL, WebGPU, or WebAssembly backen… | 23 | 1103 | maintenance |
| rlancemartin/auto-evaluator A lightweight Streamlit-based evaluation tool for LLM question-answering chains built with LangChain. It auto-generates question-answer pai… | 30 | 1101 | maintenance |
| DonTizi/rlama RLAMA is a Go CLI tool for building and querying local Retrieval-Augmented Generation (RAG) systems over your documents using Ollama models… | 38 | 1095 | maintenance |
| WangRongsheng/XrayGLM XrayGLM is the first Chinese multimodal medical large language model that generates radiology report summaries from chest X-ray images, bui… | 29 | 1082 | maintenance |
| replit/ReplitLM Official repository with inference code, configs, and guides for the ReplitLM family of code-oriented language models, such as replit-code-… | 10 | 1082 | maintenance |
| openai/automated-interpretability OpenAI's code and tools for automatically generating, simulating, and scoring explanations of neuron behavior in language models, based on … | 10 | 1081 | maintenance |
| EdVince/Stable-Diffusion-NCNN A C++ implementation of Stable Diffusion using the NCNN inference framework, supporting both txt2img and img2img. It runs on x86 Windows ex… | 23 | 1067 | maintenance |
| ThePrimeagen/99 A Neovim plugin providing an agentic AI coding experience, integrating LLM providers like OpenCode and Claude Code into the editor. It augm… | 55 | 4749 | experimental |
| OpenBMB/VisCPM VisCPM is a family of open-source bilingual (Chinese/English) multimodal large models built on the 10B CPM-Bee language model, comprising V… | 29 | 1062 | maintenance |
| obiscr/ChatGPT A JetBrains IDE plugin that integrates ChatGPT into IntelliJ-based IDEs, providing an in-editor chat interface with OpenAI's language model… | 32 | 1059 | maintenance |
| huggingface/optimum-quanto Optimum Quanto is a PyTorch quantization backend for Hugging Face Optimum that quantizes model weights (int2/int4/int8/float8) and activati… | 70 | 1053 | maintenance |
| mpaepper/llm_agents A small Python library for building agents controlled by large language models, inspired by LangChain but implemented from scratch in very … | 43 | 1053 | maintenance |
| CSHaitao/LexiLaw LexiLaw is a fine-tuned Chinese legal large language model based on ChatGLM-6B, providing legal consultation and Q&A capabilities. It inclu… | 61 | 1041 | maintenance |
| dandelionsllm/pandallm Panda is an open-source project for overseas Chinese large language models, providing PandaLLM model weights (continued pretraining of LLaM… | 30 | 1031 | maintenance |
| microsoft/Llama-2-Onnx Microsoft's optimized ONNX export of Meta's Llama 2 models (7B and 13B, pretrained and fine-tuned, float16/float32), distributed via Git su… | 28 | 1026 | maintenance |
| zabirauf/AutoGPT.js AutoGPT.js is an open-source web application that runs AutoGPT-style autonomous GPT agents directly in the browser. It supports file creati… | 30 | 1024 | maintenance |
| awslabs/multi-model-server Multi Model Server (MMS) is a tool for serving deep learning model inference over HTTP endpoints, supporting models from any ML/DL framewor… | 10 | 1024 | maintenance |
| turtlesoupy/this-word-does-not-exist A project that trains a GPT-2 variant to invent fake English words with generated definitions and example sentences, powering the thiswordd… | 72 | 1023 | maintenance |
| CaviraOSS/OpenMemory OpenMemory is a local-first, self-hosted cognitive memory engine that gives LLM applications and AI agents persistent long-term memory, wit… | 70 | 4466 | experimental |
| cbh123/narrator A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL… | 69 | 4426 | experimental |
| varunshenoy/GraphGPT GraphGPT is a web application that converts unstructured natural language text into a knowledge graph using GPT-3, visualizing entities and… | 31 | 4426 | experimental |
| fcakyon/autollm AutoLLM is a Python library for rapidly building RAG-based LLM web apps and APIs, offering a unified API over 100+ LLM providers and 20+ ve… | 10 | 1005 | maintenance |
| susiai/susi_device SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc… | 32 | 1004 | maintenance |
| aiwaves-cn/RecurrentGPT RecurrentGPT is a Python implementation of a paper-proposed framework that simulates LSTM-style recurrence using natural language and promp… | 29 | 1001 | maintenance |
| build-with-groq/g1 g1 is a Python prototype that uses prompting strategies with Llama-3.1 70b on Groq to produce o1-like visible reasoning chains for solving … | 49 | 4174 | experimental |
| keijiro/AICommand A proof-of-concept Unity Editor extension that integrates ChatGPT, letting users control the Unity Editor via natural language prompts. It … | 30 | 4103 | experimental |
| Significant-Gravitas/Auto-GPT-Plugins A collection of official and community plugins for Auto-GPT, the autonomous AI agent framework, providing additional capabilities like web … | 10 | 3827 | experimental |
| llSourcell/Doctor-Dignity Doctor Dignity is a fine-tuned Llama2 7B model that can pass the US Medical Licensing Exam, built with PyTorch, Transformers, TRL, and ONNX… | 28 | 3822 | experimental |
| 0hq/WebGPT WebGPT is a vanilla JavaScript and HTML implementation of GPT transformer inference running in the browser via WebGPU, in under ~1500 lines… | 30 | 3792 | experimental |
| genspark-ai/genoffice GenOffice is a free, open-source AI office suite for macOS, Windows, and Linux built on Electron, offering a word processor, spreadsheet, p… | 80 | 3763 | experimental |
| apache/maka Apache Maka (Incubating) is a local-first AI agent workspace that runs agents on your machine through a single Runtime Host, executing tool… | 80 | 3591 | experimental |
| axtonliu/axton-obsidian-visual-skills A pack of prompt-based skills for Claude Code that generates Obsidian Canvas, Excalidraw, and Mermaid diagrams from text descriptions. It s… | 68 | 3550 | experimental |
| xjdr-alt/entropix Entropix is a research project implementing entropy-based sampling and parallel chain-of-thought decoding for large language models, aiming… | 22 | 3432 | experimental |
| danielgross/localpilot A local proxy server that emulates GitHub Copilot by running open-source code-completion LLMs (via llama.cpp/GGML) on a Mac, so VS Code's C… | 27 | 3348 | experimental |
| oboard/claude-code-rev A restored, runnable source tree of Anthropic's Claude Code CLI, reconstructed from source maps with shims filling unrecoverable modules. I… | 58 | 3286 | experimental |
| microsoft/amplifier Amplifier is a modular, extensible AI-powered development assistant delivered as a command-line tool from Microsoft. It provides chat and a… | 62 | 3117 | experimental |
| evilsocket/cake Cake is a multimodal AI inference server written in Rust that runs text, image, and voice models on a single device or shards them across a… | 59 | 3114 | experimental |
| daveshap/OpenAI_Agent_Swarm A framework for building a Hierarchical Autonomous Agent Swarm (HAAS) on top of OpenAI's agent APIs, where a core set of governed agents ca… | 10 | 3101 | experimental |
| OpenNSWM-Lab/FAROS FAROS is a blueprint-driven AutoResearch runtime that orchestrates end-to-end AI research workflows, from idea generation through experimen… | 58 | 3001 | experimental |
| b-nnett/grok-bot-0.18-reconstructed An unofficial, source-oriented reconstruction of the Grok Bot 0.18.0 macOS Electron app, with readable TypeScript implementations of its ru… | 57 | 2974 | experimental |
| RootbeerComputer/backend-GPT An experimental project that replaces a traditional backend and database with an LLM. Users describe their app's purpose, provide an initia… | 31 | 2932 | experimental |
| wquguru/nof0 NOF0 is an open-source AI trading arena that lets multiple LLM/agent-driven strategies compete in live cryptocurrency markets, each startin… | 42 | 2751 | experimental |
| ishan0102/vimGPT vimGPT is a Python application that lets GPT-4V browse and interact with the web visually using Playwright and the Vimium keyboard-navigati… | 27 | 2649 | experimental |
| Yeachan-Heo/gajae-code Gajae-Code (gjc) is an external coding-agent harness that runs AI coding agents in any repository using the coding subscription plans you a… | 76 | 2618 | experimental |
| mshumer/gpt-author A Jupyter notebook-based tool that chains GPT-4/Claude 3 and Stable Diffusion API calls to generate complete fantasy novels from a user pro… | 29 | 2530 | experimental |
| vercel-labs/fx fx is a tiny, open-source coding agent harness and CLI written in Zig, designed for minimalism, fast cold starts, and embeddability in larg… | 79 | 2495 | experimental |
| NVIDIA-NeMo/Switchyard Switchyard is a Rust proxy and library that routes LLM traffic across models and providers while translating between OpenAI Chat, Anthropic… | 80 | 2490 | experimental |
| semanser/codel Codel is a self-hosted, fully autonomous AI agent that executes complex tasks using a sandboxed Docker environment with terminal, browser, … | 16 | 2475 | experimental |
| IlyaRice/RAG-Challenge-2 A Python implementation of the winning RAG system from the Enterprise RAG Challenge 2 competition, answering questions about company annual… | 29 | 2431 | experimental |
| keijiro/AIShader A Unity editor plugin that generates shaders from prompts using the ChatGPT API. It is a proof-of-concept tool demonstrating LLM-driven sha… | 30 | 2389 | experimental |
| context-labs/autodoc Autodoc is an experimental CLI toolkit that auto-generates documentation for git repositories using large language models like GPT-4. It in… | 30 | 2360 | experimental |
| dvmazur/mixtral-offloading A Python library enabling efficient inference of Mixtral-8x7B mixture-of-experts language models on limited hardware like Google Colab or c… | 26 | 2332 | experimental |