resource: llm-inference
368 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| nichtdax/awesome-totally-open-chatgpt A curated awesome-list cataloging fully open-source alternatives to ChatGPT, covering instruct-finetuned language model projects. Entries a… | 30 | 4785 | maintenance |
| Azure-Samples/openai A collection of Jupyter Notebook code samples and starter templates for Azure OpenAI Service and Azure AI Foundry, complementing the OpenAI… | 74 | 1342 | active |
| autonomous-ai/autonomous-computer An open-source hardware project providing complete build guides (parts lists, CAD, BIOS settings, assembly photos) for Personal AI Computer… | 62 | 1330 | active |
| pinchbench/skill PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, … | 72 | 1325 | active |
| AutoTrustAI/PaperGuru-Benchmark PaperGuru is a benchmark and research repository for Lifecycle-Aware Memory (LAM), a long-term memory primitive for long-horizon LLM agents… | 53 | 1324 | active |
| cloudflare/agents-starter A starter template from Cloudflare for building AI chat agents on Workers using the Agents SDK and Durable Objects. It ships with streaming… | 64 | 1322 | active |
| NVIDIA/dgx-spark-playbooks A collection of step-by-step playbooks (Jupyter Notebook-based guides) for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blac… | 60 | 1299 | active |
| LiveBench/LiveBench LiveBench is a contamination-free benchmark for large language models that releases new questions monthly, drawn from recent datasets, pape… | 68 | 1294 | active |
| OpenBuddy/OpenBuddy OpenBuddy is a family of open multilingual chatbot LLMs fine-tuned from Falcon and LLaMA base models with extended vocabulary and CJK token… | 41 | 1294 | active |
| hemingkx/SpeculativeDecodingPapers A regularly updated curated reading list (awesome list) of must-read papers, blogs, and tutorials on speculative decoding for efficient lar… | 68 | 1291 | active |
| huybery/Awesome-Code-LLM A curated awesome-list of code large language models, including model rankings, evaluation toolkits, leaderboards, and research papers on p… | 29 | 1289 | active |
| elyase/awesome-gpt3 A curated awesome-list collecting demos, articles, and resources about the OpenAI GPT-3 API. It catalogs example applications such as code … | 10 | 4521 | maintenance |
| datawhalechina/diy-llm A Chinese-language, code-driven course (adapted from Stanford CS336) for systematically building large language models from scratch. It cov… | 79 | 1278 | active |
| strongdm/attractor Attractor is a set of Natural Language Specifications (NLSpecs) from StrongDM describing how to build a non-interactive coding agent that c… | 47 | 1277 | active |
| hahhforest/pi-textbook An open-source Chinese-language textbook ('动手学 Pi' / Build Your Own Pi) that teaches agent engineering by building a Pi-style coding agent … | 55 | 1267 | active |
| modal-labs/modal-examples A curated collection of example programs for Modal, a serverless cloud platform, covering use cases like LLM serving, image generation, spe… | 77 | 1264 | active |
| microsoft/generative-ai-with-javascript An open-source Microsoft course teaching Generative AI concepts to JavaScript developers through a time-travel themed narrative with lesson… | 68 | 1256 | active |
| ibm-granite/granite-code-models A family of open decoder-only code foundation models from IBM, released under Apache 2.0 in base and instruct variants at 3B, 8B, 20B, and … | 10 | 1251 | active |
| AgenticHealthAI/Awesome-AI-Agents-for-Healthcare A curated awesome-list of research papers, projects, and resources on agentic AI and AI agents applied to healthcare, maintained alongside … | 61 | 1231 | active |
| ray-project/llm-numbers A reference document listing key numbers and rules of thumb that LLM developers should know for back-of-the-envelope calculations, inspired… | 29 | 4315 | maintenance |
| THUDM/LongBench LongBench is a benchmark suite (v1 and v2) for evaluating large language models on long-context understanding and reasoning tasks, with con… | 29 | 1229 | active |
| qingsongedu/Awesome-TimeSeries-SpatioTemporal-LM-LLM A curated awesome-list of papers, code, and datasets on Large Language Models and Foundation Models applied to time series, spatio-temporal… | 29 | 1222 | active |
| datawhalechina/agentic-ai A Chinese-language tutorial project translating and organizing Andrew Ng's Agentic AI course series from DeepLearning.AI. It provides bilin… | 55 | 1215 | active |
| zukixa/cool-ai-stuff A curated awesome-list of free-to-use AI APIs and websites offering OpenAI-format chat and image generation, organized into tiers. It is a … | 49 | 1214 | active |
| CrazyBoyM/llama3-Chinese-chat A community repository releasing Chinese post-trained (SFT and DPO) versions of Llama3/Llama3.1, including open model weights, quantized GG… | 55 | 4146 | maintenance |
| a16z-infra/ai-getting-started A JavaScript/TypeScript starter stack for building AI weekend projects, combining Next.js, Clerk auth, Pinecone/pgvector vector stores, Lan… | 29 | 4141 | maintenance |
| aws/deep-learning-containers AWS Deep Learning Containers are pre-built, security-patched Docker images for running AI/ML workloads on AWS services like EC2, EKS, and S… | 96 | 1186 | active |
| togethercomputer/together-cookbook A collection of Jupyter notebook recipes and guides demonstrating how to build applications with open-source models using the Together AI A… | 63 | 1182 | active |
| MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark A deployment recipe for serving DeepSeek-V4-Flash-0731 across two NVIDIA DGX Spark nodes using vLLM with tensor parallelism and DSpark spec… | 58 | 1176 | active |
| handsome-rich/Awesome-Auto-Research-Tools A curated awesome-list of open-source projects that automate scientific research, spanning autonomous research systems, deep literature syn… | 59 | 1164 | active |
| DataExpert-io/llm-driven-data-engineering A public educational repository accompanying DataExpert's bootcamp course on LLM-driven data engineering concepts. It contains Python labs … | 28 | 1156 | active |
| tylerprogramming/ai A collection of tutorial projects and example code demonstrating Microsoft AutoGen and CrewAI multi-agent frameworks, accompanying a YouTub… | 53 | 1144 | active |
| huggingface/search-and-learn Search and Learn is a collection of Python scripts and YAML recipes from Hugging Face for scaling inference-time (test-time) compute of ope… | 59 | 1131 | active |
| SaladDay/pi-from-scratch A tutorial project that walks through building a minimal 600-line TypeScript coding agent (nano-pi) modeled on the production-grade pi agen… | 57 | 1131 | active |
| rasbt/mini-coding-agent A minimal, readable Python implementation of a coding agent harness that demonstrates the six core components of coding agents (repo contex… | 49 | 1126 | active |
| EdinburghNLP/awesome-hallucination-detection A curated awesome-list of papers on hallucination detection in large language models, with summaries of metrics, datasets, and methods. It … | 70 | 1125 | active |
| gpu-mode/awesomeMLSys A curated reading list of papers, videos, and repositories for onboarding into ML Systems, covering attention mechanisms, LLM inference, an… | 55 | 1120 | active |
| mind-protocol/terminal-velocity Terminal Velocity is a 100,000-word novel written entirely by a team of 10 autonomous AI agents, with the full manuscript and development h… | 21 | 1116 | stable |
| LangChain-OpenTutorial/LangChain-OpenTutorial An open-source tutorial repository of Jupyter notebooks teaching LangChain and LangGraph, from basics to advanced features. It is maintaine… | 42 | 1108 | active |
| bird-bench/BIRD-CRITIC-1 BIRD-CRITIC 1.0 (a.k.a. SWE-SQL) is a benchmark dataset of real-world SQL user issues for evaluating whether LLMs can diagnose and fix data… | 53 | 1098 | active |
| 0xSojalSec/LLMs-local A curated awesome-list of platforms, tools, models, and resources for running large language models locally. It covers inference engines, U… | 55 | 1087 | active |
| unipds-engenharia-de-ia-aplicada/engenharia-de-software-com-ia-aplicada A companion repository of code examples, demos, and reference links for the UNIPDS postgraduate program 'Engenharia de Software com IA Apli… | 60 | 1084 | active |
| jmaczan/tiny-vllm tiny-vllm is both a minimal high-performance LLM inference engine written in C++ and CUDA, and a hands-on course that walks through buildin… | 60 | 1081 | active |
| wdndev/tiny-llm-zh An educational project that implements a small-parameter Chinese large language model from scratch, covering the full pipeline: tokenizer t… | 25 | 1078 | active |
| carlini/yet-another-applied-llm-benchmark A personal benchmark of nearly 100 applied tests for evaluating how well large language models perform on practical tasks the author has ac… | 34 | 1064 | active |
| lmarena/arena-hard-auto Arena-Hard-Auto is an automatic benchmark for evaluating instruction-tuned LLMs using challenging real-world prompts and LLM-based judges (… | 39 | 1060 | active |
| wdndev/llama3-from-scratch-zh A Chinese-language Jupyter Notebook tutorial that implements the Llama3 8B model from scratch, translating naklecha's llama3-from-scratch g… | 24 | 1054 | stable |
| daronyondem/claude-architect-exam-guide A community-maintained study guide for the Claude Certified Architect – Foundations certification exam, written in Markdown and published a… | 71 | 1047 | active |
| honestsoul/generative_ai_project A Python project template providing a modular structure for building generative AI applications, with pre-built LLM clients for Claude and … | 62 | 1047 | active |
| xiaowu0162/LongMemEval LongMemEval is a benchmark of 500 high-quality questions for evaluating the long-term memory abilities of chat assistants across five skill… | 58 | 1035 | active |
| mryab/efficient-dl-systems Course materials for the Efficient Deep Learning Systems course taught at HSE University and Yandex School of Data Analysis. It covers GPU/… | 70 | 1028 | active |
| docker/compose-for-agents A collection of ready-to-use Docker Compose examples for building and running AI agents with open-source LLMs, MCP tools, and agent runtime… | 64 | 1028 | active |
| jaymody/picoGPT A tiny, minimal implementation of GPT-2 inference in plain NumPy, with the forward pass in ~40 lines of code. It is an educational companio… | 31 | 3473 | maintenance |
| krishnaik06/Roadmap-To-Learn-Agentic-AI A curated roadmap repository linking YouTube playlists and resources for learning to build agentic AI systems, covering Python, NLP, deep l… | 37 | 1015 | active |
| bird-bench/BIRD-Interact BIRD-INTERACT is an interactive Text-to-SQL benchmark that evaluates LLMs through dynamic multi-turn interactions with a simulated user, a … | 52 | 1011 | active |
| mrdbourke/simple-local-rag A tutorial repository by mrdbourke teaching how to build a retrieval-augmented generation (RAG) pipeline from scratch that runs entirely lo… | 25 | 1011 | active |
| fufankeji/LLMs-Technology-Community-Beyondata A Chinese-language LLM technology community repository offering full-pipeline tutorials for large language models, covering environment set… | 42 | 1009 | active |
| zjunlp/Prompt4ReasoningPapers A curated paper list accompanying an ACL 2023 survey on reasoning with language model prompting. It organizes research on prompt engineerin… | 42 | 1006 | active |
| noahshinn/reflexion Reference implementation of the NeurIPS 2023 Reflexion paper, where language agents improve through verbal self-reflection instead of weigh… | 31 | 3242 | maintenance |
| CVI-SZU/Linly Linly is a project releasing Chinese-adapted open large language models, including Chinese-LLaMA 1&2, Chinese-Falcon, Linly-OpenLLaMA base … | 21 | 3044 | maintenance |
| baichuan-inc/Baichuan-13B Baichuan-13B is an open-source, commercially usable 13-billion-parameter large language model from Baichuan Intelligence, released as both … | 29 | 2926 | maintenance |
| wenge-research/YAYI2 YAYI 2 is a family of open-source multilingual large language models (30B Base and Chat variants) developed by Wenge Research, pretrained o… | 10 | 2796 | maintenance |
| luban-agi/Awesome-Domain-LLM A curated awesome-list collecting open-source domain-specific large language models, datasets, and evaluation benchmarks across verticals l… | 28 | 2581 | maintenance |
| hyp1231/awesome-llm-powered-agent A curated awesome-list collecting papers, repositories, blogs, and benchmarks about LLM-powered agents, covering autonomous task solving, m… | 37 | 2255 | maintenance |
| LinkSoul-AI/Chinese-Llama-2-7b An open-source, commercially usable Chinese-adapted LLaMA 2 7B chat model with a bilingual Chinese-English SFT dataset, following the llama… | 28 | 2203 | maintenance |
| bleedline/Awesome-gptlike-shellsite A curated awesome-list of GPT 'shell site' projects (ChatGPT-style web UIs), API providers, and cloud server resources, with FAQs and monet… | 26 | 2185 | maintenance |
| km1994/LLMsNineStoryDemonTower A curated Chinese-language tutorial collection ('Nine-Story Demon Tower') covering hands-on practice with open-source LLMs such as ChatGLM,… | 30 | 2168 | maintenance |
| huggingface/evaluation-guidebook A guidebook from Hugging Face sharing practical and theoretical knowledge about LLM evaluation, gathered while managing the Open LLM Leader… | 47 | 2143 | maintenance |
| GitHubDaily/ChatGPT-Prompt-Engineering-for-Developers-in-Chinese An unofficial Chinese-English bilingual subtitle repository for DeepLearning.AI's 'ChatGPT Prompt Engineering for Developers' course by And… | 30 | 2119 | maintenance |
| XiongjieDai/GPU-Benchmarks-on-LLM-Inference A curated benchmark dataset comparing LLM inference speeds (tokens/s) across many NVIDIA GPUs and Apple Silicon chips using llama.cpp on LL… | 28 | 1935 | maintenance |
| fastai/lm-hackers A Jupyter notebook and accompanying video guide from fast.ai explaining how language models work, including tokenization, base models, inst… | 27 | 1869 | maintenance |
| VHellendoorn/Code-LMs A guide to using pre-trained large language models of source code, centered on the PolyCoder models with instructions for Hugging Face infe… | 32 | 1842 | maintenance |
| supabase-community/nextjs-openai-doc-search A Next.js starter template for building a custom ChatGPT-style documentation search over your own .mdx content. It generates OpenAI embeddi… | 67 | 1732 | maintenance |
| charent/ChatLM-mini-Chinese ChatLM-Chinese-0.2B is a small (0.2B parameter) Chinese conversational language model with fully open-sourced training pipeline code coveri… | 28 | 1728 | maintenance |
| samwit/langchain-tutorials A collection of Jupyter Notebook tutorials for LangChain accompanying a YouTube playlist. It covers hands-on examples of building LLM appli… | 29 | 1553 | maintenance |
| easychen/openai-gpt-dev-notes-for-cn-developer A Chinese-language developer guide/notes explaining how to quickly build OpenAI/GPT applications, covering the chat completions API, stream… | 30 | 1539 | maintenance |
| openai/SWELancer-Benchmark SWE-Lancer is a benchmark dataset and evaluation harness measuring whether frontier LLMs can complete real-world freelance software enginee… | 10 | 1431 | maintenance |
| adrianhajdin/project_openai_codex A tutorial companion repository from JS Mastery showing how to build and deploy a ChatGPT-style AI coding assistant web app using JavaScrip… | 31 | 1413 | maintenance |
| datawhalechina/hugging-multi-agent A Datawhale tutorial based on the MetaGPT multi-agent framework that teaches the concepts of AI agents and multi-agent systems through hand… | 26 | 1411 | maintenance |
| a16z-infra/llm-app-stack A curated reference list from a16z cataloging tools, projects, and vendors at each layer of the LLM application stack, from data pipelines … | 28 | 1319 | maintenance |
| samwit/llm-tutorials A collection of Jupyter Notebook tutorials accompanying Sam Witteveen's YouTube channel on large language models. It covers hands-on exampl… | 29 | 1163 | maintenance |
| google-deepmind/funsearch FunSearch is the official repository accompanying DeepMind's Nature 2023 paper on discovering new mathematical results via program search w… | 27 | 1110 | maintenance |
| Troyanovsky/Local-LLM-Comparison-Colab-UI A curated comparison of open-source LLMs that run on consumer hardware, with one-click Google Colab WebUI notebooks for trying each model. … | 57 | 1101 | maintenance |
| nolanaatama/sd-1click-colab A collection of Jupyter Notebook scripts for running Stable Diffusion with one-click setup on Google Colab. It automates installing and lau… | 31 | 1078 | maintenance |
| pengchengneo/Claude-Code An unofficial reconstruction of the full TypeScript source code of Anthropic's Claude Code CLI, recovered from the npm package's source map… | 48 | 1761 | experimental |
| adrianhajdin/project_ai_mern_image_generation A tutorial companion repository for building and deploying a full-stack MERN AI image generation app, a MidJourney/DALL-E clone using React… | 31 | 1196 | experimental |
| ctlllll/LLM-ToolMaker Research code for the LATM (LLMs as Tool Makers) framework, where a powerful LLM writes reusable Python tool functions and a cheaper LLM us… | 29 | 1064 | experimental |
| majacinka/crewai-experiments A collection of experiments with the CrewAI multi-agent framework, testing local models via Ollama and API models like GPT-4 and Gemini Pro… | 26 | 1017 | experimental |
| sourcegraph/awesome-code-ai A curated awesome-list of AI coding tools including assistants, code completion, refactoring, and testing utilities. The repository has bee… | 10 | 1694 | abandoned |
| qualisero/awesome-pi-agent A curated awesome list of add-ons, hooks, tools, skills, and resources for the pi coding agent (pi-mono). The maintainer has retired the li… | 10 | 1095 | abandoned |
| jingyaogong/minimind MiniMind is an open-source project that trains tiny (64M-parameter) large language models completely from scratch in pure PyTorch, covering… | 62 | 55036 | active |
| humanlayer/12-factor-agents A guide of twelve principles for building reliable, production-grade LLM-powered agent applications, inspired by 12 Factor Apps. It include… | 39 | 25510 | active |
| mikeroyal/Self-Hosting-Guide A curated guide (awesome-list style) covering self-hosting software and hardware, including containers, VPNs, LLMs, home automation, networ… | 45 | 22621 | active |
| Infrasys-AI/AISystem An open-source Chinese-language course (with Jupyter notebooks, slides, and videos) covering the full AI systems stack: AI chips and archit… | 50 | 17667 | active |
| genlayerlabs/genlayer-project-boilerplate A boilerplate project for building GenLayer applications, including an example intelligent contract (a football bets game) with web access … | 63 | 16751 | active |
| n8n-io/self-hosted-ai-starter-kit A Docker Compose template from n8n that spins up a local, self-hosted AI development environment bundling n8n, Ollama, Qdrant, and PostgreS… | 68 | 15208 | active |
| Arindam200/awesome-ai-apps A curated awesome-list of 131 open-source example projects, tutorials, and recipes for building LLM-powered applications such as agents, RA… | 64 | 13511 | active |
| LuckyOne7777/LLM-Trading-Lab A repository documenting a live 6-month experiment where ChatGPT managed a real-money micro-cap stock portfolio, including trade logs, dail… | 58 | 7498 | active |
| georgezouq/awesome-ai-in-finance A curated awesome-list of AI in finance resources, covering LLM agents, deep learning strategies, trading systems, market data sources, pap… | 75 | 6444 | active |
| SWE-bench/SWE-bench SWE-bench is a benchmark and evaluation harness that tests whether large language models can resolve real-world GitHub issues by generating… | 72 | 5719 | active |