# xlite-dev/Awesome-LLM-Inference

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Repository: https://github.com/xlite-dev/Awesome-LLM-Inference
Canonical: https://ross.abutalabs.com/products/awesome-llm-inference
Language: Python
License: GPL-3.0
License Family: copyleft
Topics: flash-attention, paged-attention, tensorrt-llm, vllm, awesome-llm, llm-inference, deepseek, flash-attention-3, deepseek-v3, minimax-01, deepseek-r1, mla, flash-mla, qwen3
Last push: 2026-08-14T12:23:49+00:00

## Health v2 (maintenance only)
Score: 73/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 97, release rhythm 40, longevity 78
- inputs: {"age_days": 1103, "days_push": 19, "days_rel": 442, "gap_med": 9.5, "n_releases_24m": 25}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 5478, forks 429 (observed 2026-08-28T04:09:19.682807+00:00)

## What it is
A curated list of research papers with code on LLM and VLM inference optimization, covering topics like FlashAttention, PagedAttention, quantization (WINT8/4, AWQ), and parallelism. It also provides a 500-page beginner PDF and a script to download all referenced papers.

## Use cases
- find papers on LLM inference optimization
- learn how FlashAttention and PagedAttention work
- research quantization methods like AWQ and SmoothQuant
- get started with LLM inference as a beginner
- collect PDFs of inference papers for offline reading
- track state-of-the-art inference techniques like MLA and FlashMLA

## When to choose
- you want a curated, categorized reading list of LLM inference papers with code links
- you are learning GPU inference optimization from scratch
- you need to survey recent techniques like FlashAttention 3, MLA, or continuous batching

## When to avoid
- you need a runnable inference engine rather than a paper collection
- you want production serving software like vLLM or TensorRT-LLM themselves
- you need tutorials unrelated to LLM/VLM inference

## Facets
- artifact type: learning-resource
- maturity: active
- function: llm-inference, developer-tools, documentation
- domain: large-language-models, deep-learning, tutorials, awesome-lists, gpu-computing
- platform: python, cross-platform
- tags: awesome-list, curated-papers, flash-attention, paged-attention, quantization, inference-optimization, vlm, gpu

## Member repositories
- xlite-dev/Awesome-LLM-Inference (main) score 73

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:19.682807+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:56:42.330537+00:00, confidence not recorded.
  - readme: https://github.com/xlite-dev/Awesome-LLM-Inference (fetched 2026-08-28T04:09:19.682807+00:00, sha 86073d442cb0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
