Ross ROSS = Recommend OSS · open-source software intelligence for agents

xlite-dev/Awesome-LLM-Inference resource

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉 observed · 2026-08-28

github.com/xlite-dev/Awesome-LLM-Inference · Python · GPL-3.0 (copyleft) observed · 2026-08-28

Health v2 · maintenance only

73/100

  • Activity 97
  • Release rhythm 40
  • Longevity 78
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 9.5
  • age_days: 1103
  • days_rel: 442
  • days_push: 19
  • n_releases_24m: 25

Full methodology

Adoption not part of the score

5478 stars · 429 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A curated list of research papers with code on LLM and VLM inference optimization, covering topics like FlashAttention, PagedAttention, quantization (WINT8/4, AWQ), and parallelism. It also provides a 500-page beginner PDF and a script to download all referenced papers.

Use cases

  • find papers on LLM inference optimization
  • learn how FlashAttention and PagedAttention work
  • research quantization methods like AWQ and SmoothQuant
  • get started with LLM inference as a beginner
  • collect PDFs of inference papers for offline reading
  • track state-of-the-art inference techniques like MLA and FlashMLA

When to choose

  • you want a curated, categorized reading list of LLM inference papers with code links
  • you are learning GPU inference optimization from scratch
  • you need to survey recent techniques like FlashAttention 3, MLA, or continuous batching

When to avoid

  • you need a runnable inference engine rather than a paper collection
  • you want production serving software like vLLM or TensorRT-LLM themselves
  • you need tutorials unrelated to LLM/VLM inference

Facets

learning-resource · maturity active

llm-inference developer-tools documentation large-language-models deep-learning tutorials awesome-lists gpu-computing python cross-platform awesome-list curated-papers flash-attention paged-attention quantization inference-optimization vlm gpu

1 source

Member repositories

RepositoryRoleHealth v2
xlite-dev/Awesome-LLM-Inferencemain73

For agents

markdown · JSON · MCP: product_card(name="xlite-dev/Awesome-LLM-Inference")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem