# hemingkx/SpeculativeDecodingPapers

📰 Must-read papers and blogs on Speculative Decoding ⚡️

Repository: https://github.com/hemingkx/SpeculativeDecodingPapers
Canonical: https://ross.abutalabs.com/products/speculativedecodingpapers
License: Apache-2.0
License Family: permissive
Last push: 2026-06-27T00:21:39+00:00

## Health v2 (maintenance only)
Score: 68/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 89, release rhythm 35, longevity 77
- inputs: {"age_days": 1078, "days_push": 68, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1291, forks 80 (observed 2026-08-28T04:04:15.706014+00:00)

## What it is
A regularly updated curated reading list (awesome list) of must-read papers, blogs, and tutorials on speculative decoding for efficient large language model inference. It accompanies an ACL 2024 Findings survey and organizes research by drafting method, application area, and features such as multi-token prediction, MoE, and long-context decoding.

## Use cases
- learn about speculative decoding for LLM inference
- find research papers on faster LLM token generation
- catch up on multi-token prediction and draft-model techniques
- prepare a related-work section on efficient LLM decoding
- find tutorials and blogs explaining speculative decoding
- track new publications on LLM inference acceleration

## When to choose
- you are starting research or a literature review on speculative decoding
- you need a curated, categorized, regularly updated bibliography with venue and method tags
- you want tutorials, slides, and blogs alongside academic papers

## When to avoid
- you need runnable inference acceleration code or a serving library rather than a reading list
- you need production LLM serving or quantization tooling
- you need benchmark results you can run rather than references to benchmark papers

## Facets
- artifact type: learning-resource
- maturity: active
- function: llm-inference, machine-learning, deep-learning, nlp
- domain: large-language-models, artificial-intelligence, awesome-lists, tutorials, performance
- platform: -
- tags: awesome-list, speculative-decoding, paper-list, survey, efficient-inference, multi-token-prediction, draft-models, reading-list, research, natural-language-processing

## Member repositories
- hemingkx/SpeculativeDecodingPapers (main) score 68

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:15.706014+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:54:47.257650+00:00, confidence not recorded.
  - readme: https://github.com/hemingkx/SpeculativeDecodingPapers (fetched 2026-08-28T04:04:15.706014+00:00, sha 02460ac2fbdf)
- Data as of 2026-08-30T08:39:29.467469+00:00.
