vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend observed · 2026-08-28
Health v2 · maintenance only
83/100
- Activity 99
- Release rhythm 86
- Longevity 41
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 90.0
- age_days: 581
- days_rel: 17
- days_push: 7
- n_releases_24m: 7
Adoption not part of the score
2711 stars · 2121 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
vllm-ascend is a community-maintained hardware plugin that enables vLLM to run large language model inference on Huawei Ascend NPUs. It implements the vLLM hardware-pluggable interface, supporting Transformer, MoE, embedding, and multi-modal models on Atlas hardware.
Use cases
- serve llm inference on huawei ascend npu
- run vllm on atlas 800 hardware
- deploy deepseek or qwen models on ascend npus
- set up prefill-decode disaggregation with mooncake
- run multi-node distributed inference with ray on npus
- accelerate decoding with suffix speculative decoding
- serve multimodal llms on non-gpu accelerators
When to choose
- you have Ascend NPU hardware (Atlas 800 A2/A3, Atlas 300I DUO) and want vLLM serving
- you need PD disaggregation, expert parallelism, or multi-node inference on Ascend
- you want the officially recommended Ascend backend for vLLM
When to avoid
- you are running on NVIDIA/AMD GPUs - plain vLLM already supports those natively
- you cannot install the required CANN/TorchNPU software stack
- you need Triton-based features, which are not supported on this backend
Facets
plugin · maturity active
llm-inference plugin-system machine-learning large-language-models artificial-intelligence gpu-computing developer-tools python cpp vllm ascend-npu huawei hardware-plugin model-serving torch-npu pd-disaggregation speculative-decoding linux docker gpu
10 sources
- readme: https://github.com/vllm-project/vllm-ascend · fetched 2026-08-28 · 25387d25474b
- homepage: https://docs.vllm.ai/projects/ascend · fetched 2026-08-29 · 0c3bd92af16b
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/installation.html · fetched 2026-08-29 · fad6dd84f837
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/index.html · fetched 2026-08-29 · ab4ab23e427d
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_colocated_mooncake_multi_instance.html · fetched 2026-08-29 · 085f81c41e0f
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_single_node.html · fetched 2026-08-29 · 8be6c39d6639
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html · fetched 2026-08-29 · a55e81d79f27
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/dynamic_chunked_pipeline_parallel.html · fetched 2026-08-29 · 0610f65f988f
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/suffix_speculative_decoding.html · fetched 2026-08-29 · 8e16d7466199
- site_page: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/ray.html · fetched 2026-08-29 · e13acd3e2316
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| vllm-project/vllm-ascend | main | 83 |
For agents
markdown · JSON · MCP: product_card(name="vllm-project/vllm-ascend")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem