Ross ROSS = Recommend OSS · open-source software intelligence for agents

vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend observed · 2026-08-28

github.com/vllm-project/vllm-ascend · homepage · C++ · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

83/100

  • Activity 99
  • Release rhythm 86
  • Longevity 41
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 90.0
  • age_days: 581
  • days_rel: 17
  • days_push: 7
  • n_releases_24m: 7

Full methodology

Adoption not part of the score

2711 stars · 2121 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

vllm-ascend is a community-maintained hardware plugin that enables vLLM to run large language model inference on Huawei Ascend NPUs. It implements the vLLM hardware-pluggable interface, supporting Transformer, MoE, embedding, and multi-modal models on Atlas hardware.

Use cases

  • serve llm inference on huawei ascend npu
  • run vllm on atlas 800 hardware
  • deploy deepseek or qwen models on ascend npus
  • set up prefill-decode disaggregation with mooncake
  • run multi-node distributed inference with ray on npus
  • accelerate decoding with suffix speculative decoding
  • serve multimodal llms on non-gpu accelerators

When to choose

  • you have Ascend NPU hardware (Atlas 800 A2/A3, Atlas 300I DUO) and want vLLM serving
  • you need PD disaggregation, expert parallelism, or multi-node inference on Ascend
  • you want the officially recommended Ascend backend for vLLM

When to avoid

  • you are running on NVIDIA/AMD GPUs - plain vLLM already supports those natively
  • you cannot install the required CANN/TorchNPU software stack
  • you need Triton-based features, which are not supported on this backend

Facets

plugin · maturity active

llm-inference plugin-system machine-learning large-language-models artificial-intelligence gpu-computing developer-tools python cpp vllm ascend-npu huawei hardware-plugin model-serving torch-npu pd-disaggregation speculative-decoding linux docker gpu

10 sources

Member repositories

RepositoryRoleHealth v2
vllm-project/vllm-ascendmain83

For agents

markdown · JSON · MCP: product_card(name="vllm-project/vllm-ascend")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem