Ross ROSS = Recommend OSS · open-source software intelligence for agents

vllm-project/vllm-metal

Community maintained hardware plugin for vLLM on Apple Silicon observed · 2026-08-28

github.com/vllm-project/vllm-metal · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

79/100

  • Activity 99
  • Release rhythm 87
  • Longevity 18
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 0
  • age_days: 264
  • days_rel: 7
  • days_push: 7
  • n_releases_24m: 458

Full methodology

Adoption not part of the score

1647 stars · 228 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A community-maintained hardware plugin that enables vLLM to run LLM inference on Apple Silicon Macs using MLX as the primary compute backend. It unifies MLX and PyTorch under a single lowering path with optimized Metal kernels for attention and prefill.

Use cases

  • serve LLMs locally on a Mac with Apple Silicon
  • run vLLM inference on M-series chips
  • host a 27B model on a single Mac
  • use MLX as a backend for vLLM
  • self-host an OpenAI-compatible LLM server on macOS
  • benchmark LLM throughput on Apple hardware

When to choose

  • you want vLLM's serving stack on Apple Silicon instead of NVIDIA GPUs
  • you need local LLM inference on macOS with high throughput
  • you want to leverage MLX kernels and unified memory for large models

When to avoid

  • you are deploying on Linux servers with NVIDIA/AMD GPUs (use core vLLM)
  • you need x86_64 or Rosetta support
  • you run macOS earlier than 15 Sequoia or non-arm64 Python

Facets

plugin · maturity active

llm-inference machine-learning sdk large-language-models machine-learning apple-ecosystem python vllm mlx apple-silicon metal hardware-plugin llm-serving macos gpu

2 sources

Member repositories

RepositoryRoleHealth v2
vllm-project/vllm-metalmain79

For agents

markdown · JSON · MCP: product_card(name="vllm-project/vllm-metal")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem