# ROCm/FastFlowLM

Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.

Repository: https://github.com/ROCm/FastFlowLM
Canonical: https://ross.abutalabs.com/products/fastflowlm
Homepage: http://fastflowlm.com/
Language: C++
License: MIT
License Family: permissive
Topics: amd, deepseek, llama, llm, npu
Last push: 2026-08-26T19:04:56+00:00

## Health v2 (maintenance only)
Score: 85/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 99, longevity 31
- inputs: {"age_days": 443, "days_push": 7, "days_rel": 7, "gap_med": 7.0, "n_releases_24m": 51}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1809, forks 144 (observed 2026-08-28T04:05:39.193454+00:00)

## What it is
FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama-like single-command CLI experience. It supports text, vision, audio (Whisper), embedding, and MoE models with context lengths up to 256k tokens in an ultra-lightweight ~17 MB runtime.

## Use cases
- run llms locally on amd ryzen ai npu
- ollama alternative for amd npu
- run whisper speech transcription on npu
- run vision models like gemma3 on npu
- serve local llm with openai-compatible api
- power-efficient on-device llm inference without gpu
- run moe models like gpt-oss on npu

## When to choose
- you have an AMD Ryzen AI chip with XDNA2 NPU (Strix, Strix Halo, Kraken, Gorgon Point) and want maximum efficiency
- you want a simple Ollama-like CLI for local LLMs without a GPU
- power efficiency and long context (up to 256k) matter more than broad hardware support

## When to avoid
- you don't have AMD Ryzen AI NPU hardware
- you need NVIDIA/Apple Silicon or CPU inference
- you need fine-grained control over inference engines like llama.cpp or vLLM

## Facets
- artifact type: application
- maturity: active
- function: llm-inference, cli, speech-recognition, machine-learning, http-server
- domain: large-language-models, machine-learning, developer-tools, speech-processing, computer-vision
- platform: windows, cli
- tags: npu, amd-ryzen-ai, ollama-alternative, on-device-inference, local-llm, xdna2, moe, whisper, embeddings, linux, amd

## Member repositories
- ROCm/FastFlowLM (main) score 85

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:39.193454+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:21:19.028633+00:00, confidence not recorded.
  - readme: https://github.com/ROCm/FastFlowLM (fetched 2026-08-28T04:05:39.193454+00:00, sha bec19da4b7d8)
  - homepage: http://fastflowlm.com/ (fetched 2026-08-29T11:00:07.625918+00:00, sha 75069f1b0af9)
  - site_page: http://fastflowlm.com/docs (fetched 2026-08-29T11:00:07.635071+00:00, sha 2c4810fbe895)
  - site_page: http://fastflowlm.com/docs/models (fetched 2026-08-29T11:00:07.636926+00:00, sha e505b7649e30)
  - site_page: http://fastflowlm.com/docs/benchmarks (fetched 2026-08-29T11:00:07.638497+00:00, sha 159e32955ebf)
- Data as of 2026-08-30T08:39:29.467469+00:00.
