# lemonade-sdk/lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Repository: https://github.com/lemonade-sdk/lemonade
Canonical: https://ross.abutalabs.com/products/lemonade
Homepage: https://lemonade-server.ai/
Language: C++
License: Apache-2.0
License Family: permissive
Topics: amd, llama, llm, llm-inference, local-server, mistral, npu, onnxruntime, qwen, openai-api, mcp, mcp-server, gpu, radeon, ryzen, vulkan, ai, genai, rocm
Last push: 2026-08-26T20:32:46+00:00

## Health v2 (maintenance only)
Score: 82/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 99, release rhythm 87, longevity 33
- inputs: {"age_days": 475, "days_push": 7, "days_rel": 7, "gap_med": 6.0, "n_releases_24m": 69}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 5472, forks 478 (observed 2026-08-28T04:09:19.527501+00:00)

## What it is
Lemonade is a local AI server that runs optimized LLMs (plus image, speech, and embedding models) on your own GPU and NPU, exposing OpenAI-, Anthropic-, and Ollama-compatible APIs plus an MCP gateway. It ships as an installable service with GUI/CLI, an embeddable portable binary for bundling into apps, and includes AMD-specific optimizations for Ryzen AI, Radeon, and Strix Halo hardware.

## Use cases
- run local llm server with openai-compatible api on my pc
- serve llama or qwen models on amd radeon gpu or ryzen ai npu
- self-host a private chatgpt alternative with no telemetry
- embed a local ai runtime into my desktop application
- connect local llm to apps like open webui or coding agents
- run an mcp server backed by local models
- generate images and speech locally on windows
- manage and download local llm models with a gui

## When to choose
- you want cloud-API-style local inference that works out-of-box with hundreds of existing apps
- you're on AMD Ryzen AI, Radeon, or Strix Halo hardware and want vendor-optimized performance
- you need multi-modal local AI (chat, vision, image, speech, embeddings) from one service
- you want to bundle a private, customizable AI runtime into your own application
- you need OpenAI, Anthropic, and Ollama API compatibility plus MCP from a single server

## When to avoid
- you need large-scale multi-node or cloud inference rather than single-PC local serving
- you're locked into a specific inference stack like vLLM or TensorRT-LLM server features
- you only need a lightweight library call rather than a running HTTP service
- your target hardware lacks GPU/NPU acceleration and CPU-only speed is insufficient

## Facets
- artifact type: service
- maturity: active
- function: llm-inference, http-server, mcp, sdk, cli, gui, speech-recognition, tts, image-processing, rag
- domain: large-language-models, artificial-intelligence, self-hosted, developer-tools, machine-learning
- platform: windows, self-hosted, cpp, python
- tags: local-ai, openai-api-compatible, npu, amd-ryzen-ai, radeon, vulkan, rocm, llama-cpp, onnxruntime, ollama-compatible, anthropic-compatible, model-manager, privacy, embeddable, linux, macos, docker, gpu

## Member repositories
- lemonade-sdk/lemonade (main) score 82

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:19.527501+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:56:43.691007+00:00, confidence not recorded.
  - readme: https://github.com/lemonade-sdk/lemonade (fetched 2026-08-28T04:09:19.527501+00:00, sha e494f4b3ee33)
  - homepage: https://lemonade-server.ai/ (fetched 2026-08-29T08:51:49.889466+00:00, sha c1887400b3f8)
  - site_page: https://lemonade-server.ai/docs/embeddable (fetched 2026-08-29T08:51:49.898303+00:00, sha 9f6295bedfa5)
  - site_page: https://lemonade-server.ai/docs/api (fetched 2026-08-29T08:51:49.900248+00:00, sha cce9a1f0241f)
  - site_page: https://lemonade-server.ai/docs/dev/philosophy (fetched 2026-08-29T08:51:49.901943+00:00, sha 1462410a7ef1)
  - site_page: https://lemonade-server.ai/docs/dev/contribute (fetched 2026-08-29T08:51:49.904125+00:00, sha 8cced5839567)
  - site_page: https://lemonade-server.ai/docs/dev/release (fetched 2026-08-29T08:51:49.906076+00:00, sha 440b404c0fb2)
- Data as of 2026-08-30T08:39:29.467469+00:00.
