Ross ROSS = Recommend OSS · open-source software intelligence for agents

theroyallab/tabbyAPI

The official API server for Exllama. OAI compatible, lightweight, and fast. observed · 2026-08-28

github.com/theroyallab/tabbyAPI · Python · AGPL-3.0 (copyleft) observed · 2026-08-28

Health v2 · maintenance only

71/100

  • Activity 99
  • Release rhythm 35
  • Longevity 73

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1027
  • days_rel: n/a
  • days_push: 7
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1324 stars · 179 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

TabbyAPI is a FastAPI-based OpenAI-compatible API server for running large language models locally using the ExllamaV3 backend on NVIDIA GPUs. It serves as the official API backend for Exllama, supporting text generation and tool calling.

Use cases

  • serve local LLMs with an OpenAI-compatible API
  • run ExllamaV3 models on a GPU
  • self-host a lightweight LLM inference server
  • add tool calling to a local language model
  • generate text via HTTP API from local models
  • replace OpenAI endpoints with a self-hosted backend

When to choose

  • you want a fast, lightweight OpenAI-compatible server for Exllama models
  • you run local LLM inference on NVIDIA GPUs
  • you need tool calling support with a local backend
  • you prefer a hobbyist-friendly rolling-release server over heavyweight production stacks

When to avoid

  • you need a production-grade, high-concurrency inference server
  • you want to run GGUF models (use its sister project YALS instead)
  • you need ExLlamaV2 support on the main branch
  • you lack an NVIDIA GPU

Facets

application · maturity active

llm-inference http-server api-framework large-language-models artificial-intelligence self-hosted windows python self-hosted openai-compatible exllama fastapi text-generation local-llm tool-calling linux docker gpu

1 source

Member repositories

RepositoryRoleHealth v2
theroyallab/tabbyAPImain71

For agents

markdown · JSON · MCP: product_card(name="theroyallab/tabbyAPI")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem