# getumbrel/llama-gpt

A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!

Repository: https://github.com/getumbrel/llama-gpt
Canonical: https://ross.abutalabs.com/products/llama-gpt
Homepage: https://apps.umbrel.com/app/llama-gpt
Language: TypeScript
License: MIT
License Family: permissive
Topics: ai, chatgpt, gpt, gpt-4, gpt4all, llama, llama-2, llama-cpp, llama2, llamacpp, llm, localai, openai, self-hosted, code-llama, codellama
Last push: 2024-04-23T18:56:06+00:00

## Health v2 (maintenance only)
Score: 28/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 81
- inputs: {"age_days": 1138, "days_push": 862, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 10940, forks 706 (observed 2026-08-28T04:10:44.760507+00:00)

## What it is
LlamaGPT is a self-hosted, offline ChatGPT-like chatbot powered by Llama 2 and Code Llama models via llama.cpp, with all data staying on the user's device. It ships as a Docker-deployable app with an OpenAI-compatible API and one-click install on umbrelOS home servers.

## Use cases
- run a private chatgpt alternative on my own hardware
- self-host a local llm chatbot with no data leaving my device
- chat with llama 2 offline
- serve an openai-compatible api from a local model
- run code llama locally for coding help
- install a chatbot on my umbrel home server

## When to choose
- you want a fully private, offline ChatGPT-like experience on your own hardware
- you run an umbrelOS home server and want one-click local AI
- you need an OpenAI-compatible API backed by a local Llama 2 or Code Llama model
- you have an M1/M2 Mac or Nvidia GPU and enough RAM for 7B-70B quantized models

## When to avoid
- you need the latest models, custom model support, or frequent updates - the project appears in maintenance with limited recent activity
- you lack the RAM/GPU required for the quantized models
- you want multi-user or cloud-scale serving rather than a personal assistant
- you prefer actively developed alternatives like Ollama or Open WebUI

## Facets
- artifact type: application
- maturity: maintenance
- function: chatbot, llm-inference, self-hosted, chat-interface, api-framework
- domain: large-language-models, chatbots, artificial-intelligence, self-hosted, privacy
- platform: self-hosted
- tags: llama-2, code-llama, llama-cpp, offline-ai, local-llm, openai-compatible-api, umbrel, docker, macos, linux, web-server, kubernetes, gpu

## Member repositories
- getumbrel/llama-gpt (main) score 28

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:44.760507+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:17:19.456167+00:00, confidence not recorded.
  - readme: https://github.com/getumbrel/llama-gpt (fetched 2026-08-28T04:10:44.760507+00:00, sha ae60cffc357d)
  - homepage: https://apps.umbrel.com/app/llama-gpt (fetched 2026-08-29T08:16:22.428706+00:00, sha 5dd2b71a1452)
  - site_page: https://apps.umbrel.com/app/am-i-exposed (fetched 2026-08-29T08:16:22.437984+00:00, sha 7ce9901028cc)
- Data as of 2026-08-30T08:39:29.467469+00:00.
