Ross ROSS = Recommend OSS · open-source software intelligence for agents

google-ai-edge/LiteRT-LM

LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices. observed · 2026-08-28

github.com/google-ai-edge/LiteRT-LM · homepage · C++ · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

86/100

  • Activity 99
  • Release rhythm 98
  • Longevity 36
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 12.5
  • age_days: 506
  • days_rel: 15
  • days_push: 7
  • n_releases_24m: 15

Full methodology

Adoption not part of the score

6298 stars · 697 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

LiteRT-LM is Google's production-ready, high-performance open-source framework for running large language models on edge devices, built as an orchestration layer over LiteRT. It provides C, Python, Swift, JavaScript, Kotlin, and Flutter APIs with GPU/NPU acceleration across Android, iOS, Web, desktop, and IoT platforms.

Use cases

  • run llms on-device on android or ios
  • deploy gemma locally on a raspberry pi
  • run a local llm in the browser with webgpu
  • integrate on-device ai into a flutter app
  • run llm inference with gpu or npu acceleration
  • build an offline chatbot without a server
  • add function calling to an on-device ai agent

When to choose

  • you need production-grade on-device LLM inference across mobile, web, and desktop
  • you want hardware-accelerated (GPU/NPU) inference on edge devices
  • you need multimodal inputs (vision, audio) and tool use locally
  • you want to run Gemma, Llama, Phi-4, or Qwen models offline

When to avoid

  • you need server-scale inference with large models on datacenter GPUs
  • you only need cloud API access to LLMs without local deployment
  • you need fine-tuning or training rather than inference
  • you require a model format other than LiteRT/.litertlm

Facets

library · maturity active

llm-inference machine-learning sdk cli gpu-computing large-language-models artificial-intelligence mobile-development cross-platform embedded-systems cross-platform python cpp cli embedded windows edge-ai on-device-llm litert gemma hardware-acceleration multimodal function-calling webgpu raspberry-pi android ios web-server gpu macos linux

2 sources

Member repositories

RepositoryRoleHealth v2
google-ai-edge/LiteRT-LMmain86

For agents

markdown · JSON · MCP: product_card(name="google-ai-edge/LiteRT-LM")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem