Ross ROSS = Recommend OSS · open-source software intelligence for agents

cactus-compute/cactus

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots. observed · 2026-08-28

github.com/cactus-compute/cactus · homepage · C++ · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

86/100

  • Activity 99
  • Release rhythm 98
  • Longevity 35

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 7
  • age_days: 497
  • days_rel: 16
  • days_push: 7
  • n_releases_24m: 20

Full methodology

Adoption not part of the score

5934 stars · 495 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Cactus is a hybrid edge-cloud AI inference engine for mobile devices, wearables, smart home devices, and robots, built in C++ with custom quantization, CPU/GPU kernels, and a zero-copy computation graph. It exposes OpenAI-compatible APIs for text, speech, and vision, with automatic cloud handoff when on-device models are uncertain.

Use cases

  • run LLM inference on-device on Android and iOS
  • deploy AI models on wearables and microcontrollers
  • on-device speech transcription with Whisper
  • build RAG apps that run locally on phones
  • tool calling and structured extraction on tiny edge devices
  • quantize transformer models for mobile inference

When to choose

  • you need low-latency, battery-efficient LLM inference on smartphones or embedded hardware
  • you want on-device AI with automatic cloud fallback for uncertain predictions
  • you need a single engine covering text, speech, and vision at the edge

When to avoid

  • you need a fully permissively licensed dependency (license is non-standard)
  • you only target server or desktop GPUs with no edge constraints
  • you need a mature ecosystem with broad model support like llama.cpp or ONNX Runtime

Facets

framework · maturity active

llm-inference rag speech-recognition machine-learning sdk machine-learning large-language-models artificial-intelligence mobile-development embedded-systems speech-processing cross-platform cpp embedded on-device-ai edge-ai quantization wearables llamacpp whisper cloud-fallback inference-engine arm android ios macos mobile

3 sources

Member repositories

RepositoryRoleHealth v2
cactus-compute/cactusmain86

For agents

markdown · JSON · MCP: product_card(name="cactus-compute/cactus")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem