Ross ROSS = Recommend OSS · open-source software intelligence for agents

walter-grace/mac-code

mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month. observed · 2026-08-28

github.com/walter-grace/mac-code · Python observed · 2026-08-28

Health v2 · maintenance only

49/100

  • Activity 76
  • Release rhythm 35
  • Longevity 11

Flags: no_releases young no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 163
  • days_rel: n/a
  • days_push: 146
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1026 stars · 109 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A free, local coding agent for Apple Silicon Macs that runs large quantized LLMs (e.g., Qwen 35B) via llama.cpp, using techniques like flash-paging and SSD weight streaming to run models that exceed RAM. It provides a Claude Code-like terminal agent experience with no subscription cost.

Use cases

  • run a local coding agent on my mac for free
  • run a 35B model on a 16GB mac
  • replace claude code with a local model
  • run llms that don't fit in ram on apple silicon
  • stream model weights from ssd to run big models
  • self-host an ai coding assistant offline
  • get fast local llm inference on mac mini

When to choose

  • you have an Apple Silicon Mac and want a free, private, offline coding agent
  • you want to run models larger than your RAM without heavy quantization
  • you're comfortable with terminal tools, llama-server, and downloading GGUF models

When to avoid

  • you need top-tier coding quality comparable to frontier cloud models
  • you need fast inference for dense models far exceeding RAM (speeds can drop below 1 tok/s)
  • you need a permissive license or polished GUI - the repo has no license and is terminal-based

Facets

application · maturity active

llm-inference agent-framework cli chat-interface large-language-models artificial-intelligence developer-tools python cli local-llm apple-silicon llama-cpp gguf quantization flash-streaming coding-agent free off-device-ai moe ai-agents command-line macos desktop

1 source

Member repositories

RepositoryRoleHealth v2
walter-grace/mac-codemain49

For agents

markdown · JSON · MCP: product_card(name="walter-grace/mac-code")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem