Ross ROSS = Recommend OSS · open-source software intelligence for agents

XiongjieDai/GPU-Benchmarks-on-LLM-Inference resource

Multiple NVIDIA GPUs or Apple Silicon for Large Language Model Inference? observed · 2026-08-28

github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference · Jupyter Notebook observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 0
  • Release rhythm 35
  • Longevity 81

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1142
  • days_rel: n/a
  • days_push: 843
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1935 stars · 76 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A curated benchmark dataset comparing LLM inference speeds (tokens/s) across many NVIDIA GPUs and Apple Silicon chips using llama.cpp on LLaMA-family models. Results are presented as tables and Jupyter notebooks covering quantized and full-precision model sizes from 8B to 70B parameters.

Use cases

  • which gpu should i buy for running llama locally
  • compare inference speed of rtx 4090 vs a100 for llm
  • can a macbook m1 max run llama 70b
  • how fast is llama 3 8b on a 3090
  • best gpu for local llm inference on a budget
  • how many gpus do i need to run a 70b model
  • apple silicon vs nvidia for llm inference

When to choose

  • choosing hardware for local LLM inference with llama.cpp
  • comparing consumer and datacenter GPUs for token generation speed
  • checking whether a given GPU has enough VRAM for a model size and quantization

When to avoid

  • benchmarking training performance rather than inference
  • comparing inference frameworks other than llama.cpp
  • needing up-to-date results for the newest GPU generations

Facets

dataset · maturity maintenance

benchmarking llm-inference gpu-computing large-language-models gpu-computing hardware performance llama-cpp nvidia-gpus apple-silicon tokens-per-second hardware-comparison runpod gpu macos linux

1 source

Member repositories

RepositoryRoleHealth v2
XiongjieDai/GPU-Benchmarks-on-LLM-Inferencemain28

For agents

markdown · JSON · MCP: product_card(name="XiongjieDai/GPU-Benchmarks-on-LLM-Inference")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem