XiongjieDai/GPU-Benchmarks-on-LLM-Inference resource
Multiple NVIDIA GPUs or Apple Silicon for Large Language Model Inference? observed · 2026-08-28
Health v2 · maintenance only
28/100
- Activity 0
- Release rhythm 35
- Longevity 81
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1142
- days_rel: n/a
- days_push: 843
- n_releases_24m: 0
Adoption not part of the score
1935 stars · 76 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A curated benchmark dataset comparing LLM inference speeds (tokens/s) across many NVIDIA GPUs and Apple Silicon chips using llama.cpp on LLaMA-family models. Results are presented as tables and Jupyter notebooks covering quantized and full-precision model sizes from 8B to 70B parameters.
Use cases
- which gpu should i buy for running llama locally
- compare inference speed of rtx 4090 vs a100 for llm
- can a macbook m1 max run llama 70b
- how fast is llama 3 8b on a 3090
- best gpu for local llm inference on a budget
- how many gpus do i need to run a 70b model
- apple silicon vs nvidia for llm inference
When to choose
- choosing hardware for local LLM inference with llama.cpp
- comparing consumer and datacenter GPUs for token generation speed
- checking whether a given GPU has enough VRAM for a model size and quantization
When to avoid
- benchmarking training performance rather than inference
- comparing inference frameworks other than llama.cpp
- needing up-to-date results for the newest GPU generations
Facets
dataset · maturity maintenance
benchmarking llm-inference gpu-computing large-language-models gpu-computing hardware performance llama-cpp nvidia-gpus apple-silicon tokens-per-second hardware-comparison runpod gpu macos linux
1 source
- readme: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference · fetched 2026-08-28 · 34fce38c3d75
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| XiongjieDai/GPU-Benchmarks-on-LLM-Inference | main | 28 |
For agents
markdown · JSON · MCP: product_card(name="XiongjieDai/GPU-Benchmarks-on-LLM-Inference")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem