SemiAnalysisAI/InferenceX resource
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3 observed · 2026-08-28
Health v2 · maintenance only
71/100
- Activity 99
- Release rhythm 60
- Longevity 29
Flags: prerelease_only
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 417
- days_rel: 14
- days_push: 7
- n_releases_24m: 1
Adoption not part of the score
1554 stars · 273 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
InferenceX is an open-source, vendor-neutral continuous benchmarking platform that measures LLM inference performance across serving frameworks (vLLM, SGLang, TensorRT-LLM) and AI accelerators (NVIDIA Blackwell/Hopper, AMD Instinct, and upcoming TPU/Trainium). All benchmark runs execute via public GitHub Actions workflows with open recipes, full logs, and weekly database snapshots for full reproducibility.
Use cases
- compare LLM inference throughput across GB200 NVL72, B200, and MI355X hardware
- benchmark vLLM vs SGLang vs TensorRT-LLM serving performance
- measure long-context multi-turn agentic coding inference performance
- track day-0 performance of new models like DeepSeek V4 or Kimi K3
- plan GPU capacity and TCO for LLM serving infrastructure
- reproduce auditable inference benchmark results from public CI runs
- evaluate AMD ROCm vs NVIDIA CUDA inference stacks
When to choose
- you need vendor-neutral, reproducible LLM inference benchmarks on frontier hardware
- you are comparing serving frameworks or accelerators for production LLM deployment
- you want continuously updated performance data on newly released models and chips
- you need agentic long-context multi-turn workloads rather than static single-turn benchmarks
When to avoid
- you need a lightweight benchmark you can run on a single consumer GPU or laptop
- you want a general-purpose ML model quality or accuracy benchmark rather than serving performance
- you lack access to datacenter-class accelerators like NVL72 racks or MI355X
Facets
dataset · maturity active
benchmarking llm-inference monitoring developer-tools machine-learning large-language-models gpu-computing performance developer-tools python cloud llm-benchmark inference-benchmark vllm sglang tensorrt-llm nvidia-blackwell amd-instinct rocm agentic-benchmark reproducible-benchmarks github-actions-automation hardware-comparison linux docker gpu
5 sources
- readme: https://github.com/SemiAnalysisAI/InferenceX · fetched 2026-08-28 · e35f85f842f9
- homepage: https://inferencex.com/ · fetched 2026-08-29 · b05e4bb93712
- site_page: https://inferencex.semianalysis.com/about · fetched 2026-08-29 · ee2fa55b9c1f
- site_page: https://semianalysis.com/about · fetched 2026-08-29 · ec62cbbcc2d3
- site_page: https://inferencex.semianalysis.com/chips · fetched 2026-08-29 · f227503e6eb3
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| SemiAnalysisAI/InferenceX | main | 71 |
For agents
markdown · JSON · MCP: product_card(name="SemiAnalysisAI/InferenceX")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem