# SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72，即将推出™ TPUv6e/v7/Trainium2/3

Repository: https://github.com/SemiAnalysisAI/InferenceX
Canonical: https://ross.abutalabs.com/products/inferencex
Homepage: https://inferencex.com/
Language: Python
License: Apache-2.0
License Family: permissive
Topics: ai, benchmark, llm, pytorch, sglang, vllm, amd, cuda, nvidia, rocm, gb200, deepseek, gb300, glm, kimi, minimax, mi355x
Last push: 2026-08-26T19:15:22+00:00

## Health v2 (maintenance only)
Score: 71/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 60, longevity 29
- inputs: {"age_days": 417, "days_push": 7, "days_rel": 14, "gap_med": null, "n_releases_24m": 1}
- flags: prerelease_only
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1554, forks 273 (observed 2026-08-28T04:05:02.332126+00:00)

## What it is
InferenceX is an open-source, vendor-neutral continuous benchmarking platform that measures LLM inference performance across serving frameworks (vLLM, SGLang, TensorRT-LLM) and AI accelerators (NVIDIA Blackwell/Hopper, AMD Instinct, and upcoming TPU/Trainium). All benchmark runs execute via public GitHub Actions workflows with open recipes, full logs, and weekly database snapshots for full reproducibility.

## Use cases
- compare LLM inference throughput across GB200 NVL72, B200, and MI355X hardware
- benchmark vLLM vs SGLang vs TensorRT-LLM serving performance
- measure long-context multi-turn agentic coding inference performance
- track day-0 performance of new models like DeepSeek V4 or Kimi K3
- plan GPU capacity and TCO for LLM serving infrastructure
- reproduce auditable inference benchmark results from public CI runs
- evaluate AMD ROCm vs NVIDIA CUDA inference stacks

## When to choose
- you need vendor-neutral, reproducible LLM inference benchmarks on frontier hardware
- you are comparing serving frameworks or accelerators for production LLM deployment
- you want continuously updated performance data on newly released models and chips
- you need agentic long-context multi-turn workloads rather than static single-turn benchmarks

## When to avoid
- you need a lightweight benchmark you can run on a single consumer GPU or laptop
- you want a general-purpose ML model quality or accuracy benchmark rather than serving performance
- you lack access to datacenter-class accelerators like NVL72 racks or MI355X

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, llm-inference, monitoring, developer-tools
- domain: machine-learning, large-language-models, gpu-computing, performance, developer-tools
- platform: python, cloud
- tags: llm-benchmark, inference-benchmark, vllm, sglang, tensorrt-llm, nvidia-blackwell, amd-instinct, rocm, agentic-benchmark, reproducible-benchmarks, github-actions-automation, hardware-comparison, linux, docker, gpu

## Member repositories
- SemiAnalysisAI/InferenceX (main) score 71

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:02.332126+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:30:04.476407+00:00, confidence not recorded.
  - readme: https://github.com/SemiAnalysisAI/InferenceX (fetched 2026-08-28T04:05:02.332126+00:00, sha e35f85f842f9)
  - homepage: https://inferencex.com/ (fetched 2026-08-29T11:30:17.695229+00:00, sha b05e4bb93712)
  - site_page: https://inferencex.semianalysis.com/about (fetched 2026-08-29T11:30:17.699484+00:00, sha ee2fa55b9c1f)
  - site_page: https://semianalysis.com/about (fetched 2026-08-29T11:30:17.701586+00:00, sha ec62cbbcc2d3)
  - site_page: https://inferencex.semianalysis.com/chips (fetched 2026-08-29T11:30:17.703327+00:00, sha f227503e6eb3)
- Data as of 2026-08-30T08:39:29.467469+00:00.
