Ross ROSS = Recommend OSS · open-source software intelligence for agents

SemiAnalysisAI/InferenceX resource

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3 observed · 2026-08-28

github.com/SemiAnalysisAI/InferenceX · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

71/100

  • Activity 99
  • Release rhythm 60
  • Longevity 29

Flags: prerelease_only

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 417
  • days_rel: 14
  • days_push: 7
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

1554 stars · 273 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

InferenceX is an open-source, vendor-neutral continuous benchmarking platform that measures LLM inference performance across serving frameworks (vLLM, SGLang, TensorRT-LLM) and AI accelerators (NVIDIA Blackwell/Hopper, AMD Instinct, and upcoming TPU/Trainium). All benchmark runs execute via public GitHub Actions workflows with open recipes, full logs, and weekly database snapshots for full reproducibility.

Use cases

  • compare LLM inference throughput across GB200 NVL72, B200, and MI355X hardware
  • benchmark vLLM vs SGLang vs TensorRT-LLM serving performance
  • measure long-context multi-turn agentic coding inference performance
  • track day-0 performance of new models like DeepSeek V4 or Kimi K3
  • plan GPU capacity and TCO for LLM serving infrastructure
  • reproduce auditable inference benchmark results from public CI runs
  • evaluate AMD ROCm vs NVIDIA CUDA inference stacks

When to choose

  • you need vendor-neutral, reproducible LLM inference benchmarks on frontier hardware
  • you are comparing serving frameworks or accelerators for production LLM deployment
  • you want continuously updated performance data on newly released models and chips
  • you need agentic long-context multi-turn workloads rather than static single-turn benchmarks

When to avoid

  • you need a lightweight benchmark you can run on a single consumer GPU or laptop
  • you want a general-purpose ML model quality or accuracy benchmark rather than serving performance
  • you lack access to datacenter-class accelerators like NVL72 racks or MI355X

Facets

dataset · maturity active

benchmarking llm-inference monitoring developer-tools machine-learning large-language-models gpu-computing performance developer-tools python cloud llm-benchmark inference-benchmark vllm sglang tensorrt-llm nvidia-blackwell amd-instinct rocm agentic-benchmark reproducible-benchmarks github-actions-automation hardware-comparison linux docker gpu

5 sources

Member repositories

RepositoryRoleHealth v2
SemiAnalysisAI/InferenceXmain71

For agents

markdown · JSON · MCP: product_card(name="SemiAnalysisAI/InferenceX")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem