function: benchmarking
824 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| awslabs/dgl-ke DGL-KE is a high-performance Python package built on Deep Graph Library (DGL) for training, evaluating, and inferring knowledge graph embed… | 64 | 1331 | active |
| guonaihong/gout gout is a Go HTTP client library designed as a 'Swiss Army Knife' for making HTTP requests with a fluent API. It supports multiple body enc… | 46 | 1329 | active |
| lucko/spark spark is a performance profiler for Minecraft clients, servers, and proxies, combining a lightweight CPU profiler, memory inspection tools … | 77 | 1323 | active |
| giltene/wrk2 wrk2 is a constant-throughput HTTP benchmarking tool based on wrk, modified to produce accurate latency recordings using HdrHistograms. It … | 32 | 4626 | maintenance |
| openai/simple-evals A lightweight Python library from OpenAI for evaluating language models against benchmarks like MMLU, GPQA, MATH, HumanEval, SimpleQA, Heal… | 60 | 4612 | maintenance |
| pythonprofilers/memory_profiler A pure Python module for monitoring memory consumption of a process, including line-by-line analysis of memory usage in Python programs. It… | 32 | 4575 | maintenance |
| treosh/lighthouse-ci-action A GitHub Action that runs Lighthouse audits on URLs and enforces Lighthouse CI assertions or performance budgets within CI workflows. It su… | 67 | 1289 | active |
| greyhaven-ai/autocontext Autocontext is a recursive self-improving harness that runs AI agents against evaluations, retains useful lessons, and produces traces, rep… | 60 | 1286 | active |
| locuslab/TCN PyTorch implementation of Temporal Convolutional Networks (TCN) with benchmarks from the paper 'An Empirical Evaluation of Generic Convolut… | 32 | 4549 | maintenance |
| microsoft/waza Waza is a Go CLI from Microsoft for evaluating AI agent skills through structured benchmarks. Users define test cases and validation rules … | 81 | 1280 | active |
| CXWorld/CapFrameX CapFrameX is a Windows desktop application for capturing and analyzing frametimes and FPS in PC games, built on Intel's PresentMon with ove… | 95 | 1279 | stable |
| siddharthkp/bundlesize bundlesize is an npm CLI tool that checks whether your JavaScript bundle files stay under configured size limits, with gzip or brotli compr… | 66 | 4474 | maintenance |
| smart-test-ti/SoloX SoloX is a real-time performance data collection tool for Android and iOS applications, gathering metrics like CPU, memory, FPS, network, a… | 23 | 1262 | active |
| mjpost/sacrebleu SacreBLEU is a Python library and CLI tool for computing shareable, comparable, and reproducible BLEU, chrF, and TER scores for machine tra… | 76 | 1258 | active |
| JonathonLuiten/TrackEval TrackEval is a Python library for evaluating multi-object tracking (MOT) algorithms, implementing metrics such as HOTA, CLEARMOT, IDF1, VAC… | 32 | 1255 | stable |
| aras-p/ClangBuildAnalyzer A command-line tool that aggregates Clang -ftime-trace JSON reports across a whole C/C++ build and summarizes what took the most compile ti… | 28 | 1253 | stable |
| VainF/pytorch-msssim A PyTorch library providing fast, differentiable SSIM and MS-SSIM image quality metrics using separable Gaussian filtering for speed. It ca… | 23 | 1253 | stable |
| benchmark-action/github-action-benchmark A GitHub Action that collects benchmark results from many benchmarking tools and tracks them across CI runs. It stores results on a GitHub … | 87 | 1250 | active |
| automl/SMAC3 SMAC3 is a Python library for Bayesian Optimization used to tune hyperparameters of machine learning algorithms and configure arbitrary alg… | 81 | 1244 | active |
| eembc/coremark CoreMark is EEMBC's industry-standard benchmark for measuring CPU and embedded microcontroller performance, testing core processor features… | 67 | 1242 | stable |
| PauperZ/SSRSpeedN SSRSpeedN is a Python-based benchmarking platform for proxy servers, forked from SSRSpeed, that supports batch speed testing of nodes. It m… | 30 | 1241 | active |
| myzhan/boomer Boomer is a Go library that acts as a high-performance load generator for Locust, spawning thousands of goroutines to run test code concurr… | 73 | 1237 | active |
| theintern/intern Intern is a complete JavaScript testing stack supporting unit and functional testing in Node.js and browsers, with built-in code coverage, … | 23 | 4344 | maintenance |
| google/orbit Orbit is a standalone native application profiler for Windows and Linux that combines callstack sampling with dynamic instrumentation to id… | 10 | 4316 | maintenance |
| ai-dynamo/nixl NVIDIA Inference Xfer Library (NIXL) is a C++/Python library that accelerates point-to-point communication in AI inference frameworks like … | 88 | 1227 | active |
| hugoduncan/criterium Criterium is a benchmarking library for Clojure that measures expression computation time while addressing JVM benchmarking pitfalls. It ap… | 79 | 1224 | active |
| google/fuzzbench FuzzBench is a free Google-run service and framework that rigorously evaluates fuzzers on real-world benchmarks at large scale. It provides… | 61 | 1205 | active |
| AvdLee/Xcode-Build-Optimization-Agent-Skill A collection of open-source Agent Skills that benchmark and optimize Xcode build performance, covering clean and incremental builds, compil… | 73 | 1204 | active |
| NVIDIA/kvpress kvpress is a Python library from NVIDIA that implements multiple KV cache compression methods and benchmarks for long-context LLM inference… | 86 | 1201 | active |
| toshas/torch-fidelity A PyTorch library providing accurate and efficient implementations of generative model evaluation metrics such as FID, Inception Score, KID… | 70 | 1197 | active |
| kornelski/dssim DSSIM is a Rust CLI tool and library that measures perceptual (dis)similarity between PNG/JPEG images using a multi-scale variant of the SS… | 68 | 1194 | active |
| aspnet/Benchmarks A collection of benchmark applications and Docker images for measuring ASP.NET Core performance, including TechEmpower Web Framework Benchm… | 77 | 1192 | active |
| xoofx/ultra Ultra is an advanced sampling profiler for .NET applications on Windows (ETW-based) and macOS Apple Silicon (EventPipe-based). It produces … | 85 | 1186 | active |
| CNugteren/CLBlast CLBlast is a lightweight, tunable OpenCL BLAS library written in C++11 that implements basic linear algebra subprograms for vectors and mat… | 70 | 1186 | stable |
| ethereum/execution-specs An executable Python reference implementation of Ethereum's Execution Layer (EELS), including the EVM and consensus-critical behavior, with… | 95 | 1185 | active |
| pydata/bottleneck Bottleneck is a collection of fast NumPy array functions implemented as C extensions, covering NaN-aware reductions like nanmean and fast m… | 66 | 1181 | active |
| data61/MP-SPDZ MP-SPDZ is a versatile C++ framework for secure multi-party computation (MPC) supporting many protocols across various security models, inc… | 85 | 1178 | active |
| Spearfoot/disk-burnin-and-testing A POSIX-compliant shell script that automates burn-in and stress testing of new or re-purposed disk drives. It runs SMART short and extende… | 66 | 1173 | active |
| evan-kolberg/prediction-market-backtesting A Python/Rust extension for Nautilus Trader providing backtesting tools for prediction market trading strategies, with a focus on Polymarke… | 52 | 1168 | active |
| zilliztech/VectorDBBench VectorDBBench (VDBBench) is an open-source Python benchmark tool for comparing the performance and cost-effectiveness of vector databases a… | 91 | 1166 | active |
| GaParmar/clean-fid Clean-FID is a PyTorch library for computing the Frechet Inception Distance (FID) with correct image resizing and quantization steps, fixin… | 48 | 1166 | stable |
| JackHopkins/factorio-learning-environment An open-source framework for developing and evaluating LLM agents in the game of Factorio, providing an open-ended, non-saturating benchmar… | 82 | 1156 | active |
| NVIDIA-NeMo/Gym NeMo Gym is a Python library from NVIDIA for evaluating and improving LLM models and agents using environments. It provides infrastructure … | 80 | 1156 | active |
| easystats/performance An R package from the easystats ecosystem that computes indices of model quality and goodness of fit, such as R-squared, RMSE, ICC, AIC, an… | 92 | 1151 | active |
| PKU-Alignment/omnisafe OmniSafe is a PyTorch-based infrastructural framework for safe reinforcement learning research, providing a unified modular toolkit and com… | 27 | 1149 | active |
| nvdv/vprof vprof is a Python package providing rich and interactive visualizations for Python program characteristics such as running time and memory … | 23 | 3977 | maintenance |
| chengtan9907/OpenSTL OpenSTL is a comprehensive benchmark and modular framework for spatio-temporal predictive learning, covering video prediction methods acros… | 54 | 1137 | active |
| SIPp/sipp SIPp is a free, open-source SIP (Session Initiation Protocol) protocol testing tool written in C++. It can generate SIP traffic to test, be… | 82 | 1134 | active |
| needle-tools/compilation-visualizer A Unity Editor plugin that visualizes assembly compilation on a timeline, hooking into editor events to show how long each assembly takes t… | 65 | 1124 | active |
| jhawthorn/vernier Vernier is a next-generation sampling profiler for Ruby 3.2.1+ that tracks time, allocations, multiple threads, GVL activity, GC pauses, an… | 93 | 1117 | active |
| tindy2013/stairspeedtest-reborn A C++ batch speed testing tool for proxy nodes supporting Shadowsocks, ShadowsocksR, V2Ray, and Trojan. It tests subscription links, logs r… | 23 | 3839 | maintenance |
| kimwalisch/primesieve primesieve is a fast command-line program and C/C++ library for generating prime numbers and prime k-tuplets up to 2^64 using a segmented s… | 93 | 1111 | stable |
| amazon-science/RAGChecker RAGChecker is an automatic evaluation framework for diagnosing Retrieval-Augmented Generation (RAG) systems. It provides holistic and diagn… | 25 | 1110 | active |
| ccfos/huatuo HUATUO is an eBPF-based Linux kernel observability agent that provides kernel-wide metrics, event-driven context capture, autotracing, and … | 76 | 1107 | active |
| prometheus-eval/prometheus-eval Prometheus-Eval is a Python library for evaluating LLM generation outputs using the Prometheus family of open evaluator models and GPT-4 as… | 23 | 1107 | active |
| ErikEJ/SqlQueryStress SqlQueryStress is a Windows GUI tool (with a cross-platform CLI companion, sqlstresscmd) for stress-testing SQL Server queries by running t… | 94 | 1105 | active |
| kristiandupont/react-geiger React Geiger is a React library that audiolizes performance issues by playing click sounds when component re-renders exceed a configurable … | 70 | 1104 | active |
| LAMDA-CL/PyCIL PyCIL is a PyTorch-based Python toolbox for class-incremental learning, implementing the largest collection of CIL methods for reproducible… | 52 | 1098 | active |
| nolanlawson/marky A tiny (491 bytes) JavaScript timer library built on the User Timing API's performance.mark() and performance.measure(), with fallbacks to … | 63 | 1094 | stable |
| Emsu/prophet Prophet is a Python microframework for financial markets analysis that lets programmers model trading strategies, manage portfolios, and ru… | 57 | 1093 | active |
| SCLBD/DeepfakeBench DeepfakeBench is a comprehensive benchmark framework for deepfake detection, providing a unified platform for data management, implementati… | 36 | 1093 | active |
| SlugLab/CXLMemSim CXLMemSim is a simulation and emulation framework for CXL 3.0 memory systems, combining a latency/bandwidth/topology/coherency simulator wi… | 70 | 1092 | active |
| openjdk/jol Java Object Layout (JOL) is a toolbox for analyzing object layout, memory footprint, and references inside JVMs. It uses Unsafe, JVMTI, and… | 61 | 1088 | active |
| GAIR-NLP/LIMO LIMO is a research project and training framework demonstrating that large language models can achieve strong mathematical reasoning with o… | 36 | 1083 | active |
| mlr mlr/mlr3 is a machine learning framework for R providing a unified, object-oriented interface to many learning algorithms. It supports resa… | 95 | 1080 | active |
| dotnet/crank Crank is Microsoft's benchmarking infrastructure used by the .NET team to run performance benchmarks, including TechEmpower web framework s… | 77 | 1080 | active |
| inikep/lzbench lzbench is an in-memory benchmarking tool that integrates many open-source compression libraries into a single executable to compare their … | 86 | 1078 | active |
| Stonesjtu/pytorch_memlab A Python library providing line-level CUDA memory profiling and tensor inspection tools for PyTorch. It helps debug out-of-memory errors by… | 94 | 1077 | active |
| autonomousvision/navsim NAVSIM is a data-driven pseudo-simulation framework and benchmark for autonomous vehicle planning, evaluating driving agents non-reactively… | 50 | 1076 | active |
| chronoxor/CppTrader CppTrader is a C++ library of high-performance components for building trading platforms, including an ultra-fast matching engine, order bo… | 74 | 1067 | active |
| tsliwowicz/go-wrk go-wrk is an HTTP benchmarking and load-testing CLI tool written in Go, inspired by the wrk tool. It uses goroutines for concurrent async I… | 73 | 1062 | stable |
| harmonycloud/kindling Kindling is an eBPF-based cloud native monitoring tool that runs as a DaemonSet in Kubernetes to capture syscalls and tracepoints from kern… | 65 | 1062 | active |
| bytedance/SandboxFusion A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via … | 63 | 1060 | active |
| bigcode-project/bigcode-evaluation-harness A framework for evaluating autoregressive code generation language models on benchmarks like HumanEval, MBPP, MultiPL-E, and HumanEvalPack.… | 38 | 1058 | active |
| oblador/react-native-performance A toolchain implementing the Performance API for React Native, enabling measurement of app performance in development, CI pipelines, and pr… | 56 | 1055 | active |
| PreferredAI/cornac Cornac is a Python framework for building and comparing multimodal recommender systems, with convenient support for auxiliary data such as … | 96 | 1053 | active |
| segmentio/encoding A Go library providing high-performance encoders and decoders for data formats, most notably a drop-in replacement for the standard library… | 86 | 1053 | active |
| nmslib/nmslib NMSLIB is a cross-platform C++ similarity search library with Python bindings for efficient k-nearest-neighbor search in generic and non-me… | 57 | 3588 | maintenance |
| NeuroTechX/moabb MOABB (Mother of All BCI Benchmarks) is a Python library for reproducible benchmarking of machine-learning algorithms on EEG-based brain-co… | 96 | 1048 | active |
| redis/memtier_benchmark memtier_benchmark is a command-line load generation and benchmarking tool for NoSQL key-value databases, developed by Redis. It supports bo… | 97 | 1047 | active |
| centerforaisafety/HarmBench HarmBench is a standardized, open-source evaluation framework for automated red teaming of large language models, comparing attack methods … | 26 | 1033 | active |
| romkatv/zsh-bench zsh-bench is a benchmarking tool that measures user-visible latency of interactive zsh, such as input lag and command lag. It also includes… | 68 | 1028 | active |
| python/pyperformance pyperformance is the official Python Performance Benchmark Suite, providing an authoritative set of real-world benchmarks for Python implem… | 84 | 1027 | active |
| numpy/x86-simd-sort A C++ template library providing high-performance SIMD-accelerated sorting routines (qsort, qselect, partial sort, argsort, key-value sort)… | 63 | 1026 | active |
| airspeed-velocity/asv Airspeed Velocity (asv) is a Python benchmarking tool that tracks a project's performance over its git history. It generates an interactive… | 84 | 1014 | active |
| authorjapps/zerocode Zerocode TDD is an open-source, no-code automated testing framework for REST/SOAP APIs, Kafka data streams, databases, ETL pipelines, and l… | 96 | 1012 | active |
| linux-rdma/perftest A collection of C-based performance micro-benchmark tools for InfiniBand and RoCE networks, built on the ibverbs API. It measures bandwidth… | 85 | 1012 | active |
| Cloudslab/cloudsim CloudSim is a Java-based framework for modeling and simulating cloud computing infrastructures and services, including data centers, virtua… | 53 | 1012 | active |
| djkoloski/rust_serialization_benchmark A benchmark suite comparing Rust serialization frameworks on serialize, deserialize, borrow, size, and compression metrics, including zero-… | 77 | 1009 | active |
| davidtvs/pytorch-lr-finder A PyTorch library implementing the learning rate range test from Leslie Smith's cyclical learning rates paper, including the fastai-tweaked… | 35 | 1008 | active |
| aragozin/jvm-tools Swiss Java Knife (SJK) is a command-line tool for JVM diagnostics, troubleshooting, and profiling built on standard JVM diagnostic interfac… | 32 | 3340 | maintenance |
| google/model_search Model Search is a Google AutoML framework that implements neural architecture search algorithms at scale to find optimal DNN architectures … | 10 | 3238 | maintenance |
| bombomby/optick Optick is a lightweight C++ profiler designed for games, offering instrumentation, context-switch tracking, sampling, and GPU counters. It … | 23 | 3161 | maintenance |
| GoogleChromeLabs/psi A Node.js library and CLI wrapper around Google's PageSpeed Insights v5 API that runs mobile and desktop performance tests on deployed site… | 10 | 3101 | maintenance |
| reloadware/reloadium Reloadium is a Python library and PyCharm plugin providing advanced hot reloading, profiling, and AI-assisted error fixing. It lets develop… | 32 | 2988 | maintenance |
| isaac-sim/IsaacGymEnvs A collection of example reinforcement learning environments for NVIDIA Isaac Gym, a GPU-accelerated physics simulator. It provides a Gym-st… | 10 | 2952 | maintenance |
| rail-berkeley/rlkit RLkit is a PyTorch-based reinforcement learning framework and algorithm collection from UC Berkeley RAIL. It provides reference implementat… | 32 | 2932 | maintenance |
| aappleby/smhasher SMHasher is a C++ test suite that evaluates the distribution, collision, and performance properties of non-cryptographic hash functions. It… | 72 | 2889 | maintenance |
| stanford-crfm/helm HELM is a Python framework from Stanford's CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal mo… | 87 | 2887 | maintenance |
| microsoftarchive/promptbench PromptBench is a unified Python library from Microsoft for evaluating and understanding large language models across many datasets, models,… | 10 | 2819 | maintenance |