resource: benchmarking
233 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| tastejs/todomvc TodoMVC is a collection of the same Todo application implemented in dozens of JavaScript frameworks and libraries, letting developers compa… | 78 | 28949 | active |
| sebastianruder/NLP-progress A community-maintained repository and website tracking state-of-the-art results across the most common NLP tasks, with datasets, papers, an… | 23 | 22952 | active |
| xlite-dev/LeetCUDA A collection of modern CUDA learning notes with PyTorch integration for beginners, featuring 200+ CUDA kernels, 100+ LLM/CUDA blogs, and hi… | 93 | 11836 | active |
| davidsonfellipe/awesome-wpo A curated awesome-list of Web Performance Optimization resources including tools, articles, books, documentation, and talks. It covers cate… | 75 | 9053 | active |
| atinfo/awesome-test-automation A curated awesome-list of test automation frameworks, tools, libraries, and software organized by programming language (Python, Java, Ruby,… | 57 | 7146 | active |
| jeinlee1991/chinese-llm-benchmark ReLE (formerly CLiB) is a continuously updated Chinese-language LLM capability benchmark and leaderboard covering ~400 commercial and open-… | 89 | 6401 | active |
| microsoft/promptbase A Microsoft-maintained collection of prompt engineering resources, best practices, and example Python scripts for eliciting top performance… | 26 | 5771 | active |
| fastruby/fast-ruby A curated collection of common Ruby idioms with benchmark comparisons showing which constructs run faster, inspired by Erik Michaels-Ober's… | 72 | 5734 | active |
| OpenBMB/ToolBench ToolBench is an open-source platform for training, serving, and evaluating large language models on tool use, built around a large-scale in… | 39 | 5732 | active |
| SWE-bench/SWE-bench SWE-bench is a benchmark and evaluation harness that tests whether large language models can resolve real-world GitHub issues by generating… | 72 | 5719 | active |
| KellerJordan/modded-nanogpt A collaborative speedrun project that trains a GPT-2 (124M) scale language model to 3.28 validation loss on FineWeb in under 75 seconds on … | 67 | 5707 | active |
| sirupsen/napkin-math A collection of memorized latency/throughput numbers, benchmark code, and techniques for estimating system performance from first principle… | 65 | 5671 | active |
| openai/parameter-golf An OpenAI-hosted challenge to train the best-performing language model that fits in a 16MB artifact and trains in under 10 minutes on 8xH10… | 51 | 5174 | active |
| fchollet/ARC-AGI The Abstraction and Reasoning Corpus (ARC-AGI-1), a benchmark dataset of 800 JSON grid-based tasks designed to test human-like fluid genera… | 29 | 4813 | stable |
| nikolaydubina/go-recipes A curated awesome-list of handy and lesser-known command-line tools for Go projects, covering testing, coverage, benchmarking, profiling, c… | 75 | 4500 | active |
| datacharmer/test_db A sample MySQL employees database with an integrated test suite, migrated from Launchpad, also compatible with MariaDB, Percona Server, and… | 57 | 4434 | stable |
| CLUEbenchmark/CLUE CLUE is the Chinese Language Understanding Evaluation Benchmark, providing representative datasets, baseline and pre-trained models, corpor… | 62 | 4279 | active |
| anthropics/original_performance_takehome Anthropic's open-sourced performance engineering take-home challenge, where you optimize a simulated machine's code from a slow baseline (2… | 44 | 4122 | stable |
| dendibakh/perf-ninja An online course in the form of C++ lab assignments for learning low-level performance analysis and tuning, covering issues like CPU cache … | 77 | 3825 | active |
| denji/awesome-http-benchmark A curated awesome-list of HTTP(S) benchmarking, load testing, and REST API testing/debugging tools. It catalogs tools like ab, wrk, autocan… | 76 | 3765 | active |
| THUDM/AgentBench AgentBench is a comprehensive benchmark for evaluating large language models as agents across multi-turn interactive tasks like database qu… | 58 | 3695 | active |
| waymo-research/waymo-open-dataset The Waymo Open Dataset is a large-scale public collection of autonomous driving datasets (Perception, Motion, and End-to-End Driving) accom… | 50 | 3398 | active |
| vectara/hallucination-leaderboard A public leaderboard maintained by Vectara that ranks LLMs by how often they hallucinate when summarizing short documents, computed with Ve… | 64 | 3306 | active |
| CLUEbenchmark/SuperCLUE SuperCLUE is a comprehensive benchmark for evaluating Chinese-language large language models across language understanding and generation, … | 59 | 3299 | active |
| adamsitnik/awesome-dot-net-performance A curated list of .NET performance resources including books, courses, trainings, conference talks, blogs, and notable open source contribu… | 68 | 3279 | active |
| xlang-ai/OSWorld OSWorld is a benchmark and scalable real computer environment for evaluating multimodal agents on open-ended computer tasks across operatin… | 62 | 3108 | active |
| PlummersSoftwareLLC/Primes A community-maintained collection of prime number sieve implementations in 100+ programming languages, used to benchmark and compare langua… | 73 | 3017 | active |
| Kobzol/hardware-effects A collection of small C++ proof-of-concept programs demonstrating hardware effects like cache conflicts, branch misprediction, false sharin… | 32 | 2999 | stable |
| kostya/benchmarks A collection of cross-language performance benchmarks comparing programming languages on tasks like Brainfuck interpretation, Base64, JSON,… | 74 | 2922 | active |
| Tony-Tan/CUDA_Freshman A collection of CUDA example programs accompanying a Chinese-language blog tutorial series on GPU programming, partly based on the book 'Pr… | 32 | 2792 | stable |
| dgryski/go-perfbook An open-source book by Damian Gryski collecting best practices for writing high-performance Go code, covering general optimization methodol… | 32 | 10896 | maintenance |
| asatarin/testing-distributed-systems A curated list of resources on testing distributed systems, including research papers, tools, and case studies. It covers approaches like J… | 75 | 2635 | active |
| harbor-framework/terminal-bench-1 Terminal-Bench is a benchmark and execution harness for evaluating LLM agents on complex, real-world tasks in a terminal sandbox. It combin… | 62 | 2554 | active |
| DestinyLinker/MingLi-Bench A Python benchmark dataset and CLI for evaluating large language models on Chinese traditional fortune telling, covering Bazi (八字) and Ziwe… | 50 | 2338 | active |
| PacktPublishing/The-Kaggle-Book The official code repository for 'The Kaggle Book' by Packt Publishing, containing Jupyter Notebook examples from two Kaggle Grandmasters. … | 64 | 2333 | active |
| beir-cellar/beir BEIR is a heterogeneous benchmark for information retrieval, aggregating 15+ diverse IR datasets with a common evaluation framework for NLP… | 47 | 2275 | active |
| Lifelong-Robot-Learning/LIBERO LIBERO is a benchmark for studying knowledge transfer in multitask and lifelong robot learning, built on a procedural generation pipeline f… | 35 | 2238 | active |
| AmberLJC/LLMSys-PaperList A curated list of academic papers, tutorials, and projects on Large Language Model systems, covering training, serving, agentic systems, an… | 71 | 2233 | active |
| alibaba/clusterdata A collection of production cluster trace datasets from Alibaba covering batch workloads, GPU/ML workloads, microservices, and microarchitec… | 70 | 2154 | active |
| gunnarmorling/1brc The One Billion Row Challenge (1BRC) is a community coding challenge exploring how fast a Java program can aggregate min/mean/max temperatu… | 26 | 8110 | maintenance |
| snap-stanford/ogb The Open Graph Benchmark (OGB) is a collection of benchmark datasets, data loaders, and evaluators for graph machine learning, covering nod… | 32 | 2094 | stable |
| wafer-ai/gpu-perf-engineering-resources A curated list of resources for learning AI/GPU performance engineering, ordered from GPU fundamentals through kernel optimization, inferen… | 60 | 2080 | active |
| opendatalab/OmniDocBench OmniDocBench is a comprehensive benchmark dataset and evaluation toolkit for document parsing, containing 1,651 annotated PDF pages across … | 64 | 1999 | active |
| Elanis/web-to-desktop-framework-comparison An open-source comparison repository that objectively evaluates frameworks for converting web apps into desktop applications, including Ele… | 75 | 1988 | active |
| nayuki/Project-Euler-solutions A collection of runnable solutions to Project Euler math/programming problems, written in Java, Python, Mathematica, and Haskell, with benc… | 32 | 1960 | active |
| enhancedformysql/The-Art-of-Problem-Solving-in-Software-Engineering_How-to-Make-MySQL-Better An open-source online book that uses real MySQL challenges as case studies to teach problem analysis, logical reasoning, algorithms, and da… | 48 | 1934 | active |
| Thinklab-SJTU/Bench2Drive Bench2Drive is a closed-loop benchmark and dataset for end-to-end autonomous driving, built on CARLA with an RL-based expert driver (Think2… | 68 | 1926 | active |
| ashvardanian/less_slow.cpp A tutorial-style repository of benchmarks demonstrating performance-oriented coding practices in C++20, C, CUDA, PTX, and assembly, coverin… | 96 | 1924 | active |
| tau-bench τ-Bench (tau2-bench) is a Python benchmark for evaluating LLM agents on tool-agent-user interaction in real-world domains like retail, airl… | 79 | 1882 | active |
| Farama-Foundation/Metaworld Meta-World is an open-source benchmark of 50 simulated robotic manipulation tasks built on MuJoCo and the Gymnasium API, for evaluating mul… | 84 | 1871 | active |
| hkust-nlp/ceval C-Eval is a comprehensive Chinese evaluation benchmark for foundation models, consisting of 13,948 multiple-choice questions across 52 disc… | 44 | 1867 | stable |
| cfregly/ai-performance-engineering Code, labs, and resources accompanying the O'Reilly book 'AI Systems Performance Engineering', covering GPU optimization, distributed train… | 63 | 1863 | active |
| karolpiczak/ESC-50 ESC-50 is a labeled dataset of 2000 five-second environmental audio recordings organized into 50 classes for benchmarking environmental sou… | 32 | 1859 | stable |
| petergpt/bullshit-benchmark BullshitBench is a benchmark dataset and evaluation harness that tests whether AI models detect and push back on nonsensical prompts rather… | 59 | 1836 | active |
| zanfranceschi/rinha-de-backend-2024-q1 The repository for the second edition of Rinha de Backend, a Brazilian community backend performance challenge focused on concurrency contr… | 26 | 1834 | active |
| bddicken/languages A community-driven collection of microbenchmarks comparing the performance of many programming languages, with shell scripts to compile and… | 38 | 1819 | active |
| epicweb-dev/react-performance An Epic Web workshop repository teaching how to diagnose, profile, and fix performance problems in React applications using the Browser Per… | 76 | 1814 | active |
| mlcommons/training Reference implementations of the MLPerf Training benchmark suite maintained by MLCommons, covering models from LLMs to recommendation and v… | 67 | 1771 | active |
| benchflow-ai/skillsbench SkillsBench is the first benchmark for evaluating how effectively AI agents use skills—modular folders of instructions, scripts, and resour… | 76 | 1725 | active |
| kuangliu/pytorch-cifar A PyTorch reference repository for training image classification models on the CIFAR10 dataset, with implementations of many popular archit… | 32 | 6423 | maintenance |
| faster-cpython/ideas A discussion and work-tracking repository for the Faster CPython project, where performance improvement ideas for the CPython interpreter a… | 59 | 1724 | active |
| brucethemoose/Minecraft-Performance-Flags-Benchmarks A benchmarked guide of Java flags and JVM tweaks for optimizing Minecraft client and server performance, with a Python benchmarking script … | 32 | 1723 | active |
| openai/mle-bench MLE-bench is an open-source benchmark from OpenAI that measures how well AI agents perform machine learning engineering tasks, built from 7… | 58 | 1720 | active |
| KEV0143/Comparative-analysis-of-hourly-load-forecasting-using-PatchTST-TFT-NHiTS-and-CatBoost A Python benchmark project comparing deep learning time-series architectures (PatchTST, Temporal Fusion Transformer, N-HiTS) against CatBoo… | 63 | 1676 | active |
| centerforaisafety/hle Humanity's Last Exam (HLE) is a multi-modal benchmark of 2,500 expert-written questions across dozens of academic subjects, designed to tes… | 63 | 1665 | active |
| StanfordVL/BEHAVIOR-1K BEHAVIOR-1K is a simulation benchmark for embodied AI agents covering 1,000 everyday household activities across 50 interactive scenes, bui… | 99 | 1660 | active |
| eranyanay/1m-go-websockets A reference implementation and case study from a Gophercon Israel 2019 talk demonstrating a pure Go server handling 1M+ WebSocket connectio… | 32 | 5995 | maintenance |
| plutov/practice-go A collection of Go programming challenges where each challenge folder contains a README and test files to implement. Contributors submit pu… | 65 | 1625 | active |
| alecthomas/go_serialization_benchmarks A benchmark suite comparing the performance of many Go serialization libraries (JSON, Protobuf, MessagePack, gob, etc.) on a representative… | 45 | 1623 | active |
| zanfranceschi/rinha-de-backend-2025 Rinha de Backend 2025 is the third edition of a community backend coding challenge where participants build a payment-intermediation servic… | 35 | 1620 | stable |
| privatenumber/minification-benchmarks An open-source benchmark suite comparing JavaScript minifiers such as esbuild, terser, swc, oxc-minify, uglify-js, and Google Closure Compi… | 75 | 1619 | active |
| MLGroupJLU/LLM-eval-survey The official repository accompanying the survey paper 'A Survey on Evaluation of Large Language Models', curating papers and resources on L… | 71 | 1609 | active |
| BytedanceSpeech/seed-tts-eval An objective evaluation test set and metric scripts from ByteDance's seed-TTS project for benchmarking zero-shot text-to-speech and voice c… | 24 | 1595 | stable |
| ubisoft/ubisoft-laforge-animation-dataset The Ubisoft La Forge Animation Dataset (LAFAN1) is a motion capture dataset of 5 subjects, 77 sequences, and ~4.6 hours of character animat… | 32 | 1568 | stable |
| llm2014/llm_benchmark A personal, long-running benchmark that tracks large language model performance on logic, math, programming, and intuition tasks using a pr… | 65 | 1567 | active |
| SemiAnalysisAI/InferenceX InferenceX is an open-source, vendor-neutral continuous benchmarking platform that measures LLM inference performance across serving framew… | 71 | 1554 | active |
| datacurve-ai/deep-swe DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks drawn from active open-source repositories, used to measure… | 57 | 1499 | active |
| android/performance-samples Official Android sample collection demonstrating performance libraries like Macrobenchmark, Microbenchmark, and JankStats. It shows best pr… | 77 | 1440 | active |
| djiangtw/data-structures-in-practice-public An open-source technical book teaching data structures from a hardware-aware perspective, covering cache behavior, memory hierarchy, and re… | 40 | 1397 | active |
| hendrycks/math The MATH Dataset is a benchmark of 12,500 competition mathematics problems with full step-by-step solutions, used to measure mathematical p… | 50 | 1386 | stable |
| OpenDriveLab/Birds-eye-view-Perception An awesome-list and survey companion for bird's-eye-view (BEV) perception in autonomous driving, paired with an open-source PyTorch BEV too… | 37 | 1381 | active |
| AnghelLeonard/Hibernate-SpringBoot A collection of 300+ runnable sample applications demonstrating Java persistence performance best practices using Hibernate 5/6 and Spring … | 66 | 1372 | active |
| zzli2022/Awesome-System2-Reasoning-LLM A curated awesome-list tracking the latest advances in System-2 reasoning for large language models, accompanying a survey paper on reasoni… | 32 | 1353 | active |
| Liu-xiandong/How_to_optimize_in_GPU A tutorial series repository teaching CUDA kernel optimization with detailed walkthroughs of elementwise, reduce, sgemv, and sgemm kernels.… | 32 | 1350 | stable |
| pinchbench/skill PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, … | 72 | 1325 | active |
| AutoTrustAI/PaperGuru-Benchmark PaperGuru is a benchmark and research repository for Lifecycle-Aware Memory (LAM), a long-term memory primitive for long-horizon LLM agents… | 53 | 1324 | active |
| LiveBench/LiveBench LiveBench is a contamination-free benchmark for large language models that releases new questions monthly, drawn from recent datasets, pape… | 68 | 1294 | active |
| siboehm/SGEMM_CUDA An educational repository demonstrating step-by-step optimization of a CUDA SGEMM (matrix multiplication) kernel from a naive implementatio… | 50 | 1294 | stable |
| openai/frontier-evals A collection of benchmark evaluations from OpenAI for measuring frontier LLM capabilities, including PaperBench (AI paper replication), SWE… | 55 | 1287 | active |
| mims-harvard/TDC Therapeutics Data Commons (TDC) is an open-science initiative and Python library providing AI-ready datasets, machine learning tasks, and c… | 46 | 1276 | active |
| sail-sg/understand-r1-zero A research codebase and paper reproduction for critically analyzing R1-Zero-like LLM training, examining the roles of base models and reinf… | 37 | 1273 | active |
| harveyai/harvey-labs Harvey LAB (Legal Agent Benchmark) is an open-source benchmark from Harvey AI for evaluating LLM agents on realistic legal work, spanning 2… | 59 | 1261 | active |
| penberg/awesome-low-latency A curated awesome-list collecting patterns, blogs, publications, and books about low-latency programming. It codifies developer folklore on… | 48 | 1260 | active |
| webcomponents/custom-elements-everywhere A project that runs Karma test suites against major JavaScript frameworks to evaluate how well they interoperate with Custom Elements (Web … | 76 | 1256 | active |
| mlc-ai/modern-gpu-programming-for-mlsys An open online book from the MLC team teaching modern GPU kernel programming for machine learning systems, progressing from GPU hardware fu… | 59 | 1239 | active |
| BoringBoredom/PC-Optimization-Hub A curated collection of resources, guides, and tools for optimizing PC performance and reducing input lag, primarily aimed at gamers. It co… | 72 | 1237 | active |
| 0burak/imperial_hft A C++ repository of low-latency programming techniques for high-frequency trading, including cache warming, lock-free programming, loop unr… | 65 | 1237 | active |
| THUDM/LongBench LongBench is a benchmark suite (v1 and v2) for evaluating large language models on long-context understanding and reasoning tasks, with con… | 29 | 1229 | active |
| ScalingIntelligence/KernelBench KernelBench is a benchmark and toolkit from Stanford's Scaling Intelligence Lab that evaluates whether LLMs can generate correct and effici… | 55 | 1214 | active |
| BIT-DataLab/LakeBench LakeBench is a large-scale benchmark for evaluating table discovery methods (joinable and unionable table search) in data lakes, containing… | 37 | 1211 | active |
page 1 / 3 next →