Ross ROSS = Recommend OSS · open-source software intelligence for agents

resource: benchmarking

233 resources, primary matches first, then adoption-weighted; health v2 shown.

ResourceHealth v2StarsMaturity
tastejs/todomvc
TodoMVC is a collection of the same Todo application implemented in dozens of JavaScript frameworks and libraries, letting developers compa…
7828949active
sebastianruder/NLP-progress
A community-maintained repository and website tracking state-of-the-art results across the most common NLP tasks, with datasets, papers, an…
2322952active
xlite-dev/LeetCUDA
A collection of modern CUDA learning notes with PyTorch integration for beginners, featuring 200+ CUDA kernels, 100+ LLM/CUDA blogs, and hi…
9311836active
davidsonfellipe/awesome-wpo
A curated awesome-list of Web Performance Optimization resources including tools, articles, books, documentation, and talks. It covers cate…
759053active
atinfo/awesome-test-automation
A curated awesome-list of test automation frameworks, tools, libraries, and software organized by programming language (Python, Java, Ruby,…
577146active
jeinlee1991/chinese-llm-benchmark
ReLE (formerly CLiB) is a continuously updated Chinese-language LLM capability benchmark and leaderboard covering ~400 commercial and open-…
896401active
microsoft/promptbase
A Microsoft-maintained collection of prompt engineering resources, best practices, and example Python scripts for eliciting top performance…
265771active
fastruby/fast-ruby
A curated collection of common Ruby idioms with benchmark comparisons showing which constructs run faster, inspired by Erik Michaels-Ober's…
725734active
OpenBMB/ToolBench
ToolBench is an open-source platform for training, serving, and evaluating large language models on tool use, built around a large-scale in…
395732active
SWE-bench/SWE-bench
SWE-bench is a benchmark and evaluation harness that tests whether large language models can resolve real-world GitHub issues by generating…
725719active
KellerJordan/modded-nanogpt
A collaborative speedrun project that trains a GPT-2 (124M) scale language model to 3.28 validation loss on FineWeb in under 75 seconds on …
675707active
sirupsen/napkin-math
A collection of memorized latency/throughput numbers, benchmark code, and techniques for estimating system performance from first principle…
655671active
openai/parameter-golf
An OpenAI-hosted challenge to train the best-performing language model that fits in a 16MB artifact and trains in under 10 minutes on 8xH10…
515174active
fchollet/ARC-AGI
The Abstraction and Reasoning Corpus (ARC-AGI-1), a benchmark dataset of 800 JSON grid-based tasks designed to test human-like fluid genera…
294813stable
nikolaydubina/go-recipes
A curated awesome-list of handy and lesser-known command-line tools for Go projects, covering testing, coverage, benchmarking, profiling, c…
754500active
datacharmer/test_db
A sample MySQL employees database with an integrated test suite, migrated from Launchpad, also compatible with MariaDB, Percona Server, and…
574434stable
CLUEbenchmark/CLUE
CLUE is the Chinese Language Understanding Evaluation Benchmark, providing representative datasets, baseline and pre-trained models, corpor…
624279active
anthropics/original_performance_takehome
Anthropic's open-sourced performance engineering take-home challenge, where you optimize a simulated machine's code from a slow baseline (2…
444122stable
dendibakh/perf-ninja
An online course in the form of C++ lab assignments for learning low-level performance analysis and tuning, covering issues like CPU cache …
773825active
denji/awesome-http-benchmark
A curated awesome-list of HTTP(S) benchmarking, load testing, and REST API testing/debugging tools. It catalogs tools like ab, wrk, autocan…
763765active
THUDM/AgentBench
AgentBench is a comprehensive benchmark for evaluating large language models as agents across multi-turn interactive tasks like database qu…
583695active
waymo-research/waymo-open-dataset
The Waymo Open Dataset is a large-scale public collection of autonomous driving datasets (Perception, Motion, and End-to-End Driving) accom…
503398active
vectara/hallucination-leaderboard
A public leaderboard maintained by Vectara that ranks LLMs by how often they hallucinate when summarizing short documents, computed with Ve…
643306active
CLUEbenchmark/SuperCLUE
SuperCLUE is a comprehensive benchmark for evaluating Chinese-language large language models across language understanding and generation, …
593299active
adamsitnik/awesome-dot-net-performance
A curated list of .NET performance resources including books, courses, trainings, conference talks, blogs, and notable open source contribu…
683279active
xlang-ai/OSWorld
OSWorld is a benchmark and scalable real computer environment for evaluating multimodal agents on open-ended computer tasks across operatin…
623108active
PlummersSoftwareLLC/Primes
A community-maintained collection of prime number sieve implementations in 100+ programming languages, used to benchmark and compare langua…
733017active
Kobzol/hardware-effects
A collection of small C++ proof-of-concept programs demonstrating hardware effects like cache conflicts, branch misprediction, false sharin…
322999stable
kostya/benchmarks
A collection of cross-language performance benchmarks comparing programming languages on tasks like Brainfuck interpretation, Base64, JSON,…
742922active
Tony-Tan/CUDA_Freshman
A collection of CUDA example programs accompanying a Chinese-language blog tutorial series on GPU programming, partly based on the book 'Pr…
322792stable
dgryski/go-perfbook
An open-source book by Damian Gryski collecting best practices for writing high-performance Go code, covering general optimization methodol…
3210896maintenance
asatarin/testing-distributed-systems
A curated list of resources on testing distributed systems, including research papers, tools, and case studies. It covers approaches like J…
752635active
harbor-framework/terminal-bench-1
Terminal-Bench is a benchmark and execution harness for evaluating LLM agents on complex, real-world tasks in a terminal sandbox. It combin…
622554active
DestinyLinker/MingLi-Bench
A Python benchmark dataset and CLI for evaluating large language models on Chinese traditional fortune telling, covering Bazi (八字) and Ziwe…
502338active
PacktPublishing/The-Kaggle-Book
The official code repository for 'The Kaggle Book' by Packt Publishing, containing Jupyter Notebook examples from two Kaggle Grandmasters. …
642333active
beir-cellar/beir
BEIR is a heterogeneous benchmark for information retrieval, aggregating 15+ diverse IR datasets with a common evaluation framework for NLP…
472275active
Lifelong-Robot-Learning/LIBERO
LIBERO is a benchmark for studying knowledge transfer in multitask and lifelong robot learning, built on a procedural generation pipeline f…
352238active
AmberLJC/LLMSys-PaperList
A curated list of academic papers, tutorials, and projects on Large Language Model systems, covering training, serving, agentic systems, an…
712233active
alibaba/clusterdata
A collection of production cluster trace datasets from Alibaba covering batch workloads, GPU/ML workloads, microservices, and microarchitec…
702154active
gunnarmorling/1brc
The One Billion Row Challenge (1BRC) is a community coding challenge exploring how fast a Java program can aggregate min/mean/max temperatu…
268110maintenance
snap-stanford/ogb
The Open Graph Benchmark (OGB) is a collection of benchmark datasets, data loaders, and evaluators for graph machine learning, covering nod…
322094stable
wafer-ai/gpu-perf-engineering-resources
A curated list of resources for learning AI/GPU performance engineering, ordered from GPU fundamentals through kernel optimization, inferen…
602080active
opendatalab/OmniDocBench
OmniDocBench is a comprehensive benchmark dataset and evaluation toolkit for document parsing, containing 1,651 annotated PDF pages across …
641999active
Elanis/web-to-desktop-framework-comparison
An open-source comparison repository that objectively evaluates frameworks for converting web apps into desktop applications, including Ele…
751988active
nayuki/Project-Euler-solutions
A collection of runnable solutions to Project Euler math/programming problems, written in Java, Python, Mathematica, and Haskell, with benc…
321960active
enhancedformysql/The-Art-of-Problem-Solving-in-Software-Engineering_How-to-Make-MySQL-Better
An open-source online book that uses real MySQL challenges as case studies to teach problem analysis, logical reasoning, algorithms, and da…
481934active
Thinklab-SJTU/Bench2Drive
Bench2Drive is a closed-loop benchmark and dataset for end-to-end autonomous driving, built on CARLA with an RL-based expert driver (Think2…
681926active
ashvardanian/less_slow.cpp
A tutorial-style repository of benchmarks demonstrating performance-oriented coding practices in C++20, C, CUDA, PTX, and assembly, coverin…
961924active
tau-bench
τ-Bench (tau2-bench) is a Python benchmark for evaluating LLM agents on tool-agent-user interaction in real-world domains like retail, airl…
791882active
Farama-Foundation/Metaworld
Meta-World is an open-source benchmark of 50 simulated robotic manipulation tasks built on MuJoCo and the Gymnasium API, for evaluating mul…
841871active
hkust-nlp/ceval
C-Eval is a comprehensive Chinese evaluation benchmark for foundation models, consisting of 13,948 multiple-choice questions across 52 disc…
441867stable
cfregly/ai-performance-engineering
Code, labs, and resources accompanying the O'Reilly book 'AI Systems Performance Engineering', covering GPU optimization, distributed train…
631863active
karolpiczak/ESC-50
ESC-50 is a labeled dataset of 2000 five-second environmental audio recordings organized into 50 classes for benchmarking environmental sou…
321859stable
petergpt/bullshit-benchmark
BullshitBench is a benchmark dataset and evaluation harness that tests whether AI models detect and push back on nonsensical prompts rather…
591836active
zanfranceschi/rinha-de-backend-2024-q1
The repository for the second edition of Rinha de Backend, a Brazilian community backend performance challenge focused on concurrency contr…
261834active
bddicken/languages
A community-driven collection of microbenchmarks comparing the performance of many programming languages, with shell scripts to compile and…
381819active
epicweb-dev/react-performance
An Epic Web workshop repository teaching how to diagnose, profile, and fix performance problems in React applications using the Browser Per…
761814active
mlcommons/training
Reference implementations of the MLPerf Training benchmark suite maintained by MLCommons, covering models from LLMs to recommendation and v…
671771active
benchflow-ai/skillsbench
SkillsBench is the first benchmark for evaluating how effectively AI agents use skills—modular folders of instructions, scripts, and resour…
761725active
kuangliu/pytorch-cifar
A PyTorch reference repository for training image classification models on the CIFAR10 dataset, with implementations of many popular archit…
326423maintenance
faster-cpython/ideas
A discussion and work-tracking repository for the Faster CPython project, where performance improvement ideas for the CPython interpreter a…
591724active
brucethemoose/Minecraft-Performance-Flags-Benchmarks
A benchmarked guide of Java flags and JVM tweaks for optimizing Minecraft client and server performance, with a Python benchmarking script …
321723active
openai/mle-bench
MLE-bench is an open-source benchmark from OpenAI that measures how well AI agents perform machine learning engineering tasks, built from 7…
581720active
KEV0143/Comparative-analysis-of-hourly-load-forecasting-using-PatchTST-TFT-NHiTS-and-CatBoost
A Python benchmark project comparing deep learning time-series architectures (PatchTST, Temporal Fusion Transformer, N-HiTS) against CatBoo…
631676active
centerforaisafety/hle
Humanity's Last Exam (HLE) is a multi-modal benchmark of 2,500 expert-written questions across dozens of academic subjects, designed to tes…
631665active
StanfordVL/BEHAVIOR-1K
BEHAVIOR-1K is a simulation benchmark for embodied AI agents covering 1,000 everyday household activities across 50 interactive scenes, bui…
991660active
eranyanay/1m-go-websockets
A reference implementation and case study from a Gophercon Israel 2019 talk demonstrating a pure Go server handling 1M+ WebSocket connectio…
325995maintenance
plutov/practice-go
A collection of Go programming challenges where each challenge folder contains a README and test files to implement. Contributors submit pu…
651625active
alecthomas/go_serialization_benchmarks
A benchmark suite comparing the performance of many Go serialization libraries (JSON, Protobuf, MessagePack, gob, etc.) on a representative…
451623active
zanfranceschi/rinha-de-backend-2025
Rinha de Backend 2025 is the third edition of a community backend coding challenge where participants build a payment-intermediation servic…
351620stable
privatenumber/minification-benchmarks
An open-source benchmark suite comparing JavaScript minifiers such as esbuild, terser, swc, oxc-minify, uglify-js, and Google Closure Compi…
751619active
MLGroupJLU/LLM-eval-survey
The official repository accompanying the survey paper 'A Survey on Evaluation of Large Language Models', curating papers and resources on L…
711609active
BytedanceSpeech/seed-tts-eval
An objective evaluation test set and metric scripts from ByteDance's seed-TTS project for benchmarking zero-shot text-to-speech and voice c…
241595stable
ubisoft/ubisoft-laforge-animation-dataset
The Ubisoft La Forge Animation Dataset (LAFAN1) is a motion capture dataset of 5 subjects, 77 sequences, and ~4.6 hours of character animat…
321568stable
llm2014/llm_benchmark
A personal, long-running benchmark that tracks large language model performance on logic, math, programming, and intuition tasks using a pr…
651567active
SemiAnalysisAI/InferenceX
InferenceX is an open-source, vendor-neutral continuous benchmarking platform that measures LLM inference performance across serving framew…
711554active
datacurve-ai/deep-swe
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks drawn from active open-source repositories, used to measure…
571499active
android/performance-samples
Official Android sample collection demonstrating performance libraries like Macrobenchmark, Microbenchmark, and JankStats. It shows best pr…
771440active
djiangtw/data-structures-in-practice-public
An open-source technical book teaching data structures from a hardware-aware perspective, covering cache behavior, memory hierarchy, and re…
401397active
hendrycks/math
The MATH Dataset is a benchmark of 12,500 competition mathematics problems with full step-by-step solutions, used to measure mathematical p…
501386stable
OpenDriveLab/Birds-eye-view-Perception
An awesome-list and survey companion for bird's-eye-view (BEV) perception in autonomous driving, paired with an open-source PyTorch BEV too…
371381active
AnghelLeonard/Hibernate-SpringBoot
A collection of 300+ runnable sample applications demonstrating Java persistence performance best practices using Hibernate 5/6 and Spring …
661372active
zzli2022/Awesome-System2-Reasoning-LLM
A curated awesome-list tracking the latest advances in System-2 reasoning for large language models, accompanying a survey paper on reasoni…
321353active
Liu-xiandong/How_to_optimize_in_GPU
A tutorial series repository teaching CUDA kernel optimization with detailed walkthroughs of elementwise, reduce, sgemv, and sgemm kernels.…
321350stable
pinchbench/skill
PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, …
721325active
AutoTrustAI/PaperGuru-Benchmark
PaperGuru is a benchmark and research repository for Lifecycle-Aware Memory (LAM), a long-term memory primitive for long-horizon LLM agents…
531324active
LiveBench/LiveBench
LiveBench is a contamination-free benchmark for large language models that releases new questions monthly, drawn from recent datasets, pape…
681294active
siboehm/SGEMM_CUDA
An educational repository demonstrating step-by-step optimization of a CUDA SGEMM (matrix multiplication) kernel from a naive implementatio…
501294stable
openai/frontier-evals
A collection of benchmark evaluations from OpenAI for measuring frontier LLM capabilities, including PaperBench (AI paper replication), SWE…
551287active
mims-harvard/TDC
Therapeutics Data Commons (TDC) is an open-science initiative and Python library providing AI-ready datasets, machine learning tasks, and c…
461276active
sail-sg/understand-r1-zero
A research codebase and paper reproduction for critically analyzing R1-Zero-like LLM training, examining the roles of base models and reinf…
371273active
harveyai/harvey-labs
Harvey LAB (Legal Agent Benchmark) is an open-source benchmark from Harvey AI for evaluating LLM agents on realistic legal work, spanning 2…
591261active
penberg/awesome-low-latency
A curated awesome-list collecting patterns, blogs, publications, and books about low-latency programming. It codifies developer folklore on…
481260active
webcomponents/custom-elements-everywhere
A project that runs Karma test suites against major JavaScript frameworks to evaluate how well they interoperate with Custom Elements (Web …
761256active
mlc-ai/modern-gpu-programming-for-mlsys
An open online book from the MLC team teaching modern GPU kernel programming for machine learning systems, progressing from GPU hardware fu…
591239active
BoringBoredom/PC-Optimization-Hub
A curated collection of resources, guides, and tools for optimizing PC performance and reducing input lag, primarily aimed at gamers. It co…
721237active
0burak/imperial_hft
A C++ repository of low-latency programming techniques for high-frequency trading, including cache warming, lock-free programming, loop unr…
651237active
THUDM/LongBench
LongBench is a benchmark suite (v1 and v2) for evaluating large language models on long-context understanding and reasoning tasks, with con…
291229active
ScalingIntelligence/KernelBench
KernelBench is a benchmark and toolkit from Stanford's Scaling Intelligence Lab that evaluates whether LLMs can generate correct and effici…
551214active
BIT-DataLab/LakeBench
LakeBench is a large-scale benchmark for evaluating table discovery methods (joinable and unionable table search) in data lakes, containing…
371211active

page 1 / 3 next →