modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. observed · 2026-08-28
Health v2 · maintenance only
93/100
- Activity 99
- Release rhythm 99
- Longevity 71
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 13.0
- age_days: 1000
- days_rel: 9
- days_push: 7
- n_releases_24m: 49
Adoption not part of the score
3312 stars · 467 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
EvalScope is a Python framework from the ModelScope community for evaluating large language models, vision-language models, embedding models, and AIGC models against built-in benchmarks like MMLU, C-Eval, and GSM8K. It also provides inference performance stress testing, agent-based evaluation with sandboxed tool use, multi-model arena battles, and a web dashboard for result visualization.
Use cases
- benchmark an LLM on MMLU and GSM8K
- stress test inference performance of a model service
- compare two models head-to-head in an arena
- evaluate a vision-language model
- evaluate RAG pipelines with rerankers and embeddings
- run agentic benchmarks like SWE-bench in a sandbox
- visualize evaluation results in a dashboard
When to choose
- you need a one-command evaluation pipeline for LLMs or VLMs via OpenAI-compatible APIs
- you want both quality benchmarks and inference performance metrics like TTFT and TPOT in one tool
- you need multi-backend support spanning OpenCompass, VLMEvalKit, and RAGEval
- you want pairwise model battles and interactive comparison reports
When to avoid
- you only need simple unit testing of application code rather than model evaluation
- you require a fully managed cloud evaluation service with no local setup
- your models are not reachable via supported API or local inference backends
Facets
framework · maturity active
benchmarking testing machine-learning llm-inference rag data-visualization large-language-models machine-learning artificial-intelligence developer-tools performance python cli cross-platform llm-evaluation vlm-evaluation stress-testing arena-mode model-benchmarking openai-api agent-evaluation
2 sources
- readme: https://github.com/modelscope/evalscope · fetched 2026-08-28 · e0dba4034242
- registry_pypi: https://pypi.org/pypi/evalscope/json · fetched 2026-08-29 · 0022424af9c8
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| modelscope/evalscope | main | 93 |
For agents
markdown · JSON · MCP: product_card(name="modelscope/evalscope")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem