Ross ROSS = Recommend OSS · open-source software intelligence for agents

sunny-glow/Auto-BenchMax

None observed · 2026-08-28

github.com/sunny-glow/Auto-BenchMax · Python observed · 2026-08-28

Health v2 · maintenance only

55/100

  • Activity 94
  • Release rhythm 35
  • Longevity 2

Flags: no_releases young no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 41
  • days_rel: n/a
  • days_push: 40
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1317 stars · 27 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Auto-BenchMax is a Python pipeline that automatically synthesizes benchmark-targeted training data for LLMs, claiming to more than double a base model's score on agentic benchmarks like MCP-Atlas and Tau2-bench. It ships a data-construction pipeline, training scripts, an evaluation environment, and a skill that lets agent tools like Claude Code drive synthesis for custom tool-use benchmarks.

Use cases

  • synthesize training data for an agentic benchmark
  • improve my model's score on MCP-Atlas
  • generate fine-tuning data targeted at a tool-use benchmark
  • reproduce benchmark score improvements with one click
  • create LLM-judged or rule-based benchmark training sets
  • fine-tune a model to double its benchmark baseline

When to choose

  • you need to boost a model's score on a specific tool-use or agentic benchmark
  • you want a reproducible data-synthesis plus training pipeline
  • you use a skill-capable agent like Claude Code and want one-sentence automation

When to avoid

  • you need genuinely general capability gains rather than benchmark-targeted optimization
  • you lack the compute to fine-tune models
  • you need a license-cleared project for commercial use, since no license is specified

Facets

library · maturity active

llm-training data-generation agent-framework machine-learning large-language-models machine-learning artificial-intelligence developer-tools python cli benchmark-optimization synthetic-data fine-tuning tool-use skill-based-pipeline llm-judge training-scripts ai-agents linux

1 source

Member repositories

RepositoryRoleHealth v2
sunny-glow/Auto-BenchMaxmain55

For agents

markdown · JSON · MCP: product_card(name="sunny-glow/Auto-BenchMax")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem