Ross ROSS = Recommend OSS · open-source software intelligence for agents

HKUDS/ClawWork

"ClawWork: OpenClaw as Your AI Coworker - 💰 $15K earned in 11 Hours" observed · 2026-08-28

github.com/HKUDS/ClawWork · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

47/100

  • Activity 70
  • Release rhythm 35
  • Longevity 14

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 199
  • days_rel: n/a
  • days_push: 183
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

8522 stars · 1101 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

ClawWork is an open-source benchmark and framework that turns AI assistants into 'AI coworkers' which must earn income by completing real professional tasks from the GDPVal dataset while paying for their own token usage. It includes an economic survival evaluation system, leaderboards across LLM agents, and a local dashboard for tracking agent income, cost, and work quality.

Use cases

  • benchmark ai agents on real-world economic tasks
  • compare llm agents by income earned and cost efficiency
  • evaluate ai work quality across 44 professions
  • run an economic survival simulation for autonomous agents
  • track agent performance on the gdpval task dataset
  • visualize agent leaderboard results in a local dashboard

When to choose

  • you want to measure LLM agents on real professional work tasks rather than static benchmarks
  • you need cost-aware agent evaluation that factors in token expenses
  • you want to compare models like Qwen, Gemini, GLM, or Kimi on economic survival metrics

When to avoid

  • you need a production AI assistant for actual business work rather than an evaluation harness
  • you only need standard agent benchmarks like SWE-bench or MMLU
  • you require guaranteed reproducible results, since live LLM APIs introduce variance and cost

Facets

framework · maturity active

agent-framework benchmarking llm-inference data-visualization analytics artificial-intelligence large-language-models data-science developer-tools python cli cross-platform economic-benchmark ai-coworker gdpval agent-evaluation llm-agents dashboard ai-agents docker

1 source

Member repositories

RepositoryRoleHealth v2
HKUDS/ClawWorkmain47

For agents

markdown · JSON · MCP: product_card(name="HKUDS/ClawWork")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem