Ross ROSS = Recommend OSS · open-source software intelligence for agents

harbor-framework/harbor

Framework for evaluating and improving agents observed · 2026-08-28

github.com/harbor-framework/harbor · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

85/100

  • Activity 99
  • Release rhythm 99
  • Longevity 28
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 4.0
  • age_days: 394
  • days_rel: 11
  • days_push: 7
  • n_releases_24m: 27

Full methodology

Adoption not part of the score

4664 stars · 1662 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Harbor is a Python framework from the creators of Terminal-Bench for evaluating and optimizing AI agents and language models in containerized sandbox environments. It provides modular interfaces for tasks, agents, and environments, integrates popular CLI agents and cloud sandbox providers, and supports generating rollouts for RL optimization.

Use cases

  • evaluate coding agents like Claude Code on Terminal-Bench 2.0
  • run SWE-Bench or Aider Polyglot benchmarks against a model
  • build and share custom agent benchmarks and task environments
  • run thousands of eval environments in parallel on cloud sandboxes
  • generate rollout traces for reinforcement learning or SFT training
  • test agents in CI/CD pipelines

When to choose

  • you need a standardized harness to benchmark agents across many tasks
  • you want to scale agent evaluations horizontally across cloud sandbox providers
  • you are building custom evals or RL environments for LLM agents
  • you want pre-integrated CLI agents like Claude Code, OpenHands, or Codex CLI

When to avoid

  • you only need simple LLM API calls without task evaluation
  • you need a GUI-based evaluation dashboard rather than a CLI harness
  • your evaluation targets are not containerizable tasks

Facets

framework · maturity active

agent-framework llm-inference llm-training benchmarking testing rag artificial-intelligence large-language-models developer-tools machine-learning python cli cross-platform cloud agent-evaluation evals terminal-bench reinforcement-learning-environments sandboxed-tasks benchmark-harness rollouts ai-agents docker

3 sources

Member repositories

RepositoryRoleHealth v2
harbor-framework/harbormain85

For agents

markdown · JSON · MCP: product_card(name="harbor-framework/harbor")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem