Ross ROSS = Recommend OSS · open-source software intelligence for agents

AMAP-ML/LongHorizon-Harness

The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration. observed · 2026-08-28

github.com/AMAP-ML/LongHorizon-Harness · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

79/100

  • Activity 98
  • Release rhythm 98
  • Longevity 2

Flags: young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 2
  • age_days: 29
  • days_rel: 13
  • days_push: 13
  • n_releases_24m: 8

Full methodology

Adoption not part of the score

1332 stars · 136 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

LongHorizon-Harness is a Python-based loop engineering harness that turns existing coding agents (Claude Code, Codex, OpenCode, DeepSeek) into long-running computer-use systems across desktop apps and the CLI. It manages a plan-act-verify-checkpoint loop with durable, independently audited task state so agents can make reliable progress on multi-hour tasks.

Use cases

  • run AI agents on long-horizon tasks for dozens of hours
  • keep an agent working across desktop apps and the terminal
  • recover agent progress after failure or context refresh
  • independently verify each step an agent completes
  • benchmark agents on WeaveBench, OSWorld, and Terminal-Bench
  • orchestrate different models for manage, execute, and audit roles
  • checkpoint and resume complex multi-step workflows

When to choose

  • you need an existing agent like Claude Code or Codex to run reliably for extended periods
  • your tasks span GUI desktop apps and CLI tools
  • you want audited, verifiable progress records outside the agent session
  • you want to improve benchmark scores without changing the underlying model

When to avoid

  • you need a simple single-shot agent run without state management
  • you want to train or fine-tune a model rather than orchestrate one
  • you need a fully managed cloud service rather than a local Python harness
  • your agent runtime is not supported (no Claude Code, Codex, OpenCode, or DeepSeek backend)

Facets

framework · maturity active

agent-framework workflow-automation developer-tools cli gui developer-tools large-language-models windows cli python loop-engineering computer-use long-horizon-agents claude-code codex agent-harness checkpointing task-orchestration mea-loop ai-agents automation linux macos

2 sources

Member repositories

RepositoryRoleHealth v2
AMAP-ML/LongHorizon-Harnessmain79

For agents

markdown · JSON · MCP: product_card(name="AMAP-ML/LongHorizon-Harness")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem