AMAP-ML/LongHorizon-Harness
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration. observed · 2026-08-28
Health v2 · maintenance only
79/100
- Activity 98
- Release rhythm 98
- Longevity 2
Flags: young
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 2
- age_days: 29
- days_rel: 13
- days_push: 13
- n_releases_24m: 8
Adoption not part of the score
1332 stars · 136 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
LongHorizon-Harness is a Python-based loop engineering harness that turns existing coding agents (Claude Code, Codex, OpenCode, DeepSeek) into long-running computer-use systems across desktop apps and the CLI. It manages a plan-act-verify-checkpoint loop with durable, independently audited task state so agents can make reliable progress on multi-hour tasks.
Use cases
- run AI agents on long-horizon tasks for dozens of hours
- keep an agent working across desktop apps and the terminal
- recover agent progress after failure or context refresh
- independently verify each step an agent completes
- benchmark agents on WeaveBench, OSWorld, and Terminal-Bench
- orchestrate different models for manage, execute, and audit roles
- checkpoint and resume complex multi-step workflows
When to choose
- you need an existing agent like Claude Code or Codex to run reliably for extended periods
- your tasks span GUI desktop apps and CLI tools
- you want audited, verifiable progress records outside the agent session
- you want to improve benchmark scores without changing the underlying model
When to avoid
- you need a simple single-shot agent run without state management
- you want to train or fine-tune a model rather than orchestrate one
- you need a fully managed cloud service rather than a local Python harness
- your agent runtime is not supported (no Claude Code, Codex, OpenCode, or DeepSeek backend)
Facets
framework · maturity active
agent-framework workflow-automation developer-tools cli gui developer-tools large-language-models windows cli python loop-engineering computer-use long-horizon-agents claude-code codex agent-harness checkpointing task-orchestration mea-loop ai-agents automation linux macos
2 sources
- readme: https://github.com/AMAP-ML/LongHorizon-Harness · fetched 2026-08-28 · 75b0aa90cb69
- homepage: https://lh-harness.pages.dev · fetched 2026-08-29 · e7b2264fdf9e
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| AMAP-ML/LongHorizon-Harness | main | 79 |
For agents
markdown · JSON · MCP: product_card(name="AMAP-ML/LongHorizon-Harness")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem