Ross ROSS = Recommend OSS · open-source software intelligence for agents

web-infra-dev/midscene

GUI Agent for E2E Testing observed · 2026-08-28

github.com/web-infra-dev/midscene · homepage · TypeScript · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

86/100

  • Activity 99
  • Release rhythm 87
  • Longevity 55
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 2.0
  • age_days: 771
  • days_rel: 6
  • days_push: 7
  • n_releases_24m: 187

Full methodology

Adoption not part of the score

14714 stars · 1134 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Midscene.js is an open-source TypeScript SDK that acts as a GUI agent for E2E testing and UI automation, using multimodal vision models to plan and operate interfaces from natural-language instructions. It works across web (Playwright/Puppeteer), Android, iOS, HarmonyOS, and desktop apps, reaching elements that DOM/accessibility-tree-based tools cannot, such as canvas and native controls.

Use cases

  • write e2e tests with natural language instead of css selectors
  • automate web forms in playwright or puppeteer using ai vision
  • test android and ios apps on real devices with plain english commands
  • automate desktop applications on macos windows and linux
  • interact with canvas elements and unlabeled buttons that selectors cannot reach
  • run yaml-scripted ui automation flows
  • automate any interface from screenshots

When to choose

  • your UI has fragile or missing selectors, canvas rendering, or cross-origin frames
  • you want one natural-language automation API across web, mobile, and desktop
  • you need AI-driven E2E tests integrated with Playwright or Puppeteer
  • you want to automate native apps without accessibility-tree access

When to avoid

  • you need fast, deterministic, low-cost tests and stable selectors already work
  • you cannot send screenshots to external or self-hosted vision models
  • your tests require strict reproducibility without AI variability
  • you need offline automation with no LLM access

Facets

library · maturity active

e2e-testing testing computer-vision agent-framework workflow-automation testing developer-tools artificial-intelligence web-development cross-platform windows cli gui-agent vision-driven-automation ui-testing playwright puppeteer multimodal-llm yaml-automation mobile-automation ai-testing automation nodejs web-server android ios macos linux

4 sources

Member repositories

RepositoryRoleHealth v2
web-infra-dev/midscenemain86

For agents

markdown · JSON · MCP: product_card(name="web-infra-dev/midscene")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem