# Tongyi-MAI/MAI-UI

Qwen-UI-Agent: Towards Next-Generation Real-World Centric Foundation GUI Agent

Repository: https://github.com/Tongyi-MAI/MAI-UI
Canonical: https://ross.abutalabs.com/products/mai-ui
Homepage: https://tongyi-mai.github.io/Qwen-UI-Agent/
Language: Jupyter Notebook
License Family: other
Topics: gui-agent, gui-grounding, gui-navigation, browser-use, cli-agent, computer-use, deepresearch, mobile-use
Last push: 2026-08-19T04:25:41+00:00

## Health v2 (maintenance only)
Score: 60/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 98, release rhythm 35, longevity 18
- inputs: {"age_days": 261, "days_push": 14, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2228, forks 202 (observed 2026-08-28T04:06:28.717045+00:00)

## What it is
Qwen-UI-Agent (MAI-UI) is a foundation GUI agent model from Alibaba's Tongyi-MAI team that unifies mobile, desktop, browser, and deep-research task execution in a single model. It combines GUI grounding and navigation with a hybrid GUI+CLI action space, trained via scalable long-horizon online reinforcement learning on real devices.

## Use cases
- automate tasks on android apps via a gui agent
- control a desktop computer with natural language instructions
- navigate websites and complete browser tasks automatically
- perform deep research combining web search and gui actions
- evaluate gui agents on real-device mobile benchmarks
- build an agent that clicks, types, and runs shell commands
- benchmark screen grounding on screenspot-pro or osworld

## When to choose
- you need a state-of-the-art open GUI agent spanning mobile, desktop, and web
- you want real-device training/evaluation environments instead of simulators
- you need hybrid GUI plus CLI action execution with batched actions
- you are researching long-horizon online RL for agents

## When to avoid
- you need a simple RPA tool without ML models
- you require a permissively licensed project (no license is specified)
- you lack GPU resources for running large vision-language models
- you only need lightweight browser automation like Playwright scripts

## Facets
- artifact type: framework
- maturity: active
- function: agent-framework, computer-vision, machine-learning, llm-inference, gui
- domain: artificial-intelligence, large-language-models, computer-vision, mobile-development, web-development
- platform: cross-platform, python
- tags: gui-agent, gui-grounding, computer-use, mobile-use, browser-use, deep-research, reinforcement-learning, benchmark, qwen, alibaba, ai-agents, automation, android, gpu, web-server

## Member repositories
- Tongyi-MAI/MAI-UI (main) score 60

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:28.717045+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:44:41.812083+00:00, confidence not recorded.
  - readme: https://github.com/Tongyi-MAI/MAI-UI (fetched 2026-08-28T04:06:28.717045+00:00, sha b6e5569c537e)
  - homepage: https://tongyi-mai.github.io/Qwen-UI-Agent/ (fetched 2026-08-29T10:25:22.634627+00:00, sha 7a37ed3deb0d)
- Data as of 2026-08-30T08:39:29.467469+00:00.
