FEATUREDAI HOT (Curated Pool)· aihot-apiZH09:45 · 08·20
→Alibaba releases Qwen-UI-Agent, a GUI agent foundation model that operates across mobile, desktop, web, and deep search
Alibaba's Qwen team released Qwen-UI-Agent, a model built to understand and operate mobile, desktop, and web interfaces. It scores 92.2% on the real-device benchmark MobileWorld-Real and 97.5% on AndroidDaily. On desktop, it hits 79.5% on OSWorld-Verified, beating GPT-5.5 and Gemini 3.1 Pro. Training used over 100 real phones and 150+ apps, and the model handles tasks longer than 100 steps. It pauses for user confirmation on payments or privacy actions. The tech report, project page, and GitHub repo are all public.
#Alibaba#Qwen#Qwen-UI-Agent
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
Alibaba's Qwen-UI-Agent hits 92.2% on real-device tests, beating GPT-5.5 and Gemini 3.1 Pro, with full tech report and code released.
sharp
This one's worth opening because it moves GUI agents from simulators to real phones. Alibaba trained and evaluated on 100+ real devices and 150+ apps — that 92.2% on MobileWorld-Real isn't a simulator score. On desktop, 79.5% on OSWorld-Verified edges out GPT-5.5.
Two things I'd watch. First, it handles tasks over 100 steps with online RL training, which matters more than single-step accuracy in real use. Second, the safety design: it pauses for user confirmation on payments or privacy actions instead of blindly executing.
What's missing: model size and inference latency. The post doesn't mention either. If it's a large model, real-device deployment won't be cheap.
HKR breakdown
hook ✓knowledge ✓resonance ✓