Qwen introduces Qwen-CUA, a native computer-use agent built on a 397B-A17B mixture-of-experts backbone that operates via screenshots and input events without relying on DOM trees or accessibility metadata. The system utilizes a scaffold to maintain up to 20 active screenshots and was trained using a cloud rollout fleet with nearly 100,000 vCPUs across approximately 40,000 verifiable tasks.
- Qwen-CUA outperforms Qwen3.7 on eight benchmarks, achieving 86.2 on OSWorld-Verified and competitive scores on OSWorld 2.0.
- Scaling the approach to over one trillion parameters yields Qwen-CUA-Max, which improves OSWorld-Verified scores to 87.6.
- The model reduces RedTeamCUA attack success rates from 36.6 to 16.4 relative to Qwen3.7.
These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction as a key direction for future development.