Open account

Huawei's New Benchmark Gives AI Agents Months of Your Life—Then Watches Them Fail

Researchers from Huawei and partner institutions released Claw-Anything, a benchmark testing AI agents on realistic, long-horizon personal-assistant tasks. Models performed poorly: OpenAI’s GPT-5.5 scored 34.5% pass@1 overall (25.9% reactive, 6.7% proactive), highlighting sizable gaps between benchmark performance and real-world assistant capabilities. The benchmark uses massive context (avg. 191,700 words per task) and multi-service, multi-device scenarios; removing cross-service tools collapses success rates. On the constructive side, the team published 2,000 training environments and an automated pipeline; fine-tuning Qwen3.5-27B on 1,500 successful trajectories raised pass@1 by 23.7%, outperforming several closed models. Market implication: core AI products and cloud/software vendors may face delayed enterprise adoption for autonomous agent features until cross-service coordination and long-horizon reasoning improve, benefiting firms that lead in scalable fine-tuning and integration.

Category

Microsoft

Sentiment

Mixed

Event

Market commentary

Reading time

1 min