Computer-Use Agents: From Demo to Deployment
If every model can click, clicking is no longer the moat.
Foreword
Computer-use agents went from demo to deployment in about a year.
- 42% → 85%. The best score on OSWorld-Verified, a benchmark of real desktop work, roughly doubled from early 2025 to 2026, past the 72% or so that a16z cites for humans.
- No API needed. An agent that can see a screen, click and type can reach the long tail of legacy portals and internal tools that only people could operate.
- Not yet reliable. On longer, realistic workflows the leading system finished only 20.6% of its tasks, and agents still take several times more steps than they need.
This edition, drawn mainly from a16z's August 2026 field study, looks at:
- where computer-use agents already work in production;
- what they cost against offshore and US back-office labor;
- why the moat is moving from the model to the context around the work.
For founders and investors the question has changed: not how smart the agent is, but how much work a customer can safely stop doing.