Skip to content
Hands at a laptop and tablet among floating AI agent and app icons

Computer-Use Agents: From Demo to Deployment

If every model can click, clicking is no longer the moat.

Foreword

Computer-use agents went from demo to deployment in about a year.

  • 42% → 85%. The best score on OSWorld-Verified, a benchmark of real desktop work, roughly doubled from early 2025 to 2026, past the 72% or so that a16z cites for humans.
  • No API needed. An agent that can see a screen, click and type can reach the long tail of legacy portals and internal tools that only people could operate.
  • Not yet reliable. On longer, realistic workflows the leading system finished only 20.6% of its tasks, and agents still take several times more steps than they need.

This edition, drawn mainly from a16z's August 2026 field study, looks at:

  • where computer-use agents already work in production;
  • what they cost against offshore and US back-office labor;
  • why the moat is moving from the model to the context around the work.

For founders and investors the question has changed: not how smart the agent is, but how much work a customer can safely stop doing.

First page of the report: Computer-Use Agents: From Demo to Deployment

New reads every month.