Loading market data...

Computer-Use Agents Leap to 85% on OSWorld, From 12% Two Years Ago

Computer-Use Agents Leap to 85% on OSWorld, From 12% Two Years Ago

Computer-use agents, the AI systems that operate software by mimicking human clicks and keystrokes, have made a striking leap on a key benchmark. They now score 85% on OSWorld, a standard test of how well agents handle real computer tasks, up from just 12% two years ago. The jump points to a near-term future where structured, repetitive workflows could be handed off to machines.

What OSWorld Measures

OSWorld drops agents into a desktop environment and asks them to complete tasks like filling out forms, navigating menus, and using applications. It's designed to reflect the messy reality of everyday computing, not just clean, isolated tests. The score represents the percentage of tasks completed successfully.

A Two-Year Leap

Two years ago, agents could barely get off the ground, completing only 12% of tasks. Today, they're at 85%. That's a more than sevenfold improvement, and it's not just incremental progress. The rapid climb suggests that the underlying models and training methods have gotten dramatically better at understanding and interacting with graphical interfaces.

The Promise of Automation

The improvement has real-world implications. Structured workflows — data entry, invoice processing, software configuration, even basic customer support — are the kinds of tasks that follow predictable steps. If agents can handle those reliably, businesses could automate a significant chunk of their back-office work. The benchmark's results suggest that's no longer a distant possibility.

The Limits That Remain

But the benchmark also has a ceiling. Complex tasks — ones that require reasoning, adaptation, or handling unexpected errors — still trip up the agents. The 85% score doesn't mean agents are ready for every scenario. It means they're good at the structured stuff, but the messy, unpredictable parts of computing remain a challenge. The gap between the two is where the next round of research will focus.

The next step for the field is clear: move beyond the structured tasks that are now within reach and tackle the complex, open-ended problems that still defeat the agents. Until then, the benchmark's score is a measure of progress, not a promise of full autonomy.