Loading market data...

AI Agents Fail Complex Instructions More Than 70% of the Time, New Benchmark Shows

AI Agents Fail Complex Instructions More Than 70% of the Time, New Benchmark Shows

A new benchmark has found that AI agents follow complex instructions less than 30% of the time. The result, which comes from a test designed to measure how well agents handle multi-step tasks, suggests that the technology is far from reliable for anything beyond the simplest commands.

The benchmark's low success rate

The benchmark, whose details have not been fully disclosed, evaluated AI agents on a series of instruction-following scenarios. In cases that involved multiple steps, dependencies, or constraints, the agents succeeded in fewer than three out of ten attempts. That means for every ten complex instructions, the agents got less than three right.

The exact tasks and the criteria for what counts as "complex" remain unclear. But the headline number is stark. It paints a picture of AI agents that can handle straightforward prompts but stumble when the instructions get complicated.

The gap between simple and complex

AI models have shown they can follow basic commands with high accuracy. Ask one to summarize a text or translate a sentence, and it usually delivers. But the benchmark suggests that when instructions require juggling multiple pieces of information, following a sequence of actions, or respecting conditional logic, the performance drops sharply.

This gap isn't surprising to anyone who has tried to get an AI agent to book a trip with specific layover rules or file a report with several formatting requirements. The new benchmark puts a number on that frustration: less than 30% success.

Businesses are already deploying AI agents for customer service, scheduling, data entry, and other tasks that often involve complex user instructions. If agents can't follow those instructions reliably, they'll make mistakes that require human intervention. That defeats the purpose of automation.

The benchmark's results suggest that, for now, AI agents are best suited to narrow, well-defined tasks. Anything that requires careful adherence to a multi-step process is risky. Companies that rely on agents for such work may need to keep humans in the loop or accept a high error rate.

What the benchmark doesn't tell us

The lack of detail about the benchmark is itself a problem. Without knowing the exact tasks, the types of agents tested, or the evaluation method, it's hard to compare results or understand what the 30% figure really means. Is it an average across many different agents? Does it include only text-based agents, or also those that act in simulated environments? Those questions remain unanswered.

For developers, the finding is a reminder that progress on simple tasks doesn't automatically translate to progress on complex ones. More work is needed on instruction comprehension, memory, and planning. Until then, the 30% success rate stands as a sobering checkpoint.

It's not clear when the benchmark's creators will release the full methodology or additional results. Until they do, the number is all we have. And it's not a good one.