Loading market data...

CertiK Wants AI Agents Treated as Staff, Not Software

CertiK Wants AI Agents Treated as Staff, Not Software

CertiK published a report this week arguing that AI agents doing security and financial-crime work should be managed like employees rather than tools. The report, called Intel3D, looks at how agents now reason through steps, pull in evidence, act in live environments, and hand results back for review with limited human input. CertiK's point isn't that the technology is coming. It's that it's already here, and most companies aren't set up to supervise it.

From flagging problems to fixing them

The shift is in what these systems actually do. Older machine-learning tools mostly helped analysts by flagging unusual logins, scoring transactions, and drafting reports — the human still made the call and did the work. Newer agents go further. CertiK describes an agent in a security operations center that gets a suspicious login alert, checks device records, location data, and threat feeds, then suspends the account and logs every step it took. A person reviews the trail afterward.

That pattern is spreading across Web3 security, according to the report — contract triage, transaction risk scoring, tracing stolen funds. The common thread is speed. Flash-loan exploits can drain a protocol in seconds, and stolen crypto can move across bridges and mixers within hours. Human teams often don't have time to investigate before the money is gone.

The staffing problem underneath

There's a less glamorous reason this is happening: not enough people. Cybersecurity and compliance roles are hard to fill, and regulatory pressure keeps adding to the workload. CertiK cites more than $900 million in AML penalties in the first half of 2025 — a figure that makes the case for automation on its own, without any help from the technology's advocates.

In finance, the report sees agents taking on AML tasks. In crypto, it points to exploit detection, fund tracing, and transaction monitoring. None of that is hypothetical at this point.

The failure modes nobody wants to talk about

CertiK is direct about the risks. Agents produce incorrect outputs. Humans, lulled by a system that usually works, scrutinize those outputs less than they should. And AI systems themselves are targets — an attacker who can feed an agent bad data or manipulate its inputs is attacking the investigation, not just the account.

The report's warning lands hardest here: without strong oversight, automation doesn't eliminate old problems. It replaces them with failures that are harder to explain.

What CertiK says to do about it

The recommendations are unglamorous and specific. Keep records of inputs and actions. Set clear limits on what an agent can decide on its own. Run adversarial testing against the agents themselves. And name a human owner for every agent — someone accountable when it goes wrong.

Treating agents as workforce participants rather than ordinary software is the framing CertiK keeps returning to. That means companies remain responsible for outcomes, even when a machine did the work. It's a management problem dressed up as a technology problem, which may be why it's getting ignored. Buying an agent is easy. Assigning it a supervisor, a scope, and a paper trail is not.

The report doesn't set a deadline or propose a standard. For now, the practical question sits with each security team: who owns the agent, and what happens when it's wrong?