AI safety firm Apollo Research is pushing for a fundamental shift in how artificial intelligence systems are evaluated: instead of waiting until a model is finished and ready for release, the company argues that evaluators should be embedded directly inside the lab during the training process itself.
The proposal, which Apollo Research has been advocating, could change how AI companies approach safety testing and potentially influence future regulations. Right now, most evaluation of AI models happens after training is complete, often just before a public launch. Apollo Research wants that to change.
Why post-training evaluations fall short
When evaluators only step in after a model is fully trained, they're essentially inspecting a finished product. The problem is that by that point, many of the model's behaviors and capabilities are already baked in. Retraining or modifying a finished model is expensive and time-consuming, and some issues might be too deeply embedded to fix easily.
Embedding evaluators during training would let them catch potential problems as they emerge. That could mean identifying dangerous capabilities early, before they're fully developed. It could also give labs a chance to adjust the training process on the fly, rather than discovering a flaw only after millions of dollars and months of work have already been spent.
Apollo Research hasn't laid out a specific technical blueprint in its public advocacy, but the core idea is straightforward: make safety evaluation a continuous, integrated part of training rather than a final checkpoint.
A shift that could reshape industry norms
If labs adopt this approach, it could become a new standard for how AI companies operate. Right now, there's no universal requirement for when or how AI models get evaluated. Some companies do internal testing throughout development, but there's no industry-wide norm, and regulators have mostly focused on post-deployment oversight.
Putting evaluators inside the lab during training would create a very different dynamic. It would give safety teams more influence over the development process, and it would generate a continuous record of how a model behaves as it learns. That kind of transparency could be valuable not just for internal safety but also for external audits and regulatory reviews.
The idea also raises practical questions. Who would these embedded evaluators be? Would they be employees of the AI lab, independent contractors, or government regulators? How much access would they have to proprietary training data and algorithms? And what happens if an evaluator flags a problem that the lab disagrees with?
Apollo Research hasn't answered all of those questions, but its advocacy sets a direction. The company is essentially arguing that safety can't be an afterthought — it has to be part of the process from the start.
Regulatory implications remain uncertain
Regulators in the US and Europe have been debating how to oversee AI development, but most proposals focus on transparency requirements, risk assessments, and post-market monitoring. Embedding evaluators during training would go further, inserting oversight into the development phase itself.
That could be a harder sell politically and operationally. AI labs guard their training processes closely, and many are reluctant to share details about their models before release. Bringing in outside evaluators during training would require a level of openness that few companies have shown so far.
Still, Apollo Research's push adds a concrete proposal to a debate that's often abstract. Instead of just calling for "more safety," the company is pointing to a specific structural change: put evaluators where the training happens, not just where the product ships.
Whether labs and regulators pick up the idea is another matter. There's no timeline for when — or if — embedded evaluation might become standard practice. For now, it's a recommendation from one safety-focused firm, and the industry hasn't responded publicly. The next step would likely be for Apollo Research to publish more detailed technical guidance or for a major lab to pilot the approach. Until then, the proposal sits as a challenge to the status quo: evaluate early, evaluate often, and don't wait for launch day.



