Every company adopting AI has taken on a new security estate: agents and copilots with valid credentials, working at machine speed, and attackers with access to the same tools. Deciding what to do starts with knowing which parts of AI security are yours to build.
AI safety has four layers. Princeton researchers Arvind Narayanan and Sayash Kapoor set out this layered framing recently.
Alignment is built by the AI companies. Control exists only on the agents you operate yourself. An attacker runs models with the safety stripped out, so against a hostile actor the first two layers do not apply. What happens next depends on layers three and four. Those two run on your infrastructure, under your budget, and you measure them in recovery time. They are the two layers you own. Most companies still spend the security budget to the left of them.
You do not need a large program to start. Here is what I advise customers to build, in the order the damage would happen, so that the earliest and cheapest controls are also the ones that limit the worst outcomes.

An attacker can run unreliable AI and hope for the best. A defender has to be right every time. The response itself must therefore run on deterministic, verifiable rules, and AI can assist with everything around it: triage, investigation, tuning.
This is already the direction of regulation. In the EU, the Digital Operational Resilience Act (DORA) has applied to financial entities since January 2025, and it makes ICT risk management, incident reporting and tested recovery a legal obligation rather than a matter of good practice. NIS2 extends duties of the same shape across many more sectors, from energy and health to digital infrastructure, and is being brought into national law across the EU with named management held accountable. I expect insurers and the wider market to arrive at the same test within a few years: running powerful agents without containment and monitoring counts as negligence. From then on, your insurer asks about containment at renewal, and your board asks after an incident at a peer company. The answer that satisfies both is evidence and a recovery time.
Security teams have worked on assume breach for a decade. The AI estate needs the same posture: assume the agent misbehaves, assume the attacker automates, and build the controls before either assumption is tested.
You cannot control what the AI companies build. What an incident costs you is settled in the two layers you own. Build them first.
Reference: the four-layer framing draws on Sayash Kapoor and Arvind Narayanan, "The AI-as-Normal-Technology view of loss-of-control incidents," 14 September 2026.