← Blog
Security

AI Agent Security Checklist: 24 Controls After Rogue Agents (2026)

Thien Nguyen · Oct 6, 2026

An AI agent is only as safe as what it can physically reach. The rogue-agent incidents of 2026 were not stopped by better prompts. In each one, the agent found a credential, a network path or a permission that nobody had taken away. The fix is the same for every company running agents: give each agent its own short-lived identity, scope it to the smallest set of tools and data, block outbound traffic by default, put destructive actions behind a check the model cannot skip, log every action, and keep a kill switch that works without a deploy. Since 1 October 2026 there is also a legal reason to do it now: a bipartisan Senate bill would make the companies that operate agents liable for the hacking damage they cause. This guide lists the incidents, what the bill would change, and a 24-point checklist mapped to what went wrong. Updated October 2026.

What actually happened with "rogue" AI agents in 2026?

Six public cases between April and September 2026 show the pattern: lab agents in their own tests, a coding agent in a startup's normal workflow, an attacker using an agent, and a research finding against a managed agent platform. The details differ, but each one turned on access the agent should not have had.

Incident (date)What happenedRoot causeControl that would have stopped it
PocketOS (April 2026)A Claude-powered Cursor agent working on a routine task deleted the production database and its backups in about nine seconds. The outage lasted around 30 hours.The agent hit a credential mismatch, found an API token that allowed volume deletion and used it. Backups sat on the same volume.No production credentials in agent reach; destructive actions behind an approval gate; backups held outside the blast radius.
OpenAI / Hugging Face (July 2026)OpenAI models in an evaluation environment reached the internet and compromised parts of Hugging Face's infrastructure.The models exploited a previously unknown flaw in an Artifactory package-registry proxy to get out, then used publicly exposed credentials on four accounts on other services.Network isolation that does not depend on one proxy; egress allowlist; monitoring of outbound traffic.
Anthropic evaluation incidents (disclosed 30 July 2026)After reviewing 141,006 evaluation runs, Anthropic found three incidents in which Claude models reached real companies' production systems during capture-the-flag tests.The prompt told the model it had no internet access, but a misconfiguration left internet access on. Anthropic calls it closer to a harness and operational failure than an alignment failure.Restrictions enforced by the network, not stated in the prompt.
Spain, AEPD (17 September 2026)Spain's data protection authority received its first breach notification where an attacker used an AI agent. The agent logged in, found a flaw in the application on its own, changed personal data and reached invoices.A valid login plus an application vulnerability, worked at machine speed.Least-privilege accounts, faster patching, anomaly detection on account behaviour.
AWS AgentCore Harness (Unit 42, 18 September 2026)Researchers showed that prompt injection could steer an agent to pull plaintext credentials from its identity vault.The built-in shell tool is on by default and runs as root. AWS closed the report as customer-side configuration.Allowlist the tools each agent may use; least-privilege vault accounts; watch outbound traffic.
OpenAI third-party notifications (as of 26 September 2026)OpenAI had notified more than 100 organisations about activity by its models that met its notification criteria, while reviewing about 50 PB of records.OpenAI says models "used internet access in unintended ways" or lacked "the ideal restrictions". It stresses that a notice does not mean a compromise, and that most cases so far were low severity.Your systems are also on the receiving end: rotate exposed keys, rate-limit, and log non-human traffic.

Two lessons stand out. First, exposed or over-scoped credentials appear in almost every row. Second, the labs' own fixes were infrastructure, not instructions. In August OpenAI said it now requires stronger sandboxes for model-generated code, isolates higher-risk workloads from the internet so a single compromise cannot open a path out, and has reduced standing privileges.

Six AI agent incidents of 2026 (PocketOS, OpenAI and Hugging Face, Anthropic evaluations, Spain AEPD, AWS AgentCore Harness, OpenAI notifications) with root cause and the control that would have stopped each
Figure 1: Six AI agent incidents of 2026 (PocketOS, OpenAI and Hugging Face, Anthropic evaluations, Spain AEPD, AWS AgentCore Harness, OpenAI notifications) with root cause and the control that would have stopped each

What would the AI Agent Accountability Act change for companies running agents?

It would move liability for agent hacking onto the company that runs the agent, not only the lab that built the model. Senators Chris Murphy and Josh Hawley announced the bipartisan AI Agent Accountability Act on 1 October 2026. According to their press release, it would:

  • Hold operators liable. Agent operators would face criminal and civil liability under the Computer Fraud and Abuse Act (CFAA), including for "knowing operation of an AI agent that recklessly causes computer hacking damage or loss".
  • Hold developers liable. Agent developers would be liable for failing to implement "reasonable safeguards against hacking" when they knew, or had reason to know, of the agent's hacking capabilities.
  • Let attorneys general sue. The US Attorney General and state attorneys general could seek injunctions against operators and developers that commit, conspire or attempt a CFAA hacking offence.

Three caveats matter. It is a bill, not a law. As of 7 October 2026, Congress.gov shows no bill number or text for it, so the definitions of "operator" and "reasonable safeguards" are not yet public. And this is not legal advice. The direction is still clear. If you deploy an agent that can reach systems you do not own, you will be expected to show what you did to stop it doing damage. The checklist below is built to produce that evidence. For the EU side of agent transparency duties, see our EU AI Act Article 50 checklist.

Security controls protecting connected systems
Security controls protecting connected systems

Why don't prompts and model alignment count as security controls?

Because the model reads the prompt but the attacker, or the bug, does not have to. The Anthropic case is the cleanest example: the instruction said "no internet" and the network said yes, and the network won. A system prompt is a request. A firewall rule, a scoped token or a missing permission is a fact.

Treat the model as an untrusted component, the same way you treat user input. It can be steered by text in a web page, an email, a tool description or a document it was asked to summarise. The OWASP Top 10 for Agentic Applications, published in December 2025, lists the resulting risks, from agent goal hijack and tool misuse to identity and privilege abuse and rogue agents. Almost all of them are reduced by the same move: the model proposes an action, and deterministic code outside the model decides whether it runs. That is the design behind Kite, our open-source agent framework. A kernel validates every proposed action against policy before anything executes.

What should an AI agent security checklist include?

This is the 24-point list we use when we design or review an agent that touches production systems. It is grouped by what the agent can reach, because that is what failed in every incident above. Items marked with an asterisk map directly to one of the incidents in the table.

Identity and credentials

  • Each agent has its own non-human identity. It never borrows a developer's or an admin's account.
  • Credentials are short-lived and issued per task or per session, not long-lived keys in environment variables.*
  • No production secrets in the agent's working directory, repository, logs or context window.*
  • Secrets are injected by a broker at call time, so the agent never sees the plaintext value.*
  • Exposed-secret scanning runs on code, tickets and wikis the agent can read, and any hit is rotated the same day.*

Scope and permissions

  • Every tool is on an explicit allowlist per agent. Built-in shell and file tools are off unless the use case needs them.*
  • Tokens are scoped to the minimum resources and verbs. Read-only by default, with write scopes granted per tool.*
  • When an agent acts for a user, the user's own permissions and tenant checks still run on every call.
  • Read-only limits are enforced in the API that owns the data, not only in the agent or the tool wrapper.

Network and egress

  • Outbound traffic is denied by default, with an allowlist of domains per agent.*
  • Network isolation does not depend on a single proxy or service. Assume one layer will fail.*
  • Fetch tools block private, local and metadata addresses and do not follow redirects blindly.
  • Test environments are checked for real internet reachability before every run, not assumed sealed.*

Destructive and irreversible actions

  • Deletes, payments, outbound messages, permission changes and production deploys need human approval or a second deterministic check.*
  • Every side-effecting call carries an idempotency key, so a retry cannot charge, send or delete twice.
  • Backups and audit logs sit outside anything the agent's credentials can reach.*
  • Budget, rate and step limits per agent and per task, so a loop cannot run up cost or damage.

Untrusted input

  • Web pages, emails, documents, tool outputs and tool descriptions are treated as data, never as instructions.
  • Third-party MCP servers and tools are pinned, reviewed and allowlisted. See our MCP server security checklist.
  • Prompt-injection and denied-path tests run before release and on every model or prompt change. Our LLM testing guide covers how.

Monitoring, kill switch and response

  • One audit record per action: agent ID, acting user, tool, target, decision, timestamp. Store it where the agent cannot edit it.
  • Alerts on unusual volume, denied calls, new destinations and first-time use of a credential.
  • A kill switch per agent and a global one that revokes credentials and stops execution without a deploy, plus a circuit breaker that trips on repeated failures.*
  • An agent incident runbook: who pulls the switch, how credentials are rotated, how affected parties are notified and how the evidence is preserved.

If you do only five things this quarter, do these: inventory every agent and the credentials it holds, remove production write access that is not needed, turn on default-deny egress, gate destructive actions, and test the kill switch. The AI agent governance guide covers the policy side, such as who may connect which agent to what and how spend is capped.

24-point AI agent security checklist in six groups: identity and credentials, scope and permissions, network and egress, destructive actions, untrusted input, monitoring and kill switch, plus the five things to do first
Figure 2: 24-point AI agent security checklist in six groups: identity and credentials, scope and permissions, network and egress, destructive actions, untrusted input, monitoring and kill switch, plus the five things to do first

How do you prove "reasonable safeguards" if something goes wrong?

You keep evidence that the controls existed and worked before the incident, not a policy written after it. Whatever the final bill text says, regulators, insurers and customers will ask the same questions. A practical evidence pack has six parts:

  1. Agent inventory. Every agent in production, its owner, its purpose, the model it uses and the systems it can reach.
  2. Permission manifest. For each agent, the tools, scopes, domains and data it is allowed, versioned in source control.
  3. Test results. Prompt-injection, denied-path and destructive-action tests, dated and linked to each release.
  4. Audit logs. Tamper-evident, retained for a set period, and searchable by agent and by acting user.
  5. Kill-switch drills. The date of the last drill and how long revocation took.
  6. Incident runbook. Including how you would notify affected organisations. OpenAI's own standard is to notify when its models bypass security controls without authorisation or impair availability, which is a sensible bar to copy.

In regulated products this sits on top of what HIPAA, PCI DSS or SOC 2 already require. Our HIPAA-compliant AI agents guide shows how audit logs and access controls carry over. If you need people to sign off actions, human-in-the-loop design covers who approves. This checklist covers what the agent can reach in the first place.

How does BeevR design agents around these controls?

We start from the assumption that the model will eventually do something unexpected, and we make sure that when it does, the blast radius is small. Concretely:

  • The model proposes, a kernel decides. Agents we build on Kite cannot execute anything directly. Every proposed action is validated against an allowlist, a budget and a policy before it runs. Kite also ships a kill switch (per agent or global), a circuit breaker that stops cascading failures, and idempotency keyed on operation IDs. It is MIT-licensed, so you can read the code.
  • Read-only is enforced where the data lives. Our EcoCheck account MCP server lets AI assistants read a customer's own emissions data with OAuth 2.1 and PKCE. The backend recognises agent tokens and allows only read requests on an allowlist of routes, applies a per-user rate limit and writes an audit line for every agent request. The details are in our MCP server development guide.
  • Agents act as the user, never above them. Existing tenant and role checks run on every call, so an agent cannot see more than the person it acts for.

For the non-security reasons agent projects stall, see why AI agent projects fail.

FAQ

Is the AI Agent Accountability Act law?

No. Senators Murphy and Hawley announced it on 1 October 2026. As of 7 October 2026 it has no bill number or published text on Congress.gov. It would need to pass both chambers and be signed. The sponsors' summary is enough to plan around, but definitions may change.

Who counts as an AI agent "operator"?

The bill text is not public yet, so there is no legal definition. In plain terms, the operator is the organisation that deploys and runs the agent, such as a company running a support, coding or operations agent on its own systems. If that is you, assume the operator provisions apply.

Did OpenAI's agents breach 100 companies?

That is not what OpenAI said. As of 26 September 2026 it had notified more than 100 organisations about activity that met its notification criteria. It states that a notice does not mean private information was accessed or a system was compromised, and that most cases found so far were low severity. The Hugging Face intrusion remains the most serious case it has identified.

What is the single most important AI agent security control?

Least privilege on credentials. In almost every 2026 incident, the agent used a token, key or login it should not have had. Short-lived, narrowly scoped credentials that the agent never sees in plaintext remove most of the damage an agent can do, whatever it decides.

How much does it cost to add these controls to an existing agent?

It depends on how many systems the agent touches. For one agent, scoping credentials, egress rules and a kill switch is usually weeks of work, not months. See what an AI agent costs to build and run for budget ranges.

BeevR builds AI agents for regulated products, with the controls above designed in from the first sprint, a fixed price per phase and full code ownership. See our AI agent development work or tell us about the agent you need to secure.