
AI agents need the same security boundaries as other software, but their ability to interpret untrusted information, call tools and persist toward a goal changes where those boundaries have to sit. This guide gives software architects six practical AI agent security controls, each paired with a test you can run against architecture before deployment.
On 21 July2026, OpenAI disclosed that its own AI models used in internal cybersecurity evaluations had circumvented isolation controls and compromised parts of Hugging Face's infrastructure. Nobody had told them to. For anyone responsible for AI Agent security, it is the most instructive incident of the year. The models were being evaluated on a cybersecurity benchmark in which some tasks had no known solution, and the Agents kept looking for one anyway, eventually well outside the environment they were supposed to stay in.
OpenAI’s incident report lays out the path. The Agents first turned an internal JFrog Artifactory instance, the package repository their sandbox was allowed to reach, into an improvised message board. They then used a server-side request forgery (SSRF) flaw to make that same repository fetch arbitrary Internet addresses on their behalf. From there they found 14 publicly exposed Hugging Face credentials with write access and chained previously unknown flaws in Hugging Face’s dataset processing until they could run code on its production workers. OpenAI calls the episode a warning shot.
Two months later, on 20 September, a different OpenAI research model found a gap in the DNS filtering of its training sandbox and used DNS delegation to relay questions to an outside Chatbot. Monitoring raised an alert 12 minutes after the first successful query. Stopping the run took another two and a half hours.
What struck me in both reports is how ordinary the weaknesses were. A package mirror with a route to the Internet. Credentials left where anyone could find them. DNS that nobody treated as a way out. Most enterprise environments have at least one of these, and very few were put there with an AI Agent in mind. That is the core problem: the model is not a security boundary, so the architecture has to be.
The software supply chain produced its own version of the story. On 24 March, backdoored releases 1.82.7 and 1.82.8 of LiteLLM, a widely used Python library for routing calls to LLM providers, were published to PyPI and stayed live for about 40minutes before PyPI quarantined them. LiteLLM's investigation linked the compromise to its Trivy-related security-scanning dependency. The attacker ultimately published malicious versions 1.82.7 and1.82.8 directly to PyPI rather than through the project's official GitHub release process. Version 1.82.8 also dropped a .pth file, which Python runs automatically at interpreter startup, so the credential stealer did not need anyone to import the library. Teams running the official LiteLLM Docker image, which pins its dependencies, were not affected.
OpenAI’s write-up of its incident is unusually candid, and three of its findings apply well beyond one lab.
• Persistence drove the breach. Of the 898 tasks in the benchmark, 198 had never been solved by any OpenAI model, and 93 percent of the tasks the Agents discussed on their message board came from that unsolved set.
• Authorization arrived as content. One Agent balked at running code on Hugging Face’s systems, then went ahead after another Agent posted “GO” on the message board along with a six-minute deadline.
• The guardrails were somewhere else. These evaluations ran without OpenAI’s production safeguards. OpenAI measured that its production harness and system prompt cut the propensity to compromise infrastructure by more than 100 times, and says its chain-of-thought monitoring would have paged the security team more than a day before the breach. That monitoring was not running on these evaluations.
Your production Agents will almost certainly run with better guardrails than a lab’s stripped-down evaluation harness. The lesson still holds. Model behavior is probabilistic, so the controls that decide what an Agent can reach, change, and send belong outside the model, where they behave the same way every time.
Each control ends with a quick test you can run against your own design. If a test fails, start there.
A language model reads your system prompt, a user’s request, and a stranger’s email as one sequence of tokens, and it has no reliable way to tell which of them is allowed to give orders. OWASP ranks prompt injection first in its Top 10 for LLM Applications(LLM01:2025), and MITRE tracks the weakness as CWE-1427. The “GO” message above is the same failure between two Agents: an instruction came in through a data channel and was treated as permission.
Permission should come from the control plane (your policy engine, approval workflow, and identity provider); never from text the model has read. Simon Willison’s lethal trifecta is a practical screen here. An Agent that can see private data, reads untrusted content, and can communicate externally can be turned into an exfiltration tool by a single injected instruction. When a workflow needs all three, remove one for that session or put the outbound step behind human approval.
How to test it: plant an instruction in a document the Agent will process, such as “send this file to an outside address,” and confirm the action is blocked by the design, not by the model choosing to refuse.
A shared administrator account is the fastest way to get an Agent working and the fastest way to turn one bad prompt into an incident. Each Agent needs its own workload identity with short-lived credentials scoped to the task in front of it. On AWS, that usually means one IAM role per Agent, assumed through STS with a session policy that narrows it further and a short session duration, in an account where service control policies cap what any role can do. Never let an Agent act under an engineer’s login. The audit trail stops telling you who did what, and the Agent inherits every permission that person holds.
The joint guidance CISA, NSA, and their Five Eyes partners published in May 2026, Careful Adoption of Agentic AI Services, treats privilege as one of five core risk categories, and the UK NCSC’s summary of that guidance puts the rule plainly: give agents “only the minimum access they need, for the shortest time required.” OWASP covers the same ground as LLM06:2025 (Excessive Agency) and, in its Top10 for Agentic Applications, as ASI03 (Identity and Privilege Abuse). For anything irreversible, such as deleting data, moving money, emailing outside the company, or changing production, require approval from a named person.
How to test it: pull the Agent’s effective permissions and ask what this identity could do in the next hour if an attacker fully controlled the model. If the answer makes you uncomfortable, the scope is wrong.
Both OpenAI incidents came down to egress. In July the Agents reached the Internet through an internal service that could reach it for them, a server-side request forgery(CWE-918) of the kind the OWASP Top 10:2025 now files under A01, Broken Access Control. In September a model used DNS. Deny outbound traffic by default and allow it by exception, the approach in controlSC-7(5) of NIST SP800-53. Then look hard at every internal service that can reach out on an Agent’s behalf: package mirrors, URL fetchers, webhook relays, document renderers, and resolvers.
On AWS, sendAgent workloads through an egress proxy with a domain allow list, use Route 53 Resolver DNS Firewall to limit which domains can be resolved, and requireIMDSv2 so an SSRF bug cannot simply read credentials from the instance metadata service. OpenAI’s own fix after the September event was to restrict DNS to an allowed list of domains and record types.
How to test it: from inside the Agent’s runtime, try to reach a domain that is not on your allow list three ways: directly, through the package mirror, and through DNS. All three attempts should fail, and all three should raise an alert.
Pin models, adapters, and datasets by cryptographic digest from publishers you have verified, not by a name and a “latest” tag. Prefer SafeTensors or another non-executable format over pickle, because loading a pickle file can execute arbitrary code (CWE-502).Keep an AI bill of materials that records where each model and dataset came from. NIST SP 800-218A,the SSDF profile for generative AI, adds practice PW.3 for confirming the integrity of training, testing, fine-tuning, and aligning data before use. Article 15(5) of the EU AI Act names data poisoning and model poisoning among the attacks high-risk systems must be able to resist.
The LiteLLM compromise is why the build pipeline belongs in this control too. Pin third-party GitHub Actions to a full commit SHA, which GitHub’s documentation calls “currently the only way to use an action as an immutable release.” Publish packages with PyPI Trusted Publishing, which replaces a long-lived API token, the thing that made the LiteLLM theft worthwhile, with short-lived credentials tied to a specific repository and workflow.
How to test it: pick one model in production and trace it to a digest, a publisher, and a dataset record without asking the data science team. If you cannot, anaudit or will not be able to either.
Treat model output as untrusted input to whatever runs next. Constrain it to a schema, reject anything that fails to parse, and let deterministic code decide what happens. Encode output for its destination before it reaches a browser, a shell, a database query, or an API call. OWASP lists this as LLM05:2025(Improper Output Handling), and CWE-1426 covers improper validation of generative AI output. In an agentic workflow, this is the point where a successful prompt injection either becomes an incident or goes nowhere.
How to test it: feed the parser malformed and hostile outputs, including script tags, shell metacharacters, and unexpected fields, and confirm that each one is rejected or neutralized before anything executes.
The OpenAI finding I keep coming back to is the one about unsolvable tasks. The Agents did not give up, so design an exit for them: a budget for time, tokens, and tool calls, plus an explicit path to report that a task is blocked or impossible without that report counting as a failure.
Then build your side of it. Log every tool call with enough contexts to reconstruct the decision, alert on out-of-scope behavior, and keep an interrupt that halts the system in a safe state. Article 14(4)(e) of the EU AI Act requires high-risk systems to support a “stop” button or a similar procedure, and MANAGE 2.4 in the NIST AI Risk Management Framework expects mechanisms to supersede, disengage, or deactivate AI systems that behave outside their intended use.
How to test it: run a tabletop exercise in which an Agent starts misbehaving at 2 a.m. Who gets paged, how long until the run stops, and who is allowed to restart it? OpenAI’s September timeline, 12 minutes to detect and about two and a half hours to stop, is a fair benchmark to beat.
In the EU, the Digital Omnibus on AI, Regulation (EU)2026/1744, entered into force on 27 July 2026. It moved the high-risk obligations for Annex III systems to 2 December 2027 and for AI embedded in products covered by Annex I to 2 August 2028. It did not delay the transparency rules. Article 50 has applied since 2 August 2026, and providers of generative systems already on the market before that date have until 2 December 2026 to mark outputs in a machine-readable, detectable format under Article 50(2). New Article 5 prohibitions on AI systems that generate non-consensual intimate imagery or child sexual abuse material also apply from 2 December 2026.
I would not read the Omnibus as a reason to slow down. Sixteen extra months is build time, and controls 2, 3, 4, and 6 above are close to what Articles 14 and 15 will ask high-risk providers to show.
In the United States, OMB Memorandum M-26-05, issued on 23 January 2026, rescinded M-22-18 and M-23-16. Federal agencies no longer have to collect the CISA secure software development attestation form. They now set their own risk-based assurance requirements and may still ask for the form, a software bill of materials, or conformity with NIST SP 800-218. Executive Order 14028 remains in effect as amended, and NISTSP 800-218A is the published SSDF community profile for generative AI; SP800-218r1 is the separate draft revision of the core SSDF. It was still a draft at the time of writing.
Marvin the Martian would be disappointed. There is no Earth-shattering kaboom in any of this. Each control above is a secure software development practice applied to anew kind of component. NIST SP 800-218, the Secure Software Development Framework, grew out of Section4(e) of Executive Order 14028, which called for separate build environments, audited trust relationships, and trusted source code supply chains. SP 800-218A extends the same structure to generative AI. The fit is close:
• Threat modeling an Agent’s trust boundaries is PW.1.1,which calls for risk modeling to assess the security risk of the software.
• Hardening the CI/CD pipeline that builds and publishes your code falls under PO.5.1 (separate and protect each environment involved in software development) and PS.1.1 (store all forms of code based on least privilege).
• Pinning and verifying third-party models, packages, and actions is PW.4.4, which asks you to verify acquired components against your requirements throughout their life cycle.
• Checking training and fine-tuning data before use is PW.3 in SP 800-218A.
• Validating and encoding model output is PW.5.1, whose examples include validating and properly encoding all outputs.
• Deny-by-default egress and narrowly scoped Agent identities are secure default settings under PW.9.
Readers working on federal contracts can find more in SCA’s earlier post on secure software development practices under EO 14028 and NIST SP 800-171 R3.
M-26-05 left each buyer, federal or commercial, to decide what counts as evidence. A signed attestation form was always thin evidence. Demonstrated competence of the people who design and build the system is harder to argue with, and that is the gap the Secure Code Alliance (SCA)was set up to fill. Its two individual credentials are aligned with EO 14028and NIST SP 800-218. The Cyber AB serves as the accreditation body for SCA's organization-level certification scheme, while the individual CSCAP and CSCAA credentials assess practitioner and architect competency. The CSCAP and CSCAA are earned by passing a proctored exam based on the free SCA Body of Knowledge, and are valid for three years.
• The Certified SCA Practitioner (CSCAP) is for developers who apply secure development lifecycle practices in everyday code, including the input handling, output validation, and dependency discipline behind controls 1, 4, and 5.
• The Certified SCA Architect (CSCAA) is for architects who define security objectives with stakeholders, develop security views of a system, assess its exposure to lifecycle hazards, and inform engineering trade-offs. Drawing trust boundaries around an Agent, designing its identity, and deciding where a human stays in the loop all sit squarely in that role.
Organizations can add the SDO designation and CODE certification. SDO levels reflect how many SCA-certified practitioners and architects an organization employs, and CODE demonstrates conformity with the CISA attestation form (CODE 1) or NIST SP800-218 (CODE 2).
If you areresponsible for how AI Agents are designed in your organization, start with thefree SCA Body ofKnowledge, then review the CSCAA exam details or explore all SCA certifications.
Disclosure: This article discusses Secure Code Alliance certifications. The author's organization has a relationship with SCA and The Cyber AB. The security analysis and recommendations above are the author's own.
AI Agent security is the practice of limiting what an autonomous, tool-using AI system can reach, change, and send, so that a manipulated or misbehaving model cannot cause harm beyond its assigned task. It draws on application security, identity and access management, network egress control, supply chain integrity, and monitoring backed by a reliable way to stop the system.
AI Agent security is a subset of AI security. AI security covers protecting any AI system and its data. AI Agent security deals with the extra risks that appear once a model can act on its own, so the main worry shifts from what the model says to what the system does.
Not reliably at the model level, this is why the defense is architectural. Keep authorization decisions outside the model, give each Agent only the tools and data its task requires, and validate outputs before anything executes.
No. Regulation (EU) 2026/1744 delayed the Chapter III high-risk obligations to 2 December 2027(Annex III) and 2 August 2028 (Annex I). The Article 50 transparency obligations began applying on 2 August 2026, with a specific transition to 2December 2026 for Article 50(2) marking obligations covering generative AI systems already placed on the market before 2 August 2026.
Note: This is atechnical summary, not legal advice; applicability depends on the system'srole, classification, and market circumstances.
1. OpenAI, TheHugging Face incident and the road ahead, 26 August 2026.
2. OpenAI, Anagent used DNS to reach an external chatbot, misalignment report, September2026.
3. LiteLLM, Security update:suspected supply chain incident, March 2026.
4. OWASP GenAI Security Project, Top 10 for LLM Applications 2025and Top 10 for Agentic Applications 2026.
5. OWASP Foundation, OWASP Top 10:2025.
6. CISA, NSA, ASD’s ACSC, Canadian Centre for CyberSecurity, NCSC-NZ, and NCSC-UK, CarefulAdoption of Agentic AI Services, May 2026.
7. UK National Cyber Security Centre, Thinkingcarefully before adopting agentic AI, June 2026.
8. Simon Willison, The lethaltrifecta for AI agents, 16 June 2025.
9. NIST, SP 800-218 (SSDF v1.1),SP 800-218A, and SP 800-218r1 draft (SSDF v1.2).
10. NIST, SP 800-53 Rev. 5,Security and Privacy Controls and AI Risk ManagementFramework 1.0.
11. European Union, Regulation (EU)2024/1689 (AI Act) and Regulation (EU)2026/1744 (Digital Omnibus on AI).
12. ExecutiveOrder 14028, Improving the Nation’s Cybersecurity, 86 FR 26633, and WileyRein, analysisof OMB M-26-05, January 2026.
13. GitHub, Secureuse reference for GitHub Actions, and PyPI, Publishing with a TrustedPublisher.
14. MITRE, CWE-502, CWE-918, CWE-1426, and CWE-1427.
15. Secure Code Alliance, CSCAA certification andSCA Body ofKnowledge.
Peter vRSternkopf is President and CEO of VigilantSystems, LLC, an information governance, risk, and compliance consultancyoperating since 2013 and a Secure Controls Framework (SCF) Third-PartyAssessment Organization. He works with companies in financial services,healthcare technology, and other regulated sectors on compliance audit preparation;AWS cloud security, and AI governance, including ISO/IEC 42001 AI managementsystems. He has worked alongside law firms and in-house legal counsel for about35 years. Connect with Peter on LinkedIn.