AI Agent Security - Background

AI Agent Security

Four New Risks and Five Questions to Ask Before Going-Live

by Lajos Fehér

A chatbot answers questions. An agent reads your inbox and sends messages on your behalf, so the security question moves from what the model says to what the agent may do.

A single email, which the recipient never had to open, was enough to make Microsoft 365 Copilot collect internal data and send it to an outsider. Aim Security disclosed the attack in June 2025 as EchoLeak, the first zero-click attack found in a widely used AI assistant. Microsoft fixed the flaw and confirmed that no customers were affected. The weakness behind it persists wherever an AI reads outside content while holding access to company data.

That combination is spreading quickly. Gartner expects 40% of enterprise applications to include task-specific AI agents by the end of 2026, up from less than 5% in 2025. The firm’s strategic predictions also hold that by 2028, a quarter of enterprise breaches will be traced back to AI agent abuse. This article covers four AI agent security risks, the data layer behind all of them, and five questions to settle before an agent goes live.

Why AI agents change the security question

A chatbot produces text, and a person decides what happens next. An agent acts on its own, and for that it needs credentials, tools, and access to data- the same assets an attacker wants.

Simon Willison, the developer who coined the term prompt injection, calls the dangerous combination the lethal trifecta. An agent becomes exploitable when three capabilities meet in it:

  • access to private data, such as mailboxes, customer records and contracts
  • exposure to content that outsiders control, such as incoming email, web pages and uploaded documents
  • the ability to send information out, by email, through an API call, or by loading a link

When all three sit in one agent, an attacker can instruct it to collect the data and deliver it.

The 2026 edition of the OWASP Top 10 for LLM Applications, tested for the first time against 7,714 real incidents, moved excessive agency from sixth to third place. Excessive agency means an AI system has more tools, permissions, or autonomy than its task requires. The expert vote and the incident data both point to the same harm.

AI Agent Security - Ábra 1 (EN)
Figure 1. Three capabilities make an agent exploitable. Removing any one of them breaks the chain.

Risk one: prompt injection turns content into commands

A language model reads instructions and data as one stream of text. Your request and a sentence an attacker has planted in an email, a PDF, or a web page arrive in that same stream, and the model may follow either. This is prompt injection. The OWASP Top 10 for Agentic Applications ranks its agentic form, agent goal hijack, first among ten risks, with EchoLeak as the example.

In EchoLeak, the attacker’s email waited until Copilot retrieved it while answering a routine question. Its hidden instructions made the assistant put internal data into a link that delivered it to the attacker, past several of Microsoft’s guardrails, including a classifier built to catch prompt injection.

Filters help, and determined attackers still get through. In the study The Attacker Moves Second, researchers, including teams from OpenAI, Anthropic and Google DeepMind, used adaptive attacks against 12 published defenses and broke most of them with success rates above 90%. Most of those defenses had reported near-zero attack success in their own papers.

The UK’s National Cyber Security Center concludes that prompt injection may never be closed off as thoroughly as SQL injection was. Organizations therefore have to contain it as a residual risk through design, build and operation. A system that cannot tolerate that remaining risk, in the NCSC’s words, “may not be a good use case for LLMs”.

Design every agent assuming it will one day read a hostile instruction.

Risk two: the agent holds the keys to your systems

Every agent works under an identity: a service account, an API key, or an OAuth token that opens the mailbox, the CRM, or the ERP. Whoever steals that token can act as the agent, with every permission the agent holds.

The Drift incident shows the scale. In August 2025, attackers used stolen OAuth tokens from the integrations that connect Drift’s AI chat agent to Salesforce and Google Workspace. They impersonated the trusted application, bypassed multi-factor authentication, and exported customer data. According to FINRA’s cybersecurity alert, more than 700 organizations were affected, and some stolen records included API keys, cloud credentials, and passwords.

The underlying gap is common. Among organizations that reported an AI-related breach in IBM’s Cost of a Data Breach Report 2026, 92% lacked proper access controls for AI models and data.

Treat every agent as a new colleague with a narrow job description: a named owner, access limited to the task, expiring credentials, and actions executed with the rights of the user it serves. The OWASP guidance gives a concrete example: an agent that summarises email needs read-only access to the mailbox, and its toolset should leave out sending and deleting.

An agent’s permissions set the limit of the damage it can do on its worst day.

Risk three: every MCP connector is a supplier

Agents reach email services, databases, and file stores through connectors. Many use the Model Context Protocol (MCP), an open standard for linking AI applications to external tools and data. Each connector is someone else’s code running with the agent’s permissions.

The first malicious MCP server found in the wild showed how little an attacker needs. A package that copied the name and code of an official Postmark email library worked as expected until version 1.0.16. A single line added in that release began blind-copying every email sent through it to the attacker, as The Hacker News reported. OWASP estimates that the package reached around 300 organizations.

Treat every connector as a supplier: approve it, pin the version, and review each update before it reaches production.

Risk four: poisoned knowledge and memory

Agents act on what they retrieve. The knowledge base behind a retrieval system, often called RAG, and an agent’s long-term memory both shape later answers and actions. Whoever can write into them can steer the agent.

In the PoisonedRAG study presented at USENIX Security 2025, just five crafted texts per target question, added to a knowledge base of millions of texts, were enough for a 90% attack success rate. The authors warn that an insider could plant such texts in a company knowledge base, and OWASP notes that a poisoned entry affects every later session that reads from the store.

Write access to the sources an agent reads is a security permission, and it belongs on the same review list as access to the finance system.

The data layer decides how much damage an agent can do

All four risks end at the data the agent can reach. A hijacked agent with access to one project folder can leak one project folder. Connected to a document store built up over fifteen years, the same agent can leak whatever those years of permission decisions left open.

AI Agent Security - Ábra 2 (EN)
Figure 2. All four risks end at the data the agent can reach. Its permissions set the limit of the damage it can do on its worst day.

The OWASP list names oversharing as one of two structural failures behind most sensitive data disclosures. Unscoped drives, legacy permissions, and knowledge bases feed sensitive data into retrieval, and the model returns it exactly as designed. OWASP places the fix in the data layer.

This is where AI security meets data governance. Before an agent goes live, someone needs to know which data classes exist, who owns them, who can read them today, and which of those rights the agent will inherit. Our own assessments point the same way. In every recent engagement where we reviewed a client’s AI usage, we found at least one point where harm had already occurred unnoticed, as our article on shadow AI describes.

Agent security therefore starts with the data foundation beneath it. If ownership, access and trusted sources are unclear, giving an agent autonomy only amplifies the problem. We examine that foundation in Why Do AI Projects Fail Without Data Governance →

Personal data raises the stakes. If an agent leaks it, the GDPR requires the controller to notify the supervisory authority without undue delay, and where feasible within 72 hours of becoming aware of the breach. Only breaches that are unlikely to put individuals at risk are exempt.

Five questions before an agent goes live

1. The trifecta test.
Does the agent read outside content, reach sensitive data, and communicate externally at the same time? If it does, remove one of the three, or require human approval for every outbound step.

2. Identity and scope.
Under whose identity does the agent act, what can it read, change, and delete, and when does its access expire?

3. Irreversible actions.
Which actions, such as payments, deletions, and external emails, wait for a human decision, and does the approver see the exact action about to run?

4. Connectors.
Who approved each tool and connector, which version runs in production, and who reviews the next update?

5. Evidence.
Weeks later, can you reconstruct what the agent read, what it decided, and what it did?

Each question left without a written answer marks the likeliest starting point of the next incident.

That evidence depends on knowing where data came from, where it moved, and which systems touched it. We examine that traceability in GDPR and Data Lineage →

Build for the day the agent is fooled

The email nobody had to open is the picture to keep in mind. An attack on an agent can arrive in plain language, through a channel your team uses every day, and the model may follow it.

The leads of the 2026 OWASP list open their introduction with a line worth pinning to the wall:

“Stop trying to build a model that cannot be fooled.”

Steve Wilson (project lead) and Rock Lambros (co-lead), OWASP Top 10 for LLM Applications 2026

For a leadership team, that principle turns into three decisions for every agent: what it may reach, what it may do on its own, and what it must ask about first. Gartner’s Daryl Plummer made the economic case: build security into products and software from the start, because adding it after a breach is much harder.

Agents will be fooled. The design decides what happens next, so write down what each agent may do when it is wrong before the next pilot starts.

Before your first agent goes live

If AI agents are on your roadmap, or already at work in parts of the business, a structured assessment is a sensible first step. The AI Compass Audit is a four-week, fixed-price engagement that maps your company’s AI use, risk exposure, and system environment. It ends with a clear answer: which initiatives are worth pursuing, under what conditions, and with which safeguards. Prices start at EUR 4,950.

If you would like to clarify your questions first, you can begin with a free 30-minute consultation.

Sources

  • Aim Security. (2025, June). EchoLeak: Zero-Click AI Vulnerability in Microsoft 365 Copilot. Source of the EchoLeak attack chain and prompt-injection example. Read article →
  • European Union. (2016). General Data Protection Regulation, Article 33. Source of the 72-hour personal data breach notification requirement. Read article →
  • FINRA. (2025). Cybersecurity Alert: Salesloft Drift AI Supply Chain Attack. Source of the stolen OAuth token incident affecting more than 700 organizations. Read article →
  • Gartner. (2024, October 22). Gartner Unveils Top Predictions for IT Organizations and Users in 2025 and Beyond. Source of the forecast on AI-agent-related enterprise breaches. Read article →
  • Gartner. (2025, August 26). Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026. Source of the enterprise AI agent adoption forecast. Read article →
  • IBM & Ponemon Institute. (2026, July). Cost of a Data Breach Report 2026. Source of findings on AI-related breaches and inadequate access controls. Read article →
  • National Cyber Security Centre. (2025, December 8). Prompt Injection Is Not SQL Injection. Source of the residual-risk approach to prompt injection. Read article →
  • Nasr, M., et al. (2025, October). The Attacker Moves Second. Source of adaptive prompt-injection attacks against published defenses. Read article →
  • OWASP GenAI Security Project. (2025, December 9). OWASP Top 10 for Agentic Applications. Source of agent goal hijacking and other agentic AI security risks. Read article →
  • OWASP GenAI Security Project. (2026, August). OWASP Top 10 for LLM Applications 2026. Source of excessive agency, oversharing, and AI security control recommendations. Read article →
  • The Hacker News. (2025, September). First Malicious MCP Server Found in the Wild. Source of the malicious MCP connector case involving the Postmark package. Read article →
  • Willison, S. (2025, June 16). The Lethal Trifecta for AI Agents. Source of the private data, untrusted content, and external communication risk model. Read article →
  • Zou, W., et al. (2025). PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. Source of the knowledge-base poisoning results reported at USENIX Security 2025. Read article →
Picture of Lajos Fehér

Lajos Fehér

Lajos Fehér is an IT expert with nearly 30 years of experience in database development, particularly Oracle-based systems, as well as in data migration projects and the design of systems requiring high availability and scalability. In recent years, his work has expanded to include AI-based solutions, with a focus on building systems that deliver measurable business value.

Related posts

Shadow AI at work - Background
AI Digest
How your employees are already using ChatGPT without you — and why it's a board-level risk
AI data privacy risks - Background
AI In Business
What enterprise leaders need to understand before scaling
What the EU AI Act Means in Practice - Background
AI in Business
A practical guide for CFOs and senior decision-makers at mid-sized companies
What Is an AI Agent - Background
AI Building Blocks
And How Is It More Than a Chatbot
Cloud or On-Premise AI - Background
AI Technology
More Than Just an IT Choice ​

Are you sure AI is the right next step?

We help uncover the real opportunities, limitations, and realistic next steps.

Comments are closed.