LLMs Under Attack

How Prompt Injection Bypasses AI Security Controls

AI, AI Security, LLM Security, how to protect AI agents from prompt injection attacks, enterprise strategies for LLM security, prompt injection risks for large language models, Prompt Injection, Prompt
Facebook
X
LinkedIn
Reddit
WhatsApp

As companies increasingly deploy AI agents, their exposure to cyberattacks is also growing. The focus is shifting beyond traditional IT systems toward the AI models themselves, their prompts, and the data flows they rely on.

The latest Top 50 Cybersecurity Threats Report from Splunk, a Cisco company, shows that attacks targeting AI systems are becoming significantly more relevant. Emerging risk areas such as backdoor injection, jailbreaking, model extraction, and model poisoning are already affecting enterprise cybersecurity strategies today.

Ad

Many of these attacks directly target large language models (LLMs) and their inputs. Prompt based manipulation can bypass security safeguards, trigger harmful outputs, or expose sensitive information. In particular, prompt injection is evolving into a major security challenge as the adoption of agentic AI continues to accelerate.

When AI Follows the Wrong Instructions

How do cybercriminals exploit prompt injection? With this attack method, adversaries attempt to deliberately manipulate large language models by inserting hidden or additional instructions into prompts, documents, or other content processed by the model. If the AI system fails to identify these instructions as malicious, it interprets them as legitimate commands. As a result, it may ignore previous guidelines and perform unintended actions.

In general, two main attack categories can be distinguished:

Ad

1. Direct prompt injection targets the AI model itself. Attackers disguise malicious commands as legitimate prompt instructions in an attempt to bypass security controls and predefined rules. The objective is to make the system follow instructions that should normally be blocked. As AI systems become more powerful and autonomous, the potential impact of such manipulations on applications and business processes increases.

2. Indirect prompt injection embeds malicious instructions into external content such as emails, documents, or websites that are later analyzed by the language model. The system then interprets manipulated instructions as part of the legitimate context, potentially exposing internal data, producing inaccurate results, or triggering unwanted actions. This attack variant is particularly difficult to detect because the malicious commands are hidden inside external sources.

Prompt Injection as a Gateway to Follow Up Attacks

Prompt injection is not an isolated threat. It can serve as the starting point for a broader chain of attacks against AI systems and agentic applications. The Splunk report highlights that prompt based manipulation plays a key role in threats such as jailbreaking and content abuse. Attackers can force models to bypass safeguards, generate harmful content, or ignore internal policies.

At the same time, these manipulations can enable additional attack scenarios, for example when AI models disclose confidential information or when hidden instructions are introduced into training or fine tuning data. In such cases, the injected commands can later function like a digital backdoor.

Preventing Attacks Through Controls and Access Restrictions

Organizations should not rely solely on the built in security mechanisms of language models when defending against prompt injection. Instead, they need to consistently evaluate inputs and external content before these sources enter sensitive AI workflows.

Key measures include clearly defined rules for validating prompts and contextual data, continuous monitoring of model inputs and outputs for suspicious patterns, and additional safeguards for high risk AI applications. Especially when AI systems process information from emails, documents, or the web, separating trusted data sources from potentially manipulated content is essential.

AI applications should also be clearly isolated from critical systems and sensitive data whenever possible. Strict access controls, human reviews for sensitive processes, stronger segmentation of AI applications, and continuous testing against known attack techniques can significantly reduce risk.

For organizations, this means that protecting against prompt injection is not limited to securing the language model itself. Security must already be considered in the architecture of the application in which the model operates. This is particularly important for AI agents that independently access information or initiate actions. These systems require additional control layers, clearly defined access boundaries, and security architectures designed to limit potential failures.

Addressing Prompt Injection From the Start

As AI agents become more widespread, preventing prompt injection will become a security priority that organizations must incorporate into their AI strategies from the beginning. The deeper language models become embedded into operational workflows, data access processes, and automated decision making, the greater the potential impact of targeted manipulation.

Prompt injection demonstrates that AI security cannot focus solely on the model itself. Companies need a comprehensive approach that considers prompts, data sources, system permissions, and downstream processes as interconnected security factors.

Only by addressing these elements from the outset can organizations take advantage of agentic AI while avoiding new vulnerabilities in security critical environments.

Philipp Behre Splunk

Philipp

Behre

Field CTO & Strategic Advisor

Splunk

Ad

Artikel zu diesem Thema

Weitere Artikel