Skip to main content
All insights
AI/LLMPrompt InjectionApplication Security

Prompt Injection in the Wild: What We Learn Breaking LLM Apps

X-Gen Research TeamMay 2, 20268 min read

Why the model isn't the perimeter

When teams ask us to 'test the AI,' they usually mean the model. But a shipped LLM feature is an application: a system prompt, a retrieval layer, a set of tools it can call, and the data and permissions behind each of those. The model is one component, and rarely the weakest one.

Prompt injection matters precisely because the model cannot reliably distinguish instructions it should follow from instructions embedded in the data it processes. Everything the model reads — a web page, a document, a tool result — is potential instruction. That reframing is the whole discipline: treat every token the model ingests as untrusted input crossing a trust boundary.

Direct vs. indirect injection

Direct injection is the obvious case: a user types adversarial text to override the system prompt or extract it. It is worth testing, but well-designed apps increasingly resist the naive versions.

Indirect injection is where real damage lives. The payload is planted in content the model will later read — a support ticket, a calendar invite, a page the agent browses, a document in the knowledge base. The victim is not the person who wrote the payload; it is the user or system that later processes it. This is the class most teams under-test because it requires thinking about data provenance, not just the chat box.

When the agent has hands

The severity of prompt injection scales directly with what the model is allowed to do. A chatbot that only talks has a limited blast radius. An agent that can send email, query databases, execute code, or move money turns a text trick into a privileged action.

This is the excessive-agency problem: tools granted for convenience become an attacker's toolkit the moment injection succeeds. We look hard at what each tool can do, whether invocations are scoped and confirmed, and whether a poisoned instruction can chain tools into something the designer never intended.

Poisoning the well

Retrieval-augmented generation adds a supply chain. If an attacker can influence what lands in the vector store — a public doc that gets indexed, a user-submitted record, a scraped source — they can plant instructions that surface only when a relevant query retrieves them. The injection is dormant until the right question wakes it.

Retrieval poisoning is stealthy because it decouples the attacker from the trigger in both time and identity. Defenses have to consider not just prompts, but everything in the ingestion pipeline: what can be indexed, by whom, and with what review.

Building guardrails that hold

No single control stops prompt injection, so we push clients toward defense in depth: constrain the model's authority with least-privilege tools, keep untrusted content clearly separated from instructions, allowlist and validate tool arguments, and require human confirmation for high-impact actions.

Then monitor. Log tool invocations, watch for anomalous chains, and treat the AI feature like any other privileged system with an audit trail. Mapping the whole thing to the OWASP LLM Top 10 and MITRE ATLAS keeps the assessment honest and comparable across releases, rather than a one-off vibe check.

About X-Gen Research

Field notes from across the X-Generation Cyber Labs practice — drawn from real engagements, sanitized for publication.

Meet the team

Curious about your security posture? We'll identify potential weaknesses and walk you through them. No obligation, no sales theater.