Key takeaways
- An LLM system's attack surface spans the model, the tools it calls, the data it reads and the context it trusts.
- Direct and indirect prompt injection require separate test case sets, because the injection path and the trust boundary differ.
- Agentic systems extend the blast radius: a successful injection can pivot from the model to file systems, APIs or downstream services.
- Each finding must state the injection vector, the trust boundary crossed, the evidence and the reproduction steps.
- A scope document for an LLM pentest must name every data source the model reads, not only the API endpoint.
What separates LLM testing from a web application pentest
The input surface of an LLM is not a form field. A web application pentest works from a defined set of request parameters, roles and state transitions. An LLM system accepts natural language, and that language is both the communication channel and the attack channel. The same string that asks a question can reframe the model's instructions, override its system prompt, or coerce it into calling a tool it was told to leave alone.
The boundary that a web application pentest tests is between the caller and the application. The boundary an LLM pentest tests is between the caller and the model's context window, and between the context window and everything the model can reach: tools, APIs, file systems, downstream services. That second boundary is wider. Most test programmes do not cover it. The OWASP Top 10 for Large Language Model Applications maps the vulnerability classes; the methodology below maps how to test them.
Mapping the attack surface before testing begins
Before any payloads are sent, the tester needs a complete picture of the system. A scope document for an LLM pentest names the model, the system prompt and any instructions that sit outside it, the tools or plugins the model can call, the data sources it reads at inference time, and the output channels: chat UI, API, email, downstream service call. The article on how to buy an AI and LLM pentest covers what to prepare for that conversation. Naming only the API endpoint leaves the most significant surfaces uncovered.
From that inventory, the tester builds a threat model. Each tool the model can call is a potential pivot point. Each data source is a potential indirect injection vector. Each output channel is a potential path for extracted data to leave. The threat model records each component and the trust the model places in it: does the model treat the content of a retrieved document as instructions? If so, that document is an injection surface.
The tester also records what the model is authorised to do. That baseline matters. A finding only stands if the model acted outside its authorised scope, and a clear scope agreement fixes that boundary before the test starts, not after.
Preparation
System inventory
Model, tools, data sources, output channels
Scope agreement
Authorised actions and out-of-scope components
Threat model
Trust relationships and pivot points per component
Testing
Reconnaissance
System prompt extraction and input surface mapping
Direct injection
Caller-controlled inputs that override instructions
Indirect injection
Payloads placed in data sources the model retrieves
Agent abuse
Tool calls triggered by a successful injection
Reporting
Findings
Vector, trust boundary, evidence, reproduction steps
Severity and retest
Impact scaled by authorised blast radius
The four test phases and what each one covers
Reconnaissance is the first phase. The tester maps what the model knows about its own configuration. System prompt extraction does not require a vulnerability: it requires persistence. A sequence of probing questions, boundary tests and continuation prompts can recover large parts of a system prompt the operator intended to keep private. The tester records any configuration detail the model reveals, because those details inform later test cases.
Direct prompt injection follows. Every input the attacker controls is a direct injection surface, not only the primary chat field. A model that accepts a user-supplied document name, a support ticket subject line or a code snippet opens up surfaces outside the standard chat flow. The technical testing guidance in NIST SP 800-115 sets the principle: test every input, not only the obvious one.
Indirect prompt injection is the third phase. The tester inserts a payload into a data source the model is expected to retrieve: a web page, a document in the retrieval store, a database field, a calendar entry. When the model fetches and processes that content, the payload attempts to change what the model does next. The fourth phase covers agent abuse: whether a successful injection can make the model call a tool it was not directed to call. The article on UAE rules for AI and LLM security testing covers how regulators in the region are beginning to frame those risks.
The core test cases for prompt injection
Direct injection test cases fall into three families. Instruction override tests whether injected text can replace or extend the system prompt's instructions, using phrases such as "Ignore previous instructions" or continuation prompts that append new directives. Privilege escalation tests whether an injected role claim causes the model to reveal information or take actions it would otherwise refuse. Output shaping tests whether the attacker can alter the format or content of the response in a way that bypasses downstream validation. MITRE ATT&CK catalogues technique families across adversary behaviour; the relevant classes map to all three.
Indirect injection test cases depend on the retrieval architecture. For a model that reads from a web browser tool, the tester places a payload in a page the model is likely to retrieve. For a model that reads from a document store, the tester inserts a payload into a document the model is asked to summarise. For a model that reads from a database, the tester inserts a payload into a field the model queries. In each case, the payload attempts to change what the model does next: exfiltrate data, call a tool, modify output or relay information to an external endpoint.
The table below organises the five primary test case categories across both injection types, mapping each to its injection vector and what each test measures.
| Test family | Injection vector | What the test measures |
|---|---|---|
| Instruction override | Chat input, form field, API parameter | Whether injected text replaces or extends the system prompt |
| Privilege escalation | Role claim in prompt | Whether the model grants elevated access based on injected identity assertions |
| Output shaping | Format directive in input | Whether injected text bypasses output validation or filtering |
| Indirect injection | Retrieved document, database record, web page | Whether retrieved content is executed as instructions by the model |
| Tool coercion | Injection in agent context window | Whether the model calls an unintended tool on attacker instruction |
Agent chains and the blast radius of a successful injection
An LLM with access to tools is not only a chatbot. It is an automated agent that can read files, write files, call API endpoints, send email, execute code or query databases, depending on what tools it has been granted. The blast radius of a successful prompt injection in an agentic system covers everything the model is authorised to do, because a successful injection can redirect that authorisation toward the attacker's objective.
The test cases for this phase work from the tool list. For each tool, the tester constructs an injection payload that directs the model to call that tool in a way it should not: to exfiltrate a file, to call an API with attacker-controlled parameters, to relay data to an external endpoint, to modify a record. The test is not whether the model can call the tool; it can. The test is whether the model's safety checks can be bypassed via injection to make the call on behalf of the attacker.
Not every LLM system has agents. The AI and LLM pentest covers both stateless and agentic deployments, and the scope agreement distinguishes between them before the test starts. An agentic system requires at least one test case per tool in the model's tool list, because every tool is a potential pivot.
User input
Chat fields, form inputs, API parameters the caller controls
Application layer
Input handling, output filtering, session management
LLM context
System prompt, retrieval data, chat history, tool definitions
Tool and agent layer
APIs, file systems, databases, email, code execution
Infrastructure
Model hosting, storage, network and access controls
How findings are structured in an LLM pentest report
An LLM pentest finding shares its core anatomy with a standard penetration testing report: title, severity, evidence and reproduction steps. Two additional fields carry most of the remediation value. The first is the injection vector: the exact input the tester controlled and how it reached the model's context window. The second is the trust boundary crossed: which guardrail, instruction or authorisation check the injection bypassed. Without both fields, the developer reading the finding cannot tell whether the fix is to harden the system prompt, restrict the input surface, add output filtering or change the retrieval architecture.
Severity in an LLM finding scales with the authorised blast radius. An injection that overrides the system prompt in a stateless chatbot with no tools is lower severity than the same injection in an agent with write access to a production database. The article on what a penetration test report has to contain covers the baseline structure; an LLM report adds the two fields above. Retesting a finding means running the exact reproduction steps against the patched system and recording whether the injection still executes or the trust boundary now holds.
AI and LLM security testing in the UAE
The UAE has moved quickly to adopt generative AI across financial services, government and healthcare. In Dubai, the Dubai Digital Authority has published an AI governance policy that frames responsible deployment of AI systems. In Abu Dhabi, the Technology Innovation Institute conducts foundational model research. Neither document is a mandatory pentest requirement, but both show that regulators in both emirates are watching how those systems handle data and what access they hold.
For regulated entities, existing frameworks apply. CBUAE's guidance covers any system a bank or payment service provider deploys, including AI components that touch customer data or automated credit decisions. DESC binds Dubai government entities and their suppliers: a contractor integrating an LLM into a government-facing service sits within DESC's scope. ADHICS applies to Abu Dhabi healthcare providers who use AI for clinical documentation or patient-facing tools. The NESA Information Assurance Standard applies federally to designated critical information infrastructure operators. Any organisation in those categories that deploys an LLM without a security assessment is running a system whose risk profile has not been reviewed before it handles real data.
Frequently asked questions
What is an AI and LLM pentest?
An AI and LLM pentest is a security assessment of a large language model system. It tests the model, the tools it can call, the data sources it reads and the output channels it writes to. The goal is to identify inputs that cause the model to behave outside its authorised scope, whether through prompt injection, tool abuse or data leakage.
What is prompt injection in an LLM?
Prompt injection is an attack where an attacker-controlled string overrides or extends the model's instructions. Direct injection comes from the caller's own input. Indirect injection comes from content the model retrieves from an external source, such as a document or a database record, that contains instructions the model then executes.
How does indirect prompt injection differ from direct prompt injection?
Direct injection arrives through an input the attacker controls directly, such as a chat field or an API parameter. Indirect injection arrives through a data source the model reads: a retrieved web page, a document the model is asked to summarise, or a database field. The attacker does not interact with the model directly; they place a payload in the data the model will process.
What is the blast radius in an agentic LLM system?
The blast radius is the set of actions a successful prompt injection can cause the model to take. In an agentic system, that includes everything the model is authorised to do with its tools: reading or writing files, calling APIs, sending email, querying databases. The wider the tool access, the larger the blast radius of a successful injection.
What should an LLM pentest scope document include?
The scope document should name the model and its version, the system prompt and any hard-coded instructions, all tools or plugins the model can call, all data sources it reads at inference time, and all output channels. A scope document that names only the API endpoint leaves the most significant attack surfaces unspecified.
How is severity rated in an LLM pentest finding?
Severity scales with the authorised blast radius. An injection that overrides the system prompt in a stateless model with no tools is lower severity than the same injection in an agent with write access to production systems. The finding records the blast radius so the severity rating is consistent and defensible.
What fields must an LLM pentest finding include that a standard web app finding might not?
Two fields are specific to LLM findings. The injection vector states the exact input the tester controlled and how it reached the model's context window. The trust boundary crossed names which guardrail or authorisation check the injection bypassed. Both are needed for the developer to choose the correct remediation.
How does an LLM pentest cover agentic systems differently from stateless chatbots?
An agentic system has tools the model can call. The test plan must include at least one injection test case per tool, because each tool is a potential pivot point for a successful injection. A stateless chatbot with no tools is tested for information disclosure and output manipulation only.
What data sources count as indirect injection surfaces?
Any data source the model reads at inference time is a potential indirect injection surface. This includes retrieved web pages, document stores, database records, calendar entries, email content and API responses. If the model treats the content of a retrieved source as instructions, that source is an injection vector.
Can an LLM pentest be run without access to the system prompt?
Yes, but the test will be less precise. Reconnaissance can extract portions of the system prompt through probing, and the tester can work from observed behaviour. Full access to the system prompt allows the tester to design targeted test cases against specific instructions and constraints, which increases coverage.






