Home
VAPT Web Application PentestAPI PentestMobile App PentestInfrastructure PentestAI & LLM PentestOT / ICS PentestIoT PentestPenetration TestingAll VAPT services
Red Team Red Team EngagementAdversary SimulationAssumed BreachPurple TeamingSocial Engineering
CompanyResourcesBlogFree Consultation

How an AI and LLM pentest is run: the phases, test cases and what the report has to show

An AI and LLM pentest follows a different sequence than a web application test. This article maps the phases, the core test cases and what a report has to record.

Joel Aviad OssiJoel Aviad OssiRed Team Lead, RedTeam Security
7 min read
An isometric row of server nodes connected by glowing red flow lines, with one central node branching outward to a document database, a tool rack and an output channel, showing the components of an LLM system.

Key takeaways

  • An LLM system's attack surface spans the model, the tools it calls, the data it reads and the context it trusts.
  • Direct and indirect prompt injection require separate test case sets, because the injection path and the trust boundary differ.
  • Agentic systems extend the blast radius: a successful injection can pivot from the model to file systems, APIs or downstream services.
  • Each finding must state the injection vector, the trust boundary crossed, the evidence and the reproduction steps.
  • A scope document for an LLM pentest must name every data source the model reads, not only the API endpoint.

What separates LLM testing from a web application pentest

The input surface of an LLM is not a form field. A web application pentest works from a defined set of request parameters, roles and state transitions. An LLM system accepts natural language, and that language is both the communication channel and the attack channel. The same string that asks a question can reframe the model's instructions, override its system prompt, or coerce it into calling a tool it was told to leave alone.

The boundary that a web application pentest tests is between the caller and the application. The boundary an LLM pentest tests is between the caller and the model's context window, and between the context window and everything the model can reach: tools, APIs, file systems, downstream services. That second boundary is wider. Most test programmes do not cover it. The OWASP Top 10 for Large Language Model Applications maps the vulnerability classes; the methodology below maps how to test them.

Mapping the attack surface before testing begins

Before any payloads are sent, the tester needs a complete picture of the system. A scope document for an LLM pentest names the model, the system prompt and any instructions that sit outside it, the tools or plugins the model can call, the data sources it reads at inference time, and the output channels: chat UI, API, email, downstream service call. The article on how to buy an AI and LLM pentest covers what to prepare for that conversation. Naming only the API endpoint leaves the most significant surfaces uncovered.

From that inventory, the tester builds a threat model. Each tool the model can call is a potential pivot point. Each data source is a potential indirect injection vector. Each output channel is a potential path for extracted data to leave. The threat model records each component and the trust the model places in it: does the model treat the content of a retrieved document as instructions? If so, that document is an injection surface.

The tester also records what the model is authorised to do. That baseline matters. A finding only stands if the model acted outside its authorised scope, and a clear scope agreement fixes that boundary before the test starts, not after.

The three phases of an LLM pentest and what the tester covers inside each.

The four test phases and what each one covers

Reconnaissance is the first phase. The tester maps what the model knows about its own configuration. System prompt extraction does not require a vulnerability: it requires persistence. A sequence of probing questions, boundary tests and continuation prompts can recover large parts of a system prompt the operator intended to keep private. The tester records any configuration detail the model reveals, because those details inform later test cases.

Direct prompt injection follows. Every input the attacker controls is a direct injection surface, not only the primary chat field. A model that accepts a user-supplied document name, a support ticket subject line or a code snippet opens up surfaces outside the standard chat flow. The technical testing guidance in NIST SP 800-115 sets the principle: test every input, not only the obvious one.

Indirect prompt injection is the third phase. The tester inserts a payload into a data source the model is expected to retrieve: a web page, a document in the retrieval store, a database field, a calendar entry. When the model fetches and processes that content, the payload attempts to change what the model does next. The fourth phase covers agent abuse: whether a successful injection can make the model call a tool it was not directed to call. The article on UAE rules for AI and LLM security testing covers how regulators in the region are beginning to frame those risks.

The core test cases for prompt injection

Direct injection test cases fall into three families. Instruction override tests whether injected text can replace or extend the system prompt's instructions, using phrases such as "Ignore previous instructions" or continuation prompts that append new directives. Privilege escalation tests whether an injected role claim causes the model to reveal information or take actions it would otherwise refuse. Output shaping tests whether the attacker can alter the format or content of the response in a way that bypasses downstream validation. MITRE ATT&CK catalogues technique families across adversary behaviour; the relevant classes map to all three.

Indirect injection test cases depend on the retrieval architecture. For a model that reads from a web browser tool, the tester places a payload in a page the model is likely to retrieve. For a model that reads from a document store, the tester inserts a payload into a document the model is asked to summarise. For a model that reads from a database, the tester inserts a payload into a field the model queries. In each case, the payload attempts to change what the model does next: exfiltrate data, call a tool, modify output or relay information to an external endpoint.

The table below organises the five primary test case categories across both injection types, mapping each to its injection vector and what each test measures.

Core prompt injection test case categories, mapped to injection vector and what each test measures.
Test familyInjection vectorWhat the test measures
Instruction overrideChat input, form field, API parameterWhether injected text replaces or extends the system prompt
Privilege escalationRole claim in promptWhether the model grants elevated access based on injected identity assertions
Output shapingFormat directive in inputWhether injected text bypasses output validation or filtering
Indirect injectionRetrieved document, database record, web pageWhether retrieved content is executed as instructions by the model
Tool coercionInjection in agent context windowWhether the model calls an unintended tool on attacker instruction

Agent chains and the blast radius of a successful injection

An LLM with access to tools is not only a chatbot. It is an automated agent that can read files, write files, call API endpoints, send email, execute code or query databases, depending on what tools it has been granted. The blast radius of a successful prompt injection in an agentic system covers everything the model is authorised to do, because a successful injection can redirect that authorisation toward the attacker's objective.

The test cases for this phase work from the tool list. For each tool, the tester constructs an injection payload that directs the model to call that tool in a way it should not: to exfiltrate a file, to call an API with attacker-controlled parameters, to relay data to an external endpoint, to modify a record. The test is not whether the model can call the tool; it can. The test is whether the model's safety checks can be bypassed via injection to make the call on behalf of the attacker.

Not every LLM system has agents. The AI and LLM pentest covers both stateless and agentic deployments, and the scope agreement distinguishes between them before the test starts. An agentic system requires at least one test case per tool in the model's tool list, because every tool is a potential pivot.

The five layers of an LLM system a pentest must cover, from caller input down to infrastructure.

How findings are structured in an LLM pentest report

An LLM pentest finding shares its core anatomy with a standard penetration testing report: title, severity, evidence and reproduction steps. Two additional fields carry most of the remediation value. The first is the injection vector: the exact input the tester controlled and how it reached the model's context window. The second is the trust boundary crossed: which guardrail, instruction or authorisation check the injection bypassed. Without both fields, the developer reading the finding cannot tell whether the fix is to harden the system prompt, restrict the input surface, add output filtering or change the retrieval architecture.

Severity in an LLM finding scales with the authorised blast radius. An injection that overrides the system prompt in a stateless chatbot with no tools is lower severity than the same injection in an agent with write access to a production database. The article on what a penetration test report has to contain covers the baseline structure; an LLM report adds the two fields above. Retesting a finding means running the exact reproduction steps against the patched system and recording whether the injection still executes or the trust boundary now holds.

AI and LLM security testing in the UAE

The UAE has moved quickly to adopt generative AI across financial services, government and healthcare. In Dubai, the Dubai Digital Authority has published an AI governance policy that frames responsible deployment of AI systems. In Abu Dhabi, the Technology Innovation Institute conducts foundational model research. Neither document is a mandatory pentest requirement, but both show that regulators in both emirates are watching how those systems handle data and what access they hold.

For regulated entities, existing frameworks apply. CBUAE's guidance covers any system a bank or payment service provider deploys, including AI components that touch customer data or automated credit decisions. DESC binds Dubai government entities and their suppliers: a contractor integrating an LLM into a government-facing service sits within DESC's scope. ADHICS applies to Abu Dhabi healthcare providers who use AI for clinical documentation or patient-facing tools. The NESA Information Assurance Standard applies federally to designated critical information infrastructure operators. Any organisation in those categories that deploys an LLM without a security assessment is running a system whose risk profile has not been reviewed before it handles real data.

Frequently asked questions

What is an AI and LLM pentest?

An AI and LLM pentest is a security assessment of a large language model system. It tests the model, the tools it can call, the data sources it reads and the output channels it writes to. The goal is to identify inputs that cause the model to behave outside its authorised scope, whether through prompt injection, tool abuse or data leakage.

What is prompt injection in an LLM?

Prompt injection is an attack where an attacker-controlled string overrides or extends the model's instructions. Direct injection comes from the caller's own input. Indirect injection comes from content the model retrieves from an external source, such as a document or a database record, that contains instructions the model then executes.

How does indirect prompt injection differ from direct prompt injection?

Direct injection arrives through an input the attacker controls directly, such as a chat field or an API parameter. Indirect injection arrives through a data source the model reads: a retrieved web page, a document the model is asked to summarise, or a database field. The attacker does not interact with the model directly; they place a payload in the data the model will process.

What is the blast radius in an agentic LLM system?

The blast radius is the set of actions a successful prompt injection can cause the model to take. In an agentic system, that includes everything the model is authorised to do with its tools: reading or writing files, calling APIs, sending email, querying databases. The wider the tool access, the larger the blast radius of a successful injection.

What should an LLM pentest scope document include?

The scope document should name the model and its version, the system prompt and any hard-coded instructions, all tools or plugins the model can call, all data sources it reads at inference time, and all output channels. A scope document that names only the API endpoint leaves the most significant attack surfaces unspecified.

How is severity rated in an LLM pentest finding?

Severity scales with the authorised blast radius. An injection that overrides the system prompt in a stateless model with no tools is lower severity than the same injection in an agent with write access to production systems. The finding records the blast radius so the severity rating is consistent and defensible.

What fields must an LLM pentest finding include that a standard web app finding might not?

Two fields are specific to LLM findings. The injection vector states the exact input the tester controlled and how it reached the model's context window. The trust boundary crossed names which guardrail or authorisation check the injection bypassed. Both are needed for the developer to choose the correct remediation.

How does an LLM pentest cover agentic systems differently from stateless chatbots?

An agentic system has tools the model can call. The test plan must include at least one injection test case per tool, because each tool is a potential pivot point for a successful injection. A stateless chatbot with no tools is tested for information disclosure and output manipulation only.

What data sources count as indirect injection surfaces?

Any data source the model reads at inference time is a potential indirect injection surface. This includes retrieved web pages, document stores, database records, calendar entries, email content and API responses. If the model treats the content of a retrieved source as instructions, that source is an injection vector.

Can an LLM pentest be run without access to the system prompt?

Yes, but the test will be less precise. Reconnaissance can extract portions of the system prompt through probing, and the tester can work from observed behaviour. Full access to the system prompt allows the tester to design targeted test cases against specific instructions and constraints, which increases coverage.

The engagements this applies to

API Pentest

Object and function level authorisation, tokens, rate limits

Joel Aviad Ossi, Red Team Lead at RedTeam Security
Joel Aviad OssiRed Team Lead, RedTeam Security

Joel Aviad Ossi is Red Team Lead at RedTeam Security, the Dubai-licensed trading brand of WebSec FZCO. He scopes and runs objective-based engagements across the UAE.

Let's talk about your security

Tell us the objective you want tested. We will come back with a scope, a timeline and a quote, under NDA from the first conversation.

Email [email protected]
IFZA Business Park, Building A2
Nadd Hessa, Dubai Silicon Oasis
Dubai, United Arab Emirates