Home
VAPT Web Application PentestAPI PentestMobile App PentestInfrastructure PentestAI & LLM PentestOT / ICS PentestIoT PentestPenetration TestingAll VAPT services
Red Team Red Team EngagementAdversary SimulationAssumed BreachPurple TeamingSocial Engineering
CompanyResourcesBlogFree Consultation

How to buy an AI and LLM pentest: the layers to scope, what to hand over and what the proposal must say

Scoping an AI pentest starts with what the model can reach, not which model it is. Here is how to describe the system, what to hand over, and what a proposal has to say.

Joel Aviad OssiJoel Aviad OssiRed Team Lead, RedTeam Security
7 min read
A glowing filament sphere wired to a row of tool sockets and a stack of indexed documents, one wire lit red where it meets a locked vault under a magnifying lens

Key takeaways

  • An AI and LLM pentest assesses five layers: the system prompt and context, the retrieval layer, the tools the model can call, the permissions those calls run with, and the application underneath. The model itself is the least of it.
  • Scope is set by what the system can reach and who can put text in front of it, not by which model vendor sits behind the chat window.
  • Two properties decide the severity ceiling before any prompt is sent: whether the model acts or only answers, and whether it ingests content that people outside your organisation can shape.
  • White box is the honest default for this engagement. The system prompt, tool definitions and retrieval configuration decide what the model is permitted to do, and withholding them costs coverage without adding realism.
  • A proposal that describes sending adversarial prompts to the assistant is describing one layer of five. Ask for indirect injection, tool abuse, retrieval authorisation and the surrounding application, in writing.

What you are buying: a test of an architecture, not a chat window

An AI and LLM pentest is a test of an application built on a language model, and the model is the least of it. What is assessed is the system prompt and the context window, the retrieval layer that feeds documents into it, the tools and functions the model is allowed to call, the permissions those calls run with, and the ordinary web application and API underneath. A security test of an AI assistant or agent that stops at the chat window has tested the part with the smallest consequences.

The failure that matters is not the model saying something embarrassing. It is the model being persuaded, by text somebody other than the user supplied, to call a function it should not have called or to return a record the requester was never entitled to read. The OWASP Top 10 for LLM Applications names these directly: prompt injection is LLM01, excessive agency is LLM06, and vector and embedding weaknesses are LLM08. Underneath each of them sits an ordinary application flaw with a new delivery mechanism.

The first thing you are buying is coverage of those layers, in writing. A proposal that describes the engagement as sending adversarial prompts to the assistant is describing one layer of five, and the one with the lowest ceiling.

Scope has to name every layer, because the finding that matters usually sits below the model.

Scope by what the model can reach, not by which model it is

The scoping conversation for a conventional application starts with hosts and endpoints. Here it starts with a different question: what can this system reach, and who can put text in front of it? The model vendor matters less than whether the assistant can send email, whether it queries a database with its own credential, and whether it reads pages, files or messages that people outside your organisation can write. The method in how to scope a penetration test still applies. The inventory is simply made of different things.

Write the inventory down before you ask for a quote. MITRE ATLAS is the tactic and technique catalogue for attacks on AI-enabled systems, and it works as a checklist for what an attacker would want from each layer, in the same way ATT&CK does for a network. The table below is the shape of the answers a tester needs from you.

If any row comes back as unknown, that is a finding before the test starts. It belongs in the scope as something to discover, not something to skip.

The scope inventory for an LLM application: one row per layer, the question it answers and what the tester needs to answer it.
LayerQuestion the scope answersWhat the tester needs from you
Context and promptWhat is in the context window that assumed nobody would read it?The system prompt and any templates that wrap user input
RetrievalDoes the index enforce the requester's permissions, or filter after retrieval?Retrieval configuration, index tenancy model, a user in each tenant
Tools and functionsWhich functions can the model call, with which parameters?Tool definitions with parameter schemas
PermissionsDoes each call run as the signed-in user or as a service account?The identity and scope behind each tool credential
Ingested contentWhich sources can somebody outside the organisation write to?A list of mail, ticket, document and web sources the model reads
ApplicationWhat surrounds the model: auth, tenancy, rate limits, API surface?The same access a normal user or partner would have

Blast radius and untrusted content decide the tier

Two properties of the deployment set its severity ceiling before anyone sends a prompt. The first is agency: whether the model only answers, or whether it acts, by calling tools that change state, move money, send messages or write records. The second is exposure: whether the content the model reads is entirely yours, or whether it ingests documents, web pages, tickets or mail that somebody else can shape. Put the two on a grid and each deployment lands in one cell.

An assistant that acts on content others control is the cell that needs the full engagement, because it is where indirect prompt injection turns into an action taken with the agent's credential. An assistant that only answers questions over your own documents still needs the retrieval layer tested for authorisation and tenant separation, and it needs the surrounding interface tested like any other API, but the ceiling is disclosure rather than action. State which cell you are in and the proposal can be sized honestly.

The cell also tells you which finding to expect first. Read-only deployments produce disclosure findings: a record from another tenant, a system prompt, a key left in the context. Acting deployments produce authorisation findings: a function called with a parameter the requester could never have supplied.

Agency and exposure fix the severity ceiling before a single prompt is sent, and they size the engagement.

What to hand over, and why white box is the honest default here

For a web application, grey box with a normal user account is the usual starting position, and it covers the logic flaws a black box run never reaches. For an LLM application the case for going further is stronger. The question is not what the model will say but what it is permitted to do, and the answer lives in artefacts you already hold: the system prompt, the tool definitions with their parameter schemas, the retrieval configuration, and the identity each tool call runs under. Withholding them does not make the test more realistic. It makes the tester spend the first days recovering what you could have sent in an email, which is the coverage-against-realism trade NIST SP 800-115 sets out for security testing.

The tool definitions are the single most useful document. A worked example: an assistant for a support desk that can look up a customer and send them a message. The definition alone shows what an attacker who controls the model's input would be able to do.

Read that as the tester will. The send function takes any address and any body, the lookup takes any customer identifier, and both run with a service account rather than the signed-in agent's own permissions. Nothing in the model has to be broken for that to become a data exfiltration path. The list in what to have ready before a penetration test starts still applies; add these four artefacts to it.

Example tool definitions, as handed over
tools:
  - name: lookup_customer
    description: Return the customer record for an identifier
    parameters:
      customer_id: string   # any value accepted, no ownership check
    runs_as: svc-support-assistant   # service account, read on all customers

  - name: send_message
    description: Send an email to a customer
    parameters:
      to: string            # any address accepted
      body: string          # free text, model-composed
    runs_as: svc-support-assistant   # can send to any recipient

context_sources:
  - ticket_body           # written by the public
  - knowledge_base        # internal

Reading the proposal: a test or a demonstration

The hardest part of buying in a new category is telling a testing engagement from a demonstration that produces screenshots. The proposal has to say, in writing, that indirect injection through every ingested content source is in scope, that tool and function calling will be tested for parameter injection and unintended invocation, that the retrieval index will be tested for authorisation, and that the surrounding application is tested rather than assumed. The general questions in how to compare penetration testing proposals still separate the candidates. These are the questions specific to this engagement.

Look also for a stated method for proving findings. A jailbreak screenshot proves that the model can be made to say something. A finding here should show the function that was called, the parameter the attacker supplied and the record that came back, with the steps to reproduce it. If the proposal does not say that findings are exploited rather than asserted, ask.

Ask, too, how findings will be classified. A report written against the OWASP LLM Top 10, with ATLAS or ATT&CK references where a technique applies, can be read by the same people who read the rest of your testing. A report written as a list of clever prompts cannot.

Phrases that appear in AI pentest proposals, and the question that tells you what is actually being sold.
The proposal saysWhat to ask
Adversarial prompt testingIs indirect injection through documents, mail and tickets in scope, or only prompts typed by the tester?
Jailbreak assessmentWhat is the objective once the guardrail is bypassed: a tool call, a record, a credential? Or a screenshot?
Red teaming the modelIs the retrieval index, the tool layer and the application underneath tested, or only the model's answers?
Automated LLM scanningWhat is done by hand after the scanner runs, and is anything reported on tool output alone?
Findings mapped to a frameworkWhich one: OWASP LLM Top 10, ATLAS, ATT&CK? And is the mapping per finding or a paragraph in the summary?

What the report has to give you

The report is the deliverable, and it has to do two jobs. The first is the ordinary one described in what a penetration test report has to contain: findings with severity, reproduction steps, named owners and a remediation order that reflects the real chain rather than the raw score. The second is specific to this category. Every finding should be classified against the OWASP LLM Top 10 and, where a technique applies, mapped to ATLAS or ATT&CK, so that the AI assessment sits inside your existing risk reporting instead of beside it.

A redacted example of the shape. Finding: indirect prompt injection via a support ticket leads to an unauthorised customer lookup. Severity: high. Chain: an attacker submits a ticket whose body carries instructions; the assistant summarising the queue reads them and calls the lookup tool with a different customer's identifier; the tool runs under the service account and returns the record; the assistant includes it in the summary shown to the agent. Fix: scope the lookup tool to the requesting agent's own permissions, and treat ticket bodies as data rather than instructions.

That finding is not a model problem, and the fix is not a better prompt. It is an authorisation flaw in the layer beneath, which is why this engagement is bought alongside, not instead of, a web application pentest of the same product. The NIST AI Risk Management Framework puts testing under its Measure function, and a report in this shape is what that evidence looks like when an examiner asks for it.

Buying an AI pentest in Abu Dhabi and Dubai

None of the four UAE regimes that name security testing mentions language models, and that is the point rather than a gap. The obligation attaches to the system that holds the data and to the classification of the data it holds. A Dubai government entity, or a supplier to one, that puts an assistant in front of a case management system is running a system that DESC expects to be tested proportionately to its classification, and the standard already brings supplier connections into scope. A licensed bank in Abu Dhabi or Dubai that lets an agent read account data is running a system that CBUAE expects to be independently tested on a periodic basis, with findings tracked to closure.

The same reading holds federally under the NESA-published Information Assurance Standard, and in Abu Dhabi healthcare under ADHICS, where an assistant reading patient records is handling the most sensitive class of data the standard describes. The question for buyers in both emirates is whether your existing testing programme names the AI layer at all. If the annual test covers the web application and the API but nobody has scoped the retrieval index or the tool permissions, the examiner's evidence has a hole in it that no other engagement fills. What each UAE regulator requires is set out on our resources page. An AI assessment is evidenced under the same testing clause as everything else, and it should be written into the same programme, with the same owners and the same retest record.

Frequently asked questions

What is an AI and LLM pentest?

An AI and LLM pentest is a security assessment of an application built on a language model: the system prompt and context, the retrieval layer, the tools the model can call, the permissions those calls run with, and the web application and API underneath. It tests whether the system can be persuaded to take an action or return data the requester was not entitled to, rather than whether the model can be made to say something embarrassing.

How does an AI pentest differ from a web application pentest?

A web application pentest tests the interface, authorisation and business logic a user reaches directly. An AI pentest adds the layers a language model introduces: text supplied by third parties reaching the model as instructions, tools called with parameters the user never typed, and a retrieval index that may not enforce permissions. The application underneath still needs its own test, and the two are bought together, not as substitutes.

What is indirect prompt injection?

Indirect prompt injection is instruction text placed in content the model reads rather than in the user's own message: a document, a web page, an email or a support ticket. When the model treats that text as instructions, whoever wrote the content can steer what the model does. It is the highest-severity path wherever the system ingests content that people outside the organisation can write.

What is excessive agency in an LLM application?

Excessive agency is the OWASP LLM Top 10 term for a model that can do more than the task requires: too many tools, tools with too broad parameters, or tools running under a credential with more rights than the requesting user. It is the property that turns a prompt injection into an action, which is why the tool definitions and their permissions are the first thing a scope should name.

Do we still need an AI pentest if we use a hosted model from a cloud provider?

Yes. The provider is responsible for the model, not for your system prompt, your retrieval index, your tool definitions or the credentials your tools run under. Every finding in this category sits in those layers, which are yours regardless of where the model is hosted.

Should we hand over the system prompt and tool definitions before the test?

Yes. The question the test answers is what the model is permitted to do, and that is written in the tool definitions, the retrieval configuration and the identity each call runs under. Withholding them costs coverage without adding realism, because an attacker who reaches the tool layer does not need to read your prompt to abuse it.

How does an AI pentest differ from red teaming a model?

Red teaming a model probes what the model will say under adversarial input, and the result is a set of prompts that bypass a guardrail. An AI pentest starts where that ends: it asks what happens next, which tool is called, which record comes back, and whether the application underneath stops it. The first is a model evaluation; the second is a security test of a system.

Is jailbreaking in scope for an AI pentest?

It is a step, not the objective. A bypassed guardrail matters when it leads to a tool call, a record from another tenant or a credential from the context window, and a finding should show that consequence with steps to reproduce it. A screenshot of the model saying something it should not is not a finding on its own.

Does the test cover our RAG index and vector store?

It should, and the proposal should say so. The retrieval layer is tested for whether it enforces the requester's permissions at the index, or retrieves first and relies on the model to decline, and for whether tenants are separated in the store. A user in each tenant is what the tester needs to prove either way.

What framework should AI pentest findings be mapped to?

The OWASP Top 10 for LLM Applications for classification, and MITRE ATLAS or ATT&CK where a technique applies. Mapping per finding lets the AI assessment sit inside existing risk reporting rather than as a separate document nobody knows how to read.

The engagements this applies to

API Pentest

Object and function level authorisation, tokens, rate limits

Joel Aviad Ossi, Red Team Lead at RedTeam Security
Joel Aviad OssiRed Team Lead, RedTeam Security

Joel Aviad Ossi is Red Team Lead at RedTeam Security, the Dubai-licensed trading brand of WebSec FZCO. He scopes and runs objective-based engagements across the UAE.

Let's talk about your security

Tell us the objective you want tested. We will come back with a scope, a timeline and a quote, under NDA from the first conversation.

Email [email protected]
IFZA Business Park, Building A2
Nadd Hessa, Dubai Silicon Oasis
Dubai, United Arab Emirates