Direct and indirect prompt injection
Attempts to redirect behavior through user input or content the application retrieves.
Utilities Studio / Cybersecurity
Before you release an AI feature, test what it can reveal or do.
Assess prompt injection, RAG data exposure, and agent tool abuse. Test the permissions and trust boundaries around your AI application.
You are connecting an AI feature to customer records, internal documents, or tools that can take action. Your engineering team needs to know whether a manipulated prompt can cross those boundaries. We scope the assessment around the data the application can retrieve and the actions its tools permit.
The assessment
AI and LLM penetration testing examines whether an AI application can be manipulated into disclosing data or taking unauthorized actions. We assess the application around the model: its instructions, retrieved content, user permissions, and connected tools. Findings explain the conditions that produced the behavior and where your team can apply a control.
Attempts to redirect behavior through user input or content the application retrieves.
Whether users can retrieve information outside their authorized data sources or account boundaries.
Whether manipulated instructions can trigger actions the user or agent should not be allowed to perform.
Exposure of protected data through model responses, retrieved context, or connected application behavior.
Authentication and authorization around the AI feature, with conventional application testing included where agreed.
The OWASP guidance describes risks such as prompt injection and excessive agency. Your application architecture, retrieval permissions, and available tools determine which risks need testing in the engagement.
OWASP Top 10 for LLM ApplicationsWorking with your team
Confirm the assets, permissions, and production limits. Name the contacts and record the dates, reporting format, support arrangements, and retest terms.
Investigate the agreed attack paths and validate findings. Keep your team updated and escalate critical issues immediately through the agreed channel.
Walk your engineers through the report and remediation priorities. Carry out the agreed retesting and document which fixes worked and what remains unresolved.
Application complexity, user roles, retrieval sources, connected tools, environment access, and the depth of retesting determine assessment effort.
Testing begins after scope, authorization, and environment access are agreed. Complex agent actions and remediation retesting affect the engagement schedule.
Pentest delivery
We share validated findings during the test through the agreed secure channel. Critical issues go to your nominated contact immediately. Progress updates cover completed work, blockers, and what comes next.
Your engineers get affected assets, reproduction steps, evidence, and remediation guidance. We explain severity using the demonstrated impact. An executive summary sets out the business risk and the limits of the assessment.
A technical findings review lets your engineers discuss the evidence and recommended fixes with us. We name the technical contact and agree the support period and response arrangements before testing.
Retesting checks fixes to the original findings and records the result. Before booking, we specify the findings covered, retest rounds, time window, and any charges. New features or changed environments need a scope review.
We agree the findings format and handover method with your team. If you use Jira, Linear, or GitHub, we scope the export or ticket handover, required access, and treatment of sensitive evidence before testing.
Your proposal sets the start date, testing window, and report delivery date after we review scope and access. Bring your audit or release deadline so remediation and retesting can be planned around it.
Delivery references: NIST SP 800-115 and CREST's penetration testing programme guide.
The practitioner behind the work
Sheeraz Ali is our Head of Cybersecurity. His work spans application, cloud, network, and AI assessments. His personal track record includes leading pentests at Cobalt and building the internal pentest programme at SolarWinds.
Read Sheeraz's security backgroundSheeraz's personal track record
His website lists OSCP, CRTP, CRTE, CREST CRT and CPSA, CBBH, and CKA.
At SolarWinds, he delivered 120+ internal pentests. As CTO at Pwned Labs, he built a platform serving 40,000+ practitioners. He co-developed Mobexler, selected for Black Hat Arsenal, and presented research at Nullcon and c0c0n.
Explore his career timelineFAQ
An AI assessment examines instruction boundaries, retrieved content, generated responses, and the actions available to agents. Conventional authentication, authorization, and API issues may also affect the system. We define which layers the engagement covers so there are no gaps hidden by the service label.
Yes, within an authorized scope and suitable environment. We need to understand user roles, retrieval permissions, and tool access. Test accounts and representative data let the assessment examine whether the application respects the boundaries your business expects.
No. An assessment evaluates a defined system and scope at a point in time. New models, prompts, tools, or data sources can change its behavior. Findings should inform fixes and ongoing evaluation, with retesting of the changes that matter.
Bring the previous report and a list of changes to your code, permissions, or infrastructure. We can use those to scope the next assessment. Check which earlier findings were fixed and which fixes were verified; the date on the old report does not answer those questions.
Assess API authentication, object-level authorization, role boundaries, and sensitive data exposure. Get evidence your developers can reproduce.
Test authentication, access controls, and business logic in your web application. Get reproducible findings and remediation guidance for your engineering team.
Review privileged access, service accounts, authentication, and permission boundaries. Get an IAM assessment and practical least-privilege recommendations.
Tell us what your team needs to resolve, which systems are involved, and any deadline. We will work through the scope and reporting needs with you.