Introduction
Volkis will perform penetration testing on the Artificial Intelligence (AI) systems within the agreed scope. This will include identifying AI functionality, understanding how the AI system interacts with users and external services, identifying vulnerabilities within the AI implementation, attempting to exploit those vulnerabilities where appropriate, and analysing and reporting on the results.
The objective of the assessment is to evaluate the security of the AI system, identify weaknesses that could be exploited by an attacker, and provide practical recommendations to reduce risk while maintaining the intended functionality of the platform.
AI Identification and Enumeration
The tester will identify and enumerate the AI components that form part of the application or platform. This includes understanding the architecture, trust boundaries, available interfaces and supporting infrastructure used by the AI system.
This may include:
- AI models and providers
- Retrieval Augmented Generation (RAG) components
- Memory and conversation history
- Function calling and AI agents
- MCP servers and tool integrations
- External APIs and third-party services
- Knowledge bases and document repositories
- Authentication and authorisation mechanisms
- Administrative functionality
AI Vulnerability Identification
The tester will identify vulnerabilities within the AI implementation by reviewing the attack surface exposed by the application, its supporting services and integrated components. Testing will assess whether the AI system can be manipulated to disclose sensitive information, perform unintended actions or bypass implemented security controls.
Testing may include:
- Prompt injection
- Indirect prompt injection
- System prompt disclosure
- Context manipulation
- Retrieval poisoning
- Excessive agency
- Insecure output handling
- Insecure model configuration
Prompt Injection Testing
The tester will attempt to manipulate the AI system using crafted prompts designed to bypass safety controls, modify the model's behaviour, expose hidden instructions or influence downstream actions. Both direct and indirect prompt injection techniques will be assessed where applicable.
Data Leakage and Information Disclosure
The tester will assess whether sensitive information could be disclosed through AI responses. Testing will determine whether prompts could expose confidential data, system prompts, conversation history, embeddings, knowledge base content or other information not intended for the current user.
Access Control
The tester will assess authentication, authorisation and privilege enforcement across the AI system. Testing will determine whether users could access functionality, information or administrative capabilities beyond their intended level of access.
As part of this, the tester will also attempt to determine whether the AI can be used outside of what is intended, costing the client money and resources.
AI Safety Controls
The tester will evaluate the effectiveness of the AI system's safety controls and moderation mechanisms. Testing will assess whether safeguards designed to prevent unsafe, harmful or restricted behaviour could be bypassed.
Tool and Integration Security
The tester will assess the security of AI tool integrations, plugins, MCP servers and external services. Testing will determine whether the AI system could be manipulated into performing unintended actions or placing undue trust in user-controlled or otherwise untrusted data.
Exploitation
Where vulnerabilities are identified, the tester will attempt to exploit them to validate their impact and eliminate false positives. Exploitation may include bypassing AI safety controls, retrieving unauthorised information, manipulating model behaviour, abusing integrated tools or influencing downstream business processes.
Where full exploitation presents an unacceptable operational risk, vulnerabilities will be validated using alternative techniques and reported accordingly.
Post-Exploitation
Where exploitation is successful, the tester will assess the potential business impact. This may include determining whether sensitive information, connected business systems, privileged functionality or downstream services could be accessed or influenced. The tester will also assess whether malicious prompts, instructions or memory persisted beyond the initial interaction and whether further compromise could reasonably be achieved.
Reporting
Following completion of testing, Volkis will document all identified findings together with the associated business risk, technical details, supporting evidence and practical remediation recommendations. Findings will be prioritised according to their impact and likelihood, enabling remediation efforts to focus on the highest-risk issues first.