Skip to content
CyberSmithSECURE
Under Attack

Red Teaming

Red Teaming of AI and LLM Systems

An LLM application is a system that takes untrusted input and hands it to something with permissions. The model is rarely the interesting target; the interesting target is what it is connected to — the retrieval corpus, the tool it can call, the database it can query, the email it can send. Testing therefore concentrates on where instructions cross a trust boundary, which is a class of flaw most application security programmes have no existing coverage for.

Methodology

  1. 01

    Architecture and trust boundary mapping

    Every input source, retrieval corpus, tool, function and downstream system the model can reach, and the privilege each carries. The map is the assessment's foundation because the risk is in the connections.

  2. 02

    Direct prompt injection

    System prompt extraction, instruction override, role manipulation and guardrail bypass tested systematically rather than by trying jailbreaks found online.

  3. 03

    Indirect prompt injection

    Instructions planted in content the model will later ingest — documents, web pages, emails, database fields — testing whether data becomes instruction. This is the highest-severity class in most deployments.

  4. 04

    Tool and agent abuse

    Whether the model can be induced to call tools outside its intended purpose, chain calls to reach unintended systems, or pass attacker-controlled arguments to a privileged function.

  5. 05

    Data exposure testing

    Training and retrieval data leakage, cross-tenant retrieval in shared corpora, and whether the model returns content the requesting user is not entitled to.

  6. 06

    Output handling testing

    Whether model output is treated as trusted downstream — rendered as HTML, executed as code, passed to a shell or used in a query without validation.

  7. 07

    Resource and cost abuse

    Token exhaustion, recursive agent loops, and whether an attacker can impose unbounded cost.

  8. 08

    Reporting and retest

    Technical report and executive summary together, then a retest confirming closure — with the caveat that model behaviour is probabilistic and closure is evidenced across repeated trials.

Approach to testing

  • Findings are demonstrated across repeated trials, not once. A model is probabilistic, and a bypass that works one time in twenty is still a finding — but it is reported with its success rate rather than as a binary.
  • Guardrail bypass alone is not treated as high severity. The severity comes from what the bypass reaches: a model that can be made to say something rude is a brand issue, one that can be made to call a payment tool is a security issue.
  • Indirect injection is prioritised over direct, because the user is usually not the attacker — the attacker is whoever controls a document the model will read.
  • Agent systems are tested for privilege at the tool level, since the model's effective permission is the union of every tool it can call.
  • Model updates invalidate results. The report states the exact model version and date, because the same prompt against a newer version may behave differently.

Types of assessment

LLM application assessment (default)

A chat or completion application with retrieval. Covers injection, data exposure and output handling.

Agentic system assessment

Systems where the model calls tools and takes actions. Higher risk and a different methodology, because the blast radius is real.

RAG pipeline assessment

Focused on the retrieval corpus: poisoning, cross-tenant leakage and whether document-level permissions survive into retrieval.

Model integration review

Architecture-level review of trust boundaries, tool permissions and output handling, without adversarial testing. Appropriate pre-launch.

Frameworks and standards

OWASP Top 10 for LLM Applications
The primary taxonomy — prompt injection, insecure output handling, excessive agency, and the rest.
MITRE ATLAS
Adversarial threat landscape for AI systems, used for technique mapping and reporting.
NIST AI Risk Management Framework
Governance context where the client needs to demonstrate a managed approach to AI risk.
OWASP ASVS
The conventional application controls around the model, which are frequently weaker than the model controls.
EU AI Act
Where the client's system falls in scope, obligations are noted alongside technical findings.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Garak

Automated probing across known LLM vulnerability classes as a first pass.

PyRIT

Microsoft's risk identification toolkit for structured adversarial testing at scale.

Promptfoo

Regression testing so a fixed bypass can be verified across repeated trials and future model versions.

Burp Suite

The application around the model, which is still a web application with conventional flaws.

Custom injection corpora

Payloads written for the client's specific tools and retrieval sources, because generic jailbreaks test the model vendor, not the client.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Prompt injection

  • System prompt extraction and disclosure
  • Instruction override and role manipulation
  • Guardrail bypass, with success rate across trials
  • Indirect injection through retrieved documents
  • Indirect injection through user-supplied content stored and later read
  • Multi-turn injection building across a conversation

Excessive agency

  • Complete tool inventory and the privilege each carries
  • Tools callable outside their intended context
  • Chained tool calls reaching unintended systems
  • Attacker-controlled arguments passed to privileged functions
  • Human-in-the-loop requirements for consequential actions
  • Whether the model can modify its own instructions or configuration

Data exposure

  • Retrieval returning content outside the user's entitlement
  • Cross-tenant leakage in shared vector stores
  • Training data extraction where the model is fine-tuned on client data
  • Sensitive data in system prompts
  • Conversation history isolation between users

Output handling

  • Model output rendered as HTML without sanitisation
  • Output used in SQL, shell or code execution
  • Output passed to downstream systems as trusted input
  • Markdown and link rendering enabling exfiltration
  • Structured output validated against a schema

Resource and availability

  • Token and request rate limiting per user
  • Recursive agent loop termination
  • Cost controls and alerting
  • Context window exhaustion handling

Governance

  • Model version pinning and change management
  • Logging of prompts and completions, and its privacy implications
  • Human review for consequential actions
  • Incident process for model misbehaviour

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Model behaviour is probabilistic, so a single trial proves nothing. Agents run each payload hundreds of times and report a success rate, which is the only honest way to characterise a bypass.

  • Injection payload space is effectively unbounded. Parallel agents explore mutation and combination far beyond what a tester types by hand, and the successful payload is rarely the obvious one.

  • Multi-turn injection builds across a conversation; exploring conversation trees is combinatorial and is exactly what parallel agents are for.

  • Indirect injection requires planting content in every ingestible source and waiting to see what surfaces — a wide, patient search rather than a clever one.

Agents generate and evaluate payloads; a human assessor judges severity and confirms impact. This is the one capability where the swarm is testing a system of the same kind, and the report is explicit that automated results are filtered by a human before publication — including the many that look alarming and are not.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Severity from reach, not from output

Most AI red teaming reports jailbreaks. CSS reports what the bypass reaches — a rude response and a callable payment tool are not the same finding, and conflating them wastes the client's attention.

Success rates, not anecdotes

Every bypass is reported with a measured success rate across repeated trials, because a probabilistic system cannot honestly be described with a single example.

Indirect injection prioritised

The user is usually not the attacker. Testing concentrates on content the model ingests from elsewhere, which is where real compromise comes from and where most assessments are thin.

Regression suite delivered

Findings ship as a Promptfoo suite the client can run against future model versions, because a model update can silently reopen a closed finding.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Scope: application, model version and date, tools, retrieval sources and the test window
  • Findings mapped to OWASP LLM Top 10 and MITRE ATLAS
  • Every bypass with its measured success rate across trials
  • Trust boundary map showing what each finding reaches
  • Reproduction payloads verbatim, with the conversation context required
  • Tool inventory with effective privilege per tool
  • Explicit statement that results are bound to the tested model version
  • Retest results appended, evidenced across repeated trials

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • What an attacker could make this system do, in business terms
  • Whether the risk is reputational, data exposure, or unauthorised action — these are very different and are usually conflated
  • The three changes that most reduce blast radius
  • Regulatory position where the EU AI Act or sector guidance applies
  • The model version dependency, stated plainly, because it affects how long the assessment remains valid
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A financial services firm's customer support assistant with retrieval over internal knowledge base articles.

Finding
A knowledge base article editable by any support agent was used to plant instructions. When the assistant retrieved that article, it followed them — disclosing its system prompt and, in 34% of trials, including content from other customers' conversation summaries held in the same vector store.
Recommendation
Treat retrieved content as data rather than instruction through prompt structure and delimiters, partition the vector store per tenant, and restrict knowledge base editing to a reviewed workflow.
Outcome
Partitioning and editorial review implemented before launch. The regression suite runs in CI and caught a reoccurrence when the retrieval prompt was refactored.

A logistics company's internal agent with tools for querying shipments and sending customer emails.

Finding
A shipment note field, populated from customer-supplied data, contained injected instructions. Processing that shipment caused the agent to send an email to an attacker-supplied address containing data from the query results. The email tool had no recipient allowlist and no human approval step.
Recommendation
Restrict the email tool to verified recipient domains, require human approval for any outbound message, and sanitise customer-supplied fields before they enter the model context.
Outcome
Recipient allowlist and approval step implemented within two weeks. The finding reframed the client's whole approach: tools now start with no privilege and gain it by justification.

A healthtech startup's clinical summarisation feature.

Finding
Model output was rendered as Markdown in the clinician interface without sanitisation. An injected instruction caused the model to emit an image tag with a URL containing summarised patient data, which the browser fetched automatically — exfiltrating clinical data to an external server on render.
Recommendation
Sanitise model output before rendering, disable automatic remote resource loading in the interface, and apply a content security policy restricting outbound requests.
Outcome
All three implemented before the feature reached production. The client extended output sanitisation to every model-facing surface after the finding.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.