Red Teaming
Full Red Team Engagement
A penetration test answers what is vulnerable. A red team answers whether the organisation would notice, and what it would do. The engagement is defined by an objective — reach the payment system, exfiltrate the design files, obtain domain administrator — and every technique is judged on whether it moved towards that objective without being caught. Findings are as much about the security operations function as about the estate.
Methodology
- 01
Objective setting and rules of engagement
Agreed with a small control group: what would constitute success, what is out of bounds, escalation contacts, and the conditions under which the engagement is paused or disclosed.
- 02
Threat intelligence and adversary selection
A relevant threat actor profile chosen for the client's sector, so techniques reflect a plausible adversary rather than a tester's preferences.
- 03
Reconnaissance
External footprint, staff enumeration from public sources, technology fingerprinting, supplier relationships and leaked credentials — all without touching client infrastructure.
- 04
Initial access
Phishing, exposed service exploitation, credential reuse, or physical access where in scope. Multiple routes are prepared because the first is often burned.
- 05
Establish and maintain foothold
Command and control established with tooling tuned to evade the client's specific defences, and persistence placed and tested.
- 06
Escalation and lateral movement
Privilege escalation, credential harvesting and movement towards the objective, with each step chosen for stealth as much as effectiveness.
- 07
Objective execution
The agreed objective achieved and evidenced, without causing the harm the real adversary would.
- 08
Deconfliction and replay
Full timeline shared with the blue team, walked through step by step so they can compare against what their tooling recorded.
Approach to testing
- Objective-based, not coverage-based. Success is reaching the agreed target, not enumerating every weakness on the way.
- The blue team is not informed. Only a small control group knows, because telling the defenders measures a rehearsal rather than the organisation.
- Stealth is a constraint throughout. A technique that works loudly is recorded as available but not used, because the value of the engagement is in what goes undetected.
- Every action is logged with a timestamp so the client can reconcile against their own telemetry during replay. Without that log the debrief is two parties comparing recollections.
- Destructive actions are never taken. Where the objective is data exfiltration, a marker file is used rather than real client data.
Types of assessment
Full red team (default)
No knowledge for defenders, objective-based, multiple access vectors including social engineering. The most realistic and the most demanding.
Intelligence-led (TIBER-style)
Threat intelligence phase producing a bespoke actor profile, then emulation of that actor specifically. Required by some financial regulators.
Scenario-based
A single agreed scenario — ransomware operator, insider, supplier compromise — run to conclusion. Shorter and easier to schedule.
Red team with physical
Includes physical access attempts: tailgating, badge cloning, device placement. Requires separate written authorisation and a get-out-of-jail letter.
Frameworks and standards
- MITRE ATT&CK
- Every technique mapped, producing a coverage comparison against the client's detection capability.
- TIBER-EU / CBEST
- Structure for intelligence-led testing where a regulator requires it.
- PTES
- Execution standard for the technical phases.
- MITRE Engenuity ATT&CK Evaluations
- Reference for how detection efficacy is described, so results are comparable to vendor evaluations the client may have seen.
- Cyber Kill Chain
- Used in the executive narrative, because it is the model most boards recognise.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Cobalt Strike / Sliver / Mythic
Command and control, with profiles tuned per engagement to avoid signature detection.
Custom loaders and droppers
Written per engagement, because reused tooling gets caught by the tooling that caught it last time.
Gophish / Evilginx
Phishing campaign delivery and credential relay where MFA is in scope.
BloodHound
Attack path analysis once a domain foothold exists.
Impacket / Rubeus
Credential access and lateral movement.
Proxmark / Flipper Zero
Badge cloning where physical access is in scope.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Initial access
- Phishing resistance: click rate, credential submission, MFA relay success
- Externally exposed services and their exploitability
- Credential reuse from public breach data
- Supplier and third-party access routes
- Physical access controls where in scope
Execution and evasion
- Endpoint protection efficacy against tailored payloads
- Application control and script execution restrictions
- Email and web filtering effectiveness
- Egress filtering and command-and-control channel viability
- Detection of known tooling versus custom tooling
Persistence and privilege
- Persistence mechanisms that survive reboot undetected
- Local and domain privilege escalation paths
- Credential harvesting opportunities
- Service account abuse
Lateral movement
- Network segmentation as an actual constraint on movement
- Detection of lateral movement techniques
- Time from first movement to objective
- Jump host and administrative tier controls
Detection and response
- Which actions generated an alert, and at what stage
- Time to detection and time to response, per phase
- Whether alerts were investigated or closed
- Escalation path and out-of-hours coverage
- Containment effectiveness once detection occurred
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Reconnaissance across public sources, breach corpora and certificate transparency is breadth work; the useful finding is usually one credential or one forgotten subdomain among thousands.
Attack path analysis after a foothold is graph traversal — agents enumerate the paths, the operator selects the quiet one rather than the short one.
Detection correlation during replay requires matching every logged action against client telemetry, which is a join across two large datasets rather than a discussion.
Infrastructure preparation and payload variation benefit from parallel generation, because a single burned payload should not end an engagement.
Operational decisions in a red team stay with the operator throughout. Agents support reconnaissance, analysis and correlation; they do not execute against client infrastructure, and nothing runs on a client system without an operator's deliberate action.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Detection is the deliverable
The vulnerability list is secondary. The report states what was detected, when, by which control, and whether anyone acted — which is the question a red team exists to answer and many vendors skip.
Full replay with the blue team
A step-by-step walkthrough against the client's own telemetry, so defenders learn what their tooling saw and missed. An engagement that ends at the report teaches the defenders nothing.
Both reports, always
Technical report and executive summary together, plus an attack narrative the board can follow without a translator.
Custom tooling per engagement
Reused frameworks get caught by the signatures that caught them elsewhere, which measures the vendor's tooling rather than the client's defences.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Objective, rules of engagement, control group and the engagement window
- Complete timeline of every action taken, with timestamps for telemetry reconciliation
- Attack narrative from initial access to objective
- MITRE ATT&CK technique coverage with detection outcome per technique
- Detection and response timeline: what alerted, when, and what happened next
- Findings enabling each phase, with severity in terms of the objective
- Recommendations split between prevention and detection
- Retest results where a follow-up engagement is run
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Whether the objective was achieved, and how long it took
- Whether the organisation detected the engagement, and at which stage
- The three changes that would most have slowed or stopped the attack
- Comparison against the previous engagement, if one exists
- Regulatory position where intelligence-led testing is mandated
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A financial services firm with a 24x7 security operations centre.
- Finding
- Initial access came from a phishing email to three finance staff; one submitted credentials to a relay page and the MFA prompt was approved. From there, domain administrator took eleven days using only built-in tooling. The SOC generated four alerts during the engagement and closed all four as false positives.
- Recommendation
- Number-matching MFA to defeat prompt fatigue, detection rules for the specific living-off-the-land techniques used, and a triage process change requiring correlation before an alert is closed.
- Outcome
- MFA changed within a month. The follow-up engagement a year later was detected on day two, at the credential access stage.
A manufacturer with a recently deployed EDR platform.
- Finding
- Off-the-shelf tooling was blocked immediately and reliably. A custom loader written for the engagement executed without detection, established command and control over DNS, and persisted for the full three weeks. Egress filtering permitted DNS to any resolver.
- Recommendation
- Restrict DNS egress to internal resolvers only, add detection for anomalous DNS volume, and tune EDR toward behavioural rather than signature detection.
- Outcome
- DNS egress restricted within two weeks, which closed the channel. The behavioural tuning was a longer programme and is what the follow-up engagement will measure.
A healthcare provider, scenario-based engagement simulating a ransomware operator.
- Finding
- The objective — reaching a position from which patient systems could be encrypted — was achieved in six days. The critical enabler was a backup service account with domain administrator rights, which also meant the simulated adversary could have destroyed the backups before encrypting.
- Recommendation
- Separate backup identity from the production domain, and treat the backup platform as a tier-zero asset.
- Outcome
- Backup identity separated over a quarter. The finding also triggered a separate backup resilience review, which found two further recovery gaps.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.