Skip to content
CyberSmithSECURE
Under Attack

Red Teaming

Purple Team Exercises

A red team tells you whether you were caught. A purple team tells you why not, and fixes it the same afternoon. Techniques are executed openly with the security operations team watching their own console, so every gap is diagnosed immediately — was there no telemetry, no rule, or a rule that fired and was ignored? Those three have completely different remedies and a covert engagement cannot distinguish them.

Methodology

  1. 01

    Technique selection

    An ATT&CK technique set chosen for the client's threat profile and current coverage, agreed in advance so the exercise is targeted rather than exhaustive.

  2. 02

    Baseline coverage assessment

    Existing detection rules and log sources reviewed against the selected techniques, producing an expected-coverage map before anything is executed.

  3. 03

    Execution with observation

    Each technique executed in a controlled way with the blue team watching. Timestamped so telemetry can be located precisely.

  4. 04

    Immediate triage

    For each technique: did telemetry exist, did a rule fire, did an analyst see it, did they act. The gap is classified rather than simply recorded as a miss.

  5. 05

    Rule development

    Where a gap is a missing rule, a detection is written and tested during the exercise, then re-run to confirm it fires.

  6. 06

    Telemetry gap remediation

    Where a gap is missing telemetry, the logging change is identified and scoped — usually a longer piece of work than a rule.

  7. 07

    Re-execution and verification

    Every technique re-run after remediation, so the exercise ends with measured improvement rather than a to-do list.

  8. 08

    Reporting

    Coverage before and after, with the rules developed handed over as artefacts.

Approach to testing

  • Collaborative and overt throughout. The value is in the diagnosis, and hiding from the defenders destroys it.
  • Gaps are classified into three kinds — no telemetry, no rule, no action — because they have different owners and different costs, and a report that says 'not detected' helps nobody.
  • Detection rules are written during the exercise and verified by re-execution, not recommended for later.
  • Coverage is measured before and after, so the deliverable is an improvement figure rather than an assessment.
  • Technique selection is deliberately narrow. Twenty techniques closed properly beats a hundred and fifty recorded as gaps.

Types of assessment

Standard purple team (default)

An agreed technique set executed over several days with the blue team present, rules written as gaps are found.

Detection engineering sprint

Focused on building coverage for a specific tactic — credential access, lateral movement, exfiltration — rather than breadth.

Post-red-team purple

Run after a covert red team, working through the techniques that went undetected. The most effective sequence, because the gaps are already known.

Continuous validation

A recurring subset of techniques re-run monthly or quarterly to confirm detections still work after platform and rule changes.

Frameworks and standards

MITRE ATT&CK
The technique catalogue and the coverage model the whole exercise is scored against.
Atomic Red Team
Standardised, bounded technique implementations, so execution is reproducible and safe.
MITRE D3FEND
Mapping defensive countermeasures to the techniques they address.
Detection Engineering maturity models
Used to frame where the client's detection function sits and what to build next.
Sigma
Rules are written in Sigma where possible so they are portable between SIEM platforms.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Atomic Red Team

Bounded, documented technique execution with defined cleanup.

Caldera

Automated adversary emulation for longer chains, where a scripted sequence is more repeatable than manual execution.

Sigma / SIEM native rule languages

Writing detections during the exercise, portable where the platform allows.

Client SIEM and EDR consoles

The exercise is run against the client's own tooling, sitting beside their analysts.

VECTR

Tracking technique outcomes and coverage before and after, and producing the comparison view.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Initial access and execution

  • Phishing attachment and link execution
  • Malicious macro and script execution
  • Living-off-the-land binary usage
  • PowerShell and command-line logging coverage
  • Application control bypass techniques

Persistence and privilege escalation

  • Registry run keys and startup folder
  • Scheduled task creation
  • Service creation and modification
  • Token manipulation and UAC bypass
  • Accessibility feature abuse

Credential access

  • LSASS memory access
  • Credential dumping from registry hives
  • Kerberoasting and AS-REP roasting
  • Browser and credential manager access
  • Cached credential retrieval

Discovery and lateral movement

  • Domain and network enumeration
  • Remote service execution: WMI, WinRM, PsExec
  • Pass-the-hash and pass-the-ticket
  • Remote desktop and admin share usage
  • Scheduled task creation on remote hosts

Collection and exfiltration

  • Archive creation and staging
  • Exfiltration over HTTPS, DNS and cloud storage
  • Large-volume data transfer detection
  • Clipboard and screen capture

Response quality

  • Alert fired versus alert seen versus alert actioned
  • Time from execution to analyst acknowledgement
  • Investigation quality and escalation decisions
  • Out-of-hours coverage
  • Containment action availability and speed

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Technique variation matters more than technique coverage: a rule that catches one implementation of credential dumping frequently misses three others. Agents generate variants so the rule is tested rather than the tool.

  • Telemetry correlation across SIEM, EDR and platform logs for each execution is a join across large datasets, and doing it by hand slows the exercise to a crawl.

  • Coverage mapping against the full ATT&CK matrix, including sub-techniques, is bookkeeping that should not consume analyst time during a live exercise.

  • Re-execution after rule development is repetitive by design and benefits from automation, which is what makes same-day verification possible.

Execution stays under operator control and every technique is agreed in advance with the blue team present. Agents handle variation, correlation and bookkeeping; they do not decide what to run on a client network.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Gaps are diagnosed, not just recorded

Every miss is classified as no telemetry, no rule, or no action. Those have different owners and different costs, and 'not detected' on its own is not an actionable finding.

Rules written and verified during the exercise

Detections are developed and re-tested the same day rather than recommended for later. The client ends the engagement with working coverage, not a backlog.

Improvement is measured

Coverage before and after, with the same techniques re-run. The deliverable is a figure, not an opinion.

Both reports, always

Technical report and executive summary together, plus the rule artefacts in a portable format.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Technique set, selection rationale and the exercise window
  • Coverage before and after, per technique and per tactic
  • For each technique: telemetry present, rule fired, analyst actioned — with timings
  • Gap classification with owner and estimated effort for each
  • Rules developed during the exercise, in Sigma or the client's native format
  • Telemetry gaps requiring logging changes, scoped
  • Response quality observations, including alerts closed incorrectly

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • Detection coverage before and after, as two figures
  • Which attack stages the organisation can and cannot see
  • The three investments that would most improve detection
  • Whether the gap is tooling, configuration or staffing — they are usually not the same
  • Recommended cadence for revalidation
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A financial services firm with a SIEM and a recently hired detection engineer.

Finding
Of 32 techniques executed, 19 produced no alert. Classification showed only four were genuinely missing telemetry; eleven had the telemetry and no rule, and four fired rules that analysts had suppressed months earlier because of noise.
Recommendation
Write the eleven missing rules during the exercise, re-tune the four suppressed rules rather than leaving them off, and scope the four telemetry gaps as a logging project.
Outcome
Coverage moved from 41% to 84% within the exercise week. The suppressed-rule finding prompted a standing review of every suppression, which found nine more.

A manufacturer with an outsourced security operations provider.

Finding
Alerts fired correctly for 22 of 26 techniques. The provider acknowledged 19 and escalated 3. Median time from execution to acknowledgement was 47 minutes during business hours and over nine hours overnight, against a contracted 15 minutes.
Recommendation
Raise the response findings at the contract level with the evidence, and add out-of-hours coverage to the service agreement.
Outcome
The contract was renegotiated with measured evidence rather than impressions. The exercise is now repeated biannually as a contractual verification mechanism.

A healthcare provider building detection in-house after moving from a managed service.

Finding
Credential access techniques were entirely undetected. LSASS access telemetry was available in the EDR but was not forwarded to the SIEM, and the EDR's own detections were configured to report rather than block or alert.
Recommendation
Forward EDR telemetry to the SIEM, switch credential access detections to alerting, and build correlation rules combining EDR and directory telemetry.
Outcome
Forwarding configured during the exercise and rules written the same week. Credential access coverage went from zero to full for the tested technique set.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.