VAPT
VAPT of iOS Applications
iOS is a harder target than Android and that changes the assessment rather than shortening it. Sandboxing, code signing and a curated store remove several classes of issue, but they also encourage teams to assume the platform is doing work it is not. Keychain items with the wrong accessibility class, App Transport Security exceptions added during development and never removed, and authorisation enforced only in the client are the findings this assessment exists to catch.
Methodology
- 01
Binary and entitlement analysis
The IPA is unpacked and Info.plist, entitlements, ATS exceptions, URL schemes, associated domains, minimum OS version and the presence of PIE, stack canaries and ARC are recorded.
- 02
Static analysis
Class dump and disassembly, then review for hardcoded secrets, weak cryptography, insecure file paths, debug symbols and third-party SDK versions with known issues.
- 03
Device storage assessment
On a jailbroken device the app container is examined — Keychain, NSUserDefaults, Core Data, plists, caches, snapshots and logs — for credentials, tokens and PII, with particular attention to Keychain accessibility classes.
- 04
Dynamic analysis and traffic interception
Traffic proxied with a trusted CA. Where pinning is present it is bypassed at runtime, because the question is what the API does once pinning is defeated.
- 05
Runtime manipulation
Frida and Objection used to hook Objective-C and Swift methods, defeat jailbreak detection, bypass local biometric authentication and test whether authorisation is enforced server-side.
- 06
API and backend testing
Every endpoint the app calls tested independently for authentication, authorisation, IDOR and rate limiting.
- 07
Exploitation and evidence
Confirmed issues exploited on device and captured as reproducible steps with evidence.
- 08
Reporting and retest
Technical report and executive summary issued together, then a retest evidencing closure.
Approach to testing
- Testing runs on a jailbroken device for reach and a stock device for realism. Reporting only the jailbroken result overstates risk; reporting only the stock result hides it.
- The current iOS release and the app's declared minimum are both covered, because a platform control the app relies on may not exist on the oldest version it still supports.
- Keychain items are assessed by accessibility class, not merely by presence. An item stored with kSecAttrAccessibleAlways is available from a backup and from a locked device.
- Biometric authentication is tested for server-side enforcement. LocalAuthentication returning success is a UI event, and treating it as an authorisation decision is a common and severe finding.
- Grey box by default, with a working build distributed through TestFlight or an enterprise profile and credentials for at least two roles.
Types of assessment
Black box
Store build only, no credentials. Mirrors an attacker with a downloaded app. Limited by the platform, and limited in what it finds.
Grey box (default)
TestFlight or enterprise build, credentials for multiple roles, API documentation. Best coverage per unit of effort.
White box
Full Swift or Objective-C source and build pipeline. Adds review of cryptographic implementation and authorisation logic.
Resilience assessment
MASVS-RESILIENCE only: jailbreak detection, anti-debugging, obfuscation and integrity checks, measured as time-to-defeat. Relevant for payment and licensed content apps.
Frameworks and standards
- OWASP MASVS 2.x
- The verification standard scored against — STORAGE, CRYPTO, AUTH, NETWORK, PLATFORM, CODE, RESILIENCE.
- OWASP MASTG
- iOS-specific test procedures, so each control maps to a documented test.
- OWASP Mobile Top 10
- Executive communication.
- Apple Platform Security Guide
- The baseline for what the platform actually guarantees, which is narrower than teams assume.
- NIST SP 800-115 / PTES
- Testing lifecycle and exploitation standard.
- CWE / CVSS v3.1
- Classification and severity scoring.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Frida
Runtime hooking of Objective-C and Swift, pinning and jailbreak detection bypass, method tracing.
Objection
Keychain dumping, file system inspection and runtime exploration without bespoke scripts.
class-dump / Hopper
Header recovery and disassembly for static review.
Burp Suite Professional
Interception and API testing behind the app.
MobSF
First-pass static analysis and entitlement baseline.
Keychain-Dumper
Enumerating Keychain items and their accessibility classes on a jailbroken device.
otool / codesign
Binary protections and signing verification.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Storage (MASVS-STORAGE)
- Keychain accessibility class on every stored item
- Credentials or tokens in NSUserDefaults, plists or Core Data
- Data Protection class on files written by the app
- Sensitive data in snapshots taken on backgrounding
- Data included in iTunes and iCloud backups
- Sensitive data in device logs and crash reports
Cryptography (MASVS-CRYPTO)
- Hardcoded keys or IVs in the binary
- Use of CommonCrypto with insecure modes
- Custom cryptographic routines
- Random values from a non-cryptographic source
Authentication (MASVS-AUTH)
- LocalAuthentication result enforced server-side, not only in the UI
- Session token lifetime, rotation and invalidation
- Role separation tested across at least two accounts
- IDOR on every identifier the app transmits
Network (MASVS-NETWORK)
- App Transport Security exceptions and their justification
- Certificate pinning presence and behaviour once defeated
- TLS version and certificate validation
- Sensitive data in URLs or headers
Platform (MASVS-PLATFORM)
- Custom URL scheme handling and parameter validation
- Universal links and associated domain configuration
- WKWebView configuration and JavaScript bridges
- Pasteboard exposure of sensitive values
- Keyboard cache on sensitive fields
- Inter-app communication and extension handling
Code and resilience (MASVS-CODE, MASVS-RESILIENCE)
- PIE, stack canaries and ARC present in the release binary
- Debug symbols and logging in a production build
- Third-party SDK inventory against known issues
- Jailbreak and debugger detection, measured as time-to-defeat
- Integrity verification and response to a modified binary
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Keychain accessibility classes are easy to get wrong and invisible in the UI. Agents enumerate every stored item and its class exhaustively rather than sampling the obvious ones.
Snapshot exposure on backgrounding happens for a fraction of a second on specific screens. Driving every screen through a background transition is machine work.
URL scheme and universal link parameter space is combinatorial; enumerating it by hand misses the handler that accepts an unvalidated deep link.
Agent findings and static analysis output are cross-checked, so an unreproduced scanner result is discarded and a manual finding is not assumed unique.
Every agent finding is validated by a human tester on a device before it reaches the report. The swarm decides what to look at; it does not decide what is true.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Exploited, not scanned
Findings are demonstrated on a device with reproducible steps. Where a control holds, that is recorded as a positive result rather than dropped.
Both reports, always
Technical report and executive summary together, written separately for two audiences.
Fixation, not a backlog
Remediation worked with the development team at code level, then retested to evidence closure.
The API is in scope
Most severe mobile findings are server-side. The endpoints behind the app are tested as first-class targets, not as context.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Scope, build version, device and iOS matrix, and the test window
- Every finding with CVSS v3.1 vector, CWE reference and MASVS control
- Reproduction steps a developer can follow without the tester present
- Evidence: screenshots, request and response pairs, Frida scripts used
- Code-level remediation guidance specific to Swift or Objective-C
- Controls tested that held
- Retest results appended against each original finding
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Risk position in business terms
- Severity distribution and movement since the last assessment
- The three things that most need funding
- Regulatory exposure where relevant (DPDP Act, PCI DSS, RBI)
- Remediation timeline and retest date
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A wealth management app for a mid-sized brokerage, around 60,000 users.
- Finding
- The session token was stored in the Keychain with kSecAttrAccessibleAlways, so it was readable from an unencrypted backup and from a locked device. Combined with a token lifetime of 90 days, a single backup extraction gave long-term account access.
- Recommendation
- Move to kSecAttrAccessibleWhenUnlockedThisDeviceOnly, reduce refresh token lifetime, and bind tokens to a device identifier.
- Outcome
- All three implemented and verified. Device binding also closed token reuse from a restored backup onto a different handset.
A hospital appointment app used by patients and clinicians on the same binary.
- Finding
- Clinician features were hidden in the UI based on a role flag returned at login, but the underlying API did not check role. A patient account could call clinician endpoints directly and read any patient's record.
- Recommendation
- Enforce role server-side on every endpoint. Treat UI gating as presentation only and add authorisation tests to the API test suite.
- Outcome
- Server-side enforcement added across all endpoints before the next release. The API test suite now fails the build if an endpoint lacks an authorisation assertion.
A retail loyalty app with in-app payment, roughly 300,000 installs.
- Finding
- An App Transport Security exception allowing arbitrary loads had been added during development against a local endpoint and shipped to production. Combined with no certificate pinning, all traffic was interceptable on a hostile network.
- Recommendation
- Remove the exception, add pinning for the payment domain, and add a release check that fails the build if NSAllowsArbitraryLoads is set.
- Outcome
- Exception removed and the build check added, which is the part that prevents recurrence.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.