Skip to content
CyberSmithSECURE
Under Attack

VAPT

VAPT of Android Applications

An Android application runs on hardware the attacker owns. Every control implemented on the device — root detection, certificate pinning, obfuscation, local authentication — is a speed bump on someone else's machine, and the only question worth asking is how long it holds. Testing therefore covers the app, the device it runs on, and the APIs it talks to, because a finding in any one of them is a finding in the product.

Methodology

  1. 01

    Reconnaissance and build analysis

    The APK or AAB is unpacked and the manifest, SDK levels, permissions, exported components, network security configuration and signing scheme are recorded. This establishes what the app is permitted to do before any testing of what it actually does.

  2. 02

    Static analysis (SAST)

    Decompilation to Java and Smali, then review for hardcoded secrets, weak cryptography, insecure storage paths, debug flags, backup settings, exported providers and unsafe WebView configuration. Automated scanning is the first pass, not the finding.

  3. 03

    Device and storage assessment

    The app is exercised on a rooted device and every artefact it writes is examined — shared preferences, SQLite databases, Realm stores, the keystore, cache, logs and external storage — for credentials, tokens, PII and session material left in the clear.

  4. 04

    Dynamic analysis and traffic interception

    Traffic is proxied with a user-installed CA. Where certificate pinning is present it is bypassed at runtime, because the assessment question is not whether pinning exists but whether the API behind it is safe once it is defeated.

  5. 05

    Runtime manipulation

    Frida and Objection are used to hook methods, defeat root and emulator detection, force branch outcomes, and test whether authorisation and business logic are enforced on the server or merely on the client.

  6. 06

    API and backend testing

    Every endpoint the app calls is tested in its own right for authentication, authorisation, IDOR, mass assignment and rate limiting. Most severe mobile findings are server-side findings reached through the app.

  7. 07

    Exploitation and evidence

    Confirmed issues are exploited to demonstrate real impact, and captured as reproducible steps with screenshots, request and response pairs and, where relevant, a proof-of-concept build.

  8. 08

    Reporting and retest

    Technical report and executive summary are issued together, followed by a retest after remediation that evidences closure rather than assuming it.

Approach to testing

  • Testing runs against a real device and an emulator. Emulator-only testing misses hardware-backed keystore behaviour and device-specific storage; device-only testing makes instrumentation slower and some conditions harder to force.
  • Both a rooted and an unrooted device are used. Rooted answers what an attacker can reach; unrooted answers what a normal user is exposed to. Reporting only the rooted result overstates risk and reporting only the unrooted result hides it.
  • The most recent supported Android release and the app's declared minimum are both covered, because a control introduced in a newer API level may simply be absent on the oldest version the app still allows.
  • Grey box by default. CSS is given a working build and test credentials for at least two roles, which is what makes authorisation testing between roles possible at all.
  • Client-side controls are treated as findings only where they are the sole control. Where a server-side equivalent exists, the client-side bypass is recorded as context rather than as an issue in its own right.

Types of assessment

Black box

Production APK only, no credentials, no source. Mirrors an attacker who has downloaded the app from the store. Cheapest, and finds the least.

Grey box (default)

Working build, test credentials for multiple roles, API documentation. The best coverage per unit of effort, and the only way to test authorisation properly.

White box

Full source, build pipeline and backend access. Adds code-level review of cryptography and authorisation logic, and finds classes of issue that runtime testing cannot reach.

Resilience assessment

MASVS-RESILIENCE only: root detection, tamper detection, obfuscation and anti-instrumentation, measured as time-to-defeat rather than present or absent. Relevant where the app handles payments or licensed content.

Frameworks and standards

OWASP MASVS 2.x
The verification standard the assessment is scored against. MASVS-STORAGE, CRYPTO, AUTH, NETWORK, PLATFORM, CODE and RESILIENCE.
OWASP MASTG
The corresponding test procedures, so each MASVS control maps to a documented test rather than to tester preference.
OWASP Mobile Top 10
Used for executive communication, because it is the list a board is likely to have heard of.
NIST SP 800-115
The technical testing lifecycle the engagement runs on — planning, discovery, attack, reporting.
PTES
Execution standard for the exploitation and post-exploitation phases.
CWE / CVSS v3.1
Weakness classification and severity scoring, so findings are comparable between engagements and between vendors.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

MobSF

First-pass static and dynamic analysis, and the manifest baseline.

jadx / apktool

Decompilation to Java and Smali, and repackaging for patched builds.

Frida

Runtime hooking: pinning bypass, root detection bypass, method tracing and forced returns.

Objection

Frida-backed runtime exploration, keystore dumping and storage inspection without bespoke scripts.

Burp Suite Professional

Interception, request manipulation and API testing behind the app.

Drozer

IPC and exported component testing — activities, services, broadcast receivers and content providers.

adb

Device interaction, logcat review, backup extraction and file system access.

Semgrep

Rule-based source review where the engagement is white box.

apksigner / apkleaks

Signing scheme verification and secret discovery in packaged builds.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Storage (MASVS-STORAGE)

  • Credentials, tokens or PII in shared preferences, SQLite or Realm
  • Sensitive data written to external storage or cache
  • android:allowBackup and the resulting extractable data set
  • Keystore usage, key invalidation on biometric change, hardware backing
  • Sensitive data in logcat, crash reports and third-party analytics
  • Screenshot and task-switcher exposure of sensitive screens

Cryptography (MASVS-CRYPTO)

  • Hardcoded keys, IVs or salts in the package
  • ECB mode, static IVs, or a cipher chosen without authentication
  • Custom or home-rolled cryptographic routines
  • Random number generation from a predictable source

Authentication and authorisation (MASVS-AUTH)

  • Session token lifetime, rotation and invalidation on logout
  • Biometric and local authentication enforced server-side, not only in the UI
  • Role separation tested between at least two accounts
  • IDOR across every object identifier the app sends
  • Step-up authentication for sensitive operations

Network (MASVS-NETWORK)

  • TLS version, cipher suites and certificate validation
  • Certificate pinning presence, and behaviour once defeated
  • Cleartext traffic permitted by the network security configuration
  • Sensitive data in URLs, query strings or headers

Platform interaction (MASVS-PLATFORM)

  • Exported activities, services, receivers and content providers
  • Deep link and app link handling, including unvalidated parameters
  • WebView configuration: JavaScript, file access, addJavascriptInterface
  • Permission set against actual functional need
  • Clipboard, keyboard cache and accessibility service exposure

Code quality and resilience (MASVS-CODE, MASVS-RESILIENCE)

  • Debuggable flag and debug artefacts in a production build
  • Third-party SDK inventory and known vulnerable versions
  • Root, emulator and hooking detection, measured as time-to-defeat
  • Obfuscation coverage of security-relevant code paths
  • Integrity verification and response to a repackaged build

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • A single tester works one hypothesis at a time. Parallel agents drive storage, network, IPC and runtime tracks simultaneously against the same build, so an artefact written during one interaction is caught while another track is still exercising the UI.

  • Transient state — a token written to cache during a failed login, a key held in memory only between two screens — exists for seconds and is routinely missed by sequential manual testing.

  • Deep link and exported component space is combinatorial. Enumerating it exhaustively is machine work; judging which results matter is not.

  • Scanner output and manual findings are cross-checked against each other, so a MobSF result nobody reproduced is not reported, and a manual finding a scanner missed is not assumed unique.

Every agent finding is validated by a human tester before it reaches the report. The swarm decides what to look at; it does not decide what is true. Anything not reproduced by hand is discarded.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Exploited, not scanned

A MobSF warning is a starting point. Findings reach the report only when they have been exploited on a device with reproducible steps and evidence attached. Where a control cannot be bypassed, that is stated as a positive result rather than quietly dropped.

Both reports, always

A technical assessment report for the engineers and an executive summary for the people funding the fix, issued together. Most vendors produce one and let the other audience make do with it.

Fixation, not a backlog

CSS works remediation with the development team — not only what is wrong, but the code-level change — and retests to evidence closure. An assessment that ends at the report has moved the risk to a spreadsheet, not reduced it.

The API is in scope

Many mobile assessments stop at the client. The endpoints behind the app are tested as first-class targets, because that is where the severe findings usually are.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Scope, build hash, device and OS matrix, and the exact test window
  • Every finding with CVSS v3.1 vector, CWE reference and MASVS control
  • Reproduction steps precise enough for a developer to follow without the tester present
  • Evidence: screenshots, request and response pairs, decompiled excerpts, Frida scripts used
  • Code-level remediation guidance, not 'implement certificate pinning'
  • Controls tested that held, so the report records what is working
  • Retest results appended against each original finding

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • Risk position in business terms — what an attacker could do to the business, not to the binary
  • Severity distribution and how it compares to the previous assessment
  • The three things that most need funding, with an indication of effort
  • Regulatory and compliance exposure where relevant (DPDP Act, PCI DSS, RBI directions)
  • Remediation timeline and the retest date
  • One page. If it runs to five it will not be read by the person who signs the budget.

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A retail bank's customer app, roughly 2 million installs, tested grey box before a payments release.

Finding
Certificate pinning was implemented and defeated in under an hour with a standard Frida script. Behind it, the transaction endpoint authorised on a customer identifier supplied by the client, so changing it returned another customer's balance and transaction history.
Recommendation
Move authorisation server-side and derive the customer identifier from the session rather than the request body. Treat pinning as an attacker-delay control, not an access control, and stop relying on it.
Outcome
The IDOR was fixed before release and verified by retest. Pinning was retained but reclassified internally, which changed how the team reasoned about every subsequent client-side control.

A logistics operator's driver app, around 8,000 field devices on managed hardware.

Finding
The app wrote an OAuth refresh token to shared preferences in cleartext, and android:allowBackup was left enabled. Any device with USB debugging on — common across the fleet for support reasons — allowed the token to be extracted with adb and reused from anywhere.
Recommendation
Move token storage to the hardware-backed keystore, disable backup for the app, and bind refresh tokens to a device identifier so an extracted token is useless off the device.
Outcome
All three implemented. The device binding proved the more valuable of the changes, because it also closed token reuse from a lost or stolen handset.

A healthcare provider's patient app, tested against DPDP Act obligations.

Finding
An exported content provider, present to support an internal companion app that had been retired two years earlier, returned appointment records to any application installed on the device without a permission check.
Recommendation
Remove the provider. Where an export is genuinely required, protect it with a signature-level permission so only applications signed with the same key can reach it, and audit exported components on every release.
Outcome
Provider removed. The audit step was added to the release checklist, which subsequently caught a second exported service before it reached production.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.