Auditing Our Own Monetization Before Store Submission
A payment bypass, a CORS hole, a fail-open cache, unbounded LLM spend and a CI pipeline that was green for the wrong reasons.
Executive summary
Before submitting NihonGO! to app stores we ran a full audit of the app — money paths first, then auth, then reliability. What we found was humbling: a client-trusted subscription endpoint, receipt verification that approved anything, an unauthenticated speech endpoint, a missing Redis error listener that could crash the backend, and a test suite that was green for the wrong reasons. This is the full list, with what we changed in each case.
The Challenge
An AI tutoring app has three valuable attack surfaces: subscriptions, expensive endpoints (speech-to-text), and conversation data. None had been threat-modeled. The goal was not a theoretical review — it was to make every revenue path server-verified and every expensive operation bounded before strangers installed the app.
Key problems
- 1Subscription upgrade accepted a client-sent tier and receipt checks returned verified for any input
- 2CORS allowed any subdomain containing .supabase.co with credentials
- 3The most expensive endpoint (speech-to-text) had no authentication or rate limits
- 4Missing Redis error listener crashed the backend when the cache went down
- 5The readiness probe returned 503 forever against a misspelled table name
- 6An AI SDK update renamed maxTokens to maxOutputTokens — silently making model spend unbounded
- 7Subscriptions never expired because the expiration job was never scheduled
- 8Streaks computed in UTC broke learners in UTC+9
Constraints
- !Fix + regression test each finding before submission — no silent patches
- !Keep the existing architecture; no rewrite
- !Store review timeline: fixes had to land in days, not weeks
Threat-Model the Money First
The audit ran in one order: money, then auth, then reliability, then honesty of the test suite. Every fix shipped with a test that fails if the bug returns, and every "temporary" mitigation was replaced with a structural one before submission.
Make every payment state server-authoritative
The upgrade endpoint now derives the tier from a verified receipt on the server; client input is treated as a purchase token, never as a tier. Expiration runs on a real schedule, and downgrades are exercised in tests.
- Server-side receipt verification with provider APIs
- Subscription expiration job wired and monitored
- Entitlement checks covered by tests for upgrade, renewal and cancellation
Close auth gaps on expensive paths
Speech-to-text required no auth. It now requires a valid token, carries a per-user rate limit, and logs usage so cost anomalies are visible before the bill arrives.
- Auth + per-user rate limits on STT and LLM endpoints
- Usage logging with dashboards for spend anomalies
- CORS allowlist replaces substring matching
Fail loud, not open
The Redis client had no error listener — an outage crashed the process; worse, cached entitlement paths could fail open. Errors now degrade to database reads and alert, never to "grant access".
- Redis error listeners + circuit-breaker behavior
- Readiness probe fixed to the real table name
- Fail-closed behavior tested by killing Redis in integration tests
Make CI honest again
Tests were "green" because failing specs were being skipped or mocked incorrectly — one suite returned 401 for every case due to a missing auth mock. After unmasking, 141 analyzer issues and 32 failing tests surfaced; all are now fixed and enforced.
- Unmasked failing tests and lint: 141 analyzer issues → 0
- 263 backend tests passing (from 25 failing)
- 186 Flutter tests passing (from 7 failing)
Key technical decisions
Server-derived entitlements only
Any monetization path where the client sends the tier is a bypass waiting to happen. The receipt is the only input that matters, and it must be verified against the provider.
Fail closed on cache outages
A missing Redis listener crashed the process; the naive fix (ignore errors) would have failed open on entitlements. Degrading to the database is slower and correct.
Timezone-aware streaks, computed once
Streaks in UTC reset at 9 AM for Japanese learners. Compute in the user timezone at write time and store the resulting date — not the rule.
Implementation
Project timeline
Audit (days 1–3)
- Threat-model payments, STT and conversation data
- Reproduce each issue with an automated test or script
- Rank fixes by revenue risk
Hardening (days 4–8)
- Server-side receipt verification and entitlements
- CORS allowlist, rate limits, Redis error handling
- Cron for subscription expiration; timezone streaks
- CI unmasking and lint/test fixes
Verification (days 9–10)
- Device testing on real iOS/Android builds
- Cost anomaly dashboard check
- Submission checklist
Implementation challenges
The AI SDK renamed maxTokens to maxOutputTokens in a minor version — builds stayed green while model spend went unbounded.
Solution: Upgraded with a pinned version, added explicit token limits at the client wrapper, and a usage alert that fires when a single request exceeds its budget.
Accessibility claimed WCAG contrast compliance via a helper whose pow() implementation was this * this.
Solution: Replaced with a correct relative-luminance implementation and a test asserting known contrast pairs.
GoogleFonts CJK fonts were fetched at runtime; offline learners saw tofu boxes.
Solution: Bundled the required CJK subsets and removed the runtime fetch for offline-critical screens.
Results & Impact
Key achievements
Business impact
Lessons Learned
The threat model starts with money
The highest-impact findings were all in revenue paths. Audit the endpoints that make or spend money first; polish follows.
Green CI is a hypothesis, not evidence
Skipped tests and broken mocks made the suite green while 32 tests were failing in reality. Periodically run the suite with skips disabled and mocks off.
Fail-open is a bug, not a fallback
Convenience error handling on cache and auth paths silently grants access. Decide fail-open vs fail-closed per path and test the failure.
Timezones are user-visible bugs
A streak that resets at 9 AM local time is a churn reason. Store user-local results, not UTC assumptions.
What we’d do differently
- Wire a real payment-provider sandbox into CI
- Add synthetic cost-spike alerts per endpoint
- Schedule quarterly re-audits of revenue paths
- Write the missing integration tests for JLPT and shadowing modules
Technologies used
The project
See the full project
More details, screenshots, and information about NihonGO! AI Japanese Tutor.
View project