Merit · B2C Super App · Identity and Account · 20 August 2026

Reverse WhatsApp OTP: audit of Fayez's simulator and updated diagram

Fayez sent an interactive simulator and a redrawn solution diagram, and asked for a walkthrough before either reaches Tintash. His own 16 tests all pass. Twelve further probes found five defects, three of which are the same defect wearing three faces.

His test suite
16 / 16
all pass, run headless
Probes added
12
cases the suite does not cover
Defects found
5
3 blockers, 2 high
Code space
65,536
VERIFY_ + 4 hex
Inbound rate limit
None
100,000 guesses accepted
Recommendation. Hold the call Fayez asked for, and do not send the diagram to Tintash first. D1, D2 and D3 are one problem and should be tabled as one item. D5 is second, because it is the one that destroys an account. Tintash starts building the week of 25 August, so a diagram fix is free this week and expensive next.

What the diagram gets right

Where his message overstates it

Fayez wrote that the update "should address and resolve all the open questions." It closes four of thirteen and gives test evidence for a fifth. Six remain open, and one of those, prefill reliability across iOS, Android and WhatsApp Business, sits on the critical path. Tab 05 has the full accounting.

02 · Method

How it was tested

The simulator is not a mock-up. It carries a working engine, so the design can be executed rather than read. That is what makes it worth auditing properly.

The three steps

StepWhat was doneResult
1Extracted the engine class out of the HTML into a standalone Node module 6,593 characters of logic, no DOM dependency
2Re-ran Fayez's own 16 tests headless against that engine 16 / 16 pass
3Wrote 12 probes for cases the suite does not reach 5 defects, 2 design questions, 1 nit
Why a defect in the engine is a defect in the diagram. The engine is the design expressed as code. Every rule it enforces was drawn from the diagram, and every rule it omits is a rule the diagram never stated. So a probe that breaks the engine has found a hole in the design, not a bug in a demo.

What the audit deliberately does not claim

The simulator is a model. Where it simplifies, the real system may already be stricter. D1, D2, D3 and D5 are all visible in the diagram itself, so they hold either way. One finding, an off-by-one where a message arriving at exactly expiresAt still verifies, is a property of the model only and is recorded as a nit rather than a defect.

03 · Findings

The five defects

Three blockers and two high. The first three are one problem with three parts, and fixing any one of them defuses it. Taking all three is the safe answer.

D1Blocker

A guessed code hands a member's session to the attacker

An inbound message is matched by code alone. Nothing ties it to the device that asked for the code. So a code sent from any number verifies the session that issued it.

probe P2, registration mode
victimSession  VERIFIED
boundNumber    the attacker's number
tokenIssued    true
victim then signs in as an account that is not theirs

The victim's client is listening on the session document. When it flips to VERIFIED it signs in with the token attached, and that token belongs to the account just created for the attacker's number. Fix to decide: bind the session to the requesting client at request time and refuse an inbound whose session was opened elsewhere.

D2Blocker

No rate limit of any kind on inbound guesses

Section 4a caps the request side, 5 per number per hour and 20 per device or IP per hour, and the attempt side inside a session, 5 failures. Neither touches an inbound message that matches no session, which is exactly what a guess is.

probe P3   100,000 guesses in one run. Every one accepted, every one silent, no cap engaged.
D3Blocker

The code space is 65,536

The format is VERIFY_ plus four hex characters, so 164 codes. Small enough to enumerate, and with D2 there is nothing stopping the enumeration. Tab 04 puts numbers on it.

D4High

Registration mode has no attempt cap at all

Fayez flagged half of this himself, in the note under his test suite: the section 4a case 1 copy, "that code does not match, N attempts left", cannot be delivered in registration mode because there is no target number to compare the sender against. That is correct.

The half he did not flag is the consequence. A wrong code in registration mode never resolves to a session, so attempts never increments and the cap of five never engages.

probe P5   50 wrong codes at a live registration session  ›  status PENDING, attempts 0

US.3.3 states the attempt cap as a property of the flow. It is a property of change-number mode only.

D5High

Change-number onto an already-registered number destroys that account

The verify path for change-number does no collision check. It releases the old number, then writes the new one over whatever was already there.

probe P1, number A already belongs to another account
signedIn     true
users[A].uid  the account that ran the change
account count 2 before, 1 after

The other owner is gone, silently. Members recycle numbers and telecoms reissue them, so this will happen in production. Fix to decide: refuse the change when the target already resolves to an account, with copy that says so. A merge is the alternative and is far larger.

Two design questions, not defects

ItemBehaviourWhy it needs a stated position
The expired reply is an oracle A known but expired code gets a reply. An unknown code gets silence. Any sender can tell a real Merit code from a fake one by whether anything comes back. The reply cap of 3 per number per day limits volume, not the leak. This may be an acceptable trade for the member experience. It should be a stated trade.
The old number is released and left free After a change, the previous number maps to nothing and is claimable immediately. If intended, put it in the PRD. A cooling period is the usual answer.
04 · D1 plus D2 plus D3

The attack, in numbers

Each of the three blockers is survivable alone. Together they compose into something a bored person with a script can run. These are the arithmetic consequences of the design as drawn.

How much of the code space one 5 minute window reaches

65,536 possible codes. Coverage is guesses sent divided by the space, capped at 100 percent.
A session lives 5 minutes, so 220 guesses per second walks the entire space inside one window. Direct labels carry the values, so the chart does not depend on colour.

Expected hijacks per window

A guess does not need to hit one specific member. It only needs to hit any live session. So the number that matters is not the size of the space, it is the size of the space divided by how many sessions are open at once.

Expected hits equals guesses times live sessions divided by 65,536, over one 5 minute window.

Successful hijacks per window
2.3
3,000 guesses against 50 live sessions
What would make this wrong, and it is the first thing Fayez will say. WhatsApp itself throttles. One number cannot push 220 messages a second, so the high end of that slider needs a fleet of numbers, and that is a real cost to an attacker. It does not rescue the design. At 10 guesses a second, which one cheap sender can manage, a 50 session service still gives up roughly two accounts every five minutes, and the cost of the attempt is zero because an unmatched guess is met with silence and is never counted anywhere.

Why silence makes it worse, not better

Answering an unknown code with silence is the right call against a bot fishing for confirmation that a live namespace exists. It was chosen for that reason and it should stay. The problem is that silence is also the only response, so nothing observes the guess, nothing counts it, and nothing ever escalates. Silence needs to be paired with a counter, otherwise it is indistinguishable from not looking.

05 · Against PRD v1.2

What actually closes

Thirteen open questions were on the PRD. The updated diagram closes four and gives test evidence for a fifth. Six remain.

Closed by this artifact
5
Q1, Q3, Q4, Q5, Q11
Already closed
2
Q2, Q8
Still open
6
one on the critical path
QQuestionThe diagram's answerVerdict
Q1Code lifetime5 minutes Close 24 hours was inherited from a profile-change flow. This gates account creation.
Q3Drop phone entryRegistration takes the sender. Change-number keeps the entered target Close Correct split. Removes the wrong-number case from registration.
Q4Client side or server sideidentity-service mints the code, doc ID opaque Close
Q5Expired versus unknown replyUnknown gets silence, expired gets a reply Conditional Right call, but see the oracle question and D2.
Q11RW-09 covered by a testTest T10 asserts the number is never bound on failure or expiry Close
Q6Architecture sign-off for Firestore and Cloud FunctionsNot addressed Open
Q7Prefill reliability across iOS, Android, WhatsApp BusinessNot addressed Critical path If prefill fails on a real share of devices, the manual path becomes primary and the copy changes.
Q9Phase 1 implementation estimateNot addressed Open
Q10Members may not understand they must send, not waitNot addressed Open Measured in dogfood.
Q12Storefront deferredNot addressed Open Revisit 15 September.
Q13Which path is primary once the Notification Hub landsNot addressed Open Phase 3.
06 · Next

What to decide, and when

DecisionOwnerByCost of deciding late
Bind the session to the requesting client, or accept the risk in writingBrian with FayezBefore Tintash starts It becomes a rebuild of the verify path rather than a line on a diagram
Cap inbound guesses per sender, and count the silent onesFayez with TintashBefore Tintash starts Same, plus no data on whether it was ever attempted
Widen the code beyond four hex charactersFayezBefore Tintash starts Cheapest of the three. A format change after launch invalidates live sessions
Attempt cap for registration mode, or state that it has noneBrianBefore Tintash starts The PRD currently claims a cap that does not exist in that mode
Refuse change-number onto a taken number, or define a mergeBrian with FayezBefore Tintash starts A merge designed after launch has live data to reconcile
Q7, prefill reliability across devicesTintashCritical path If prefill fails widely, the manual path is primary and all the copy is wrong
The sequence. Call with Fayez first, diagram fixed second, Tintash third. The build starts the week of 25 August, which is the only reason any of this is urgent. Every item above is cheap this week.
Audit run 20 August 2026 against reverse_whatsapp_otp_simulator.html and reverse_whatsapp_otp_simulator-updated.png, both sent by Fayez Hilow the same day, and read against PRD B2C Reverse WhatsApp OTP v1.2 and the WhatsApp OTP Implementation minutes of 20 August. Work-tree node b2c-identity. Findings are properties of the simulator engine, which is the design expressed as code. Merit internal.