Fayez sent an interactive simulator and a redrawn solution diagram, and asked for a walkthrough before either reaches Tintash. His own 16 tests all pass. Twelve further probes found five defects, three of which are the same defect wearing three faces.
VERIFY_ + 4 hexFayez wrote that the update "should address and resolve all the open questions." It closes four of thirteen and gives test evidence for a fifth. Six remain open, and one of those, prefill reliability across iOS, Android and WhatsApp Business, sits on the critical path. Tab 05 has the full accounting.
The simulator is not a mock-up. It carries a working engine, so the design can be executed rather than read. That is what makes it worth auditing properly.
| Step | What was done | Result |
|---|---|---|
| 1 | Extracted the engine class out of the HTML into a standalone Node module | 6,593 characters of logic, no DOM dependency |
| 2 | Re-ran Fayez's own 16 tests headless against that engine | 16 / 16 pass |
| 3 | Wrote 12 probes for cases the suite does not reach | 5 defects, 2 design questions, 1 nit |
The simulator is a model. Where it simplifies, the real system may already be stricter.
D1, D2, D3 and D5 are all visible in the diagram itself, so they hold either way. One finding, an off-by-one
where a message arriving at exactly expiresAt still verifies, is a property of the model only
and is recorded as a nit rather than a defect.
Three blockers and two high. The first three are one problem with three parts, and fixing any one of them defuses it. Taking all three is the safe answer.
An inbound message is matched by code alone. Nothing ties it to the device that asked for the code. So a code sent from any number verifies the session that issued it.
The victim's client is listening on the session document. When it flips to VERIFIED it signs in with the token attached, and that token belongs to the account just created for the attacker's number. Fix to decide: bind the session to the requesting client at request time and refuse an inbound whose session was opened elsewhere.
Section 4a caps the request side, 5 per number per hour and 20 per device or IP per hour, and the attempt side inside a session, 5 failures. Neither touches an inbound message that matches no session, which is exactly what a guess is.
The format is VERIFY_ plus four hex characters, so 164 codes.
Small enough to enumerate, and with D2 there is nothing stopping the enumeration. Tab 04 puts numbers on it.
Fayez flagged half of this himself, in the note under his test suite: the section 4a case 1 copy, "that code does not match, N attempts left", cannot be delivered in registration mode because there is no target number to compare the sender against. That is correct.
The half he did not flag is the consequence. A wrong code in registration mode never
resolves to a session, so attempts never increments and the cap of five never engages.
US.3.3 states the attempt cap as a property of the flow. It is a property of change-number mode only.
The verify path for change-number does no collision check. It releases the old number, then writes the new one over whatever was already there.
The other owner is gone, silently. Members recycle numbers and telecoms reissue them, so this will happen in production. Fix to decide: refuse the change when the target already resolves to an account, with copy that says so. A merge is the alternative and is far larger.
| Item | Behaviour | Why it needs a stated position |
|---|---|---|
| The expired reply is an oracle | A known but expired code gets a reply. An unknown code gets silence. | Any sender can tell a real Merit code from a fake one by whether anything comes back. The reply cap of 3 per number per day limits volume, not the leak. This may be an acceptable trade for the member experience. It should be a stated trade. |
| The old number is released and left free | After a change, the previous number maps to nothing and is claimable immediately. | If intended, put it in the PRD. A cooling period is the usual answer. |
Each of the three blockers is survivable alone. Together they compose into something a bored person with a script can run. These are the arithmetic consequences of the design as drawn.
A guess does not need to hit one specific member. It only needs to hit any live session. So the number that matters is not the size of the space, it is the size of the space divided by how many sessions are open at once.
Expected hits equals guesses times live sessions divided by 65,536, over one 5 minute window.
Answering an unknown code with silence is the right call against a bot fishing for confirmation that a live namespace exists. It was chosen for that reason and it should stay. The problem is that silence is also the only response, so nothing observes the guess, nothing counts it, and nothing ever escalates. Silence needs to be paired with a counter, otherwise it is indistinguishable from not looking.
Thirteen open questions were on the PRD. The updated diagram closes four and gives test evidence for a fifth. Six remain.
| Q | Question | The diagram's answer | Verdict |
|---|---|---|---|
| Q1 | Code lifetime | 5 minutes | Close 24 hours was inherited from a profile-change flow. This gates account creation. |
| Q3 | Drop phone entry | Registration takes the sender. Change-number keeps the entered target | Close Correct split. Removes the wrong-number case from registration. |
| Q4 | Client side or server side | identity-service mints the code, doc ID opaque | Close |
| Q5 | Expired versus unknown reply | Unknown gets silence, expired gets a reply | Conditional Right call, but see the oracle question and D2. |
| Q11 | RW-09 covered by a test | Test T10 asserts the number is never bound on failure or expiry | Close |
| Q6 | Architecture sign-off for Firestore and Cloud Functions | Not addressed | Open |
| Q7 | Prefill reliability across iOS, Android, WhatsApp Business | Not addressed | Critical path If prefill fails on a real share of devices, the manual path becomes primary and the copy changes. |
| Q9 | Phase 1 implementation estimate | Not addressed | Open |
| Q10 | Members may not understand they must send, not wait | Not addressed | Open Measured in dogfood. |
| Q12 | Storefront deferred | Not addressed | Open Revisit 15 September. |
| Q13 | Which path is primary once the Notification Hub lands | Not addressed | Open Phase 3. |
| Decision | Owner | By | Cost of deciding late |
|---|---|---|---|
| Bind the session to the requesting client, or accept the risk in writing | Brian with Fayez | Before Tintash starts | It becomes a rebuild of the verify path rather than a line on a diagram |
| Cap inbound guesses per sender, and count the silent ones | Fayez with Tintash | Before Tintash starts | Same, plus no data on whether it was ever attempted |
| Widen the code beyond four hex characters | Fayez | Before Tintash starts | Cheapest of the three. A format change after launch invalidates live sessions |
| Attempt cap for registration mode, or state that it has none | Brian | Before Tintash starts | The PRD currently claims a cap that does not exist in that mode |
| Refuse change-number onto a taken number, or define a merge | Brian with Fayez | Before Tintash starts | A merge designed after launch has live data to reconcile |
| Q7, prefill reliability across devices | Tintash | Critical path | If prefill fails widely, the manual path is primary and all the copy is wrong |
reverse_whatsapp_otp_simulator.html and
reverse_whatsapp_otp_simulator-updated.png, both sent by Fayez Hilow the same day, and read
against PRD B2C Reverse WhatsApp OTP v1.2 and the WhatsApp OTP Implementation minutes of 20 August.
Work-tree node b2c-identity. Findings are properties of the simulator engine, which is the
design expressed as code. Merit internal.