Seller Portal · PIM · LiveOps · BRD v1.2
A seller uploads a row. Merit has to answer one question: is this a product we already have, or a new one? Today a person answers it. This is the plan to stop that, without shipping a single wrong link.
Merit cannot tell whether an uploaded product already exists, so a human decides, and the number of humans required grows with the catalogue.
Today the system compares product titles as exact strings. iPhone 17 Pro Max blue 256 and iPhone 17 Pro Max 256 blue are the same product, and the system reports no match. Every miss goes to a LiveOps reviewer, who opens the row, searches the catalogue by hand, and decides.
Ghaith stated the business consequence on 18 August: the volume of that manual review is what decides whether Ops needs another head. This is a cost question before it is a quality question.
What the system does today
The same phone, described in a different word order. Every one of these becomes manual work, and the volume of it is what decides whether Ops needs another head.
0
Confidence scores anywhere in the current flow. A near-certain match and a wild guess arrive in the same queue, looking identical
3
Numbers Merit has never measured: review volume, time per review, and how often the answer is "yes, same product"
7
Microphone variants in Ghaith's example. Three have errors. Today the risk is that all seven are rejected
Three pressures arrived together. Seller onboarding is moving into the Seller Portal and the manual Merchant Readiness Form is being retired. The Online Catalogue sync is being cut back, so the Seller Portal becomes the main way products enter the catalogue. And LiveOps capacity is already the constraint.
Merit separates two things. Everything below follows from that separation.
Product
The global definition of a thing. "iPhone 17 Pro Max, 256GB, Blue Titanium." One record, shared by every seller.
Offer
One seller's listing against that product, carrying their price, their stock and their shipping terms.
A seller never creates a Product. A seller always creates an Offer. The system decides whether that Offer attaches to a Product that already exists, or triggers a new one. That decision is the entire subject of this document.
One product, many offers
This structure is what makes price comparison and the Buy Box possible. It also breaks the moment a seller creates a second Product for a phone Merit already sells, because the offers then compete on two separate pages.
Merit has no dependable GTIN, EAN or UPC across the catalogue. There is no barcode to match on. The identity decision has to be made from the text and attributes the seller supplied, and nothing else.
Deterministic checks bracket the model on both sides. The model is the middle, not the decision.
One row, end to end
Deterministic checks bracket the model on both sides. Steps 1 and 2 are rules, step 4 is a rule, and only step 3 is probabilistic. That is what keeps a confidence score from ever being the thing that decides.
The model never rejects. It routes. Rejection is a deterministic rule before the model, or a human decision after it. Nothing is ever refused because of a score, and no seller-facing message says the model rejected an item.
This is the requirement most likely to be dropped in implementation, and it prevents the failure nobody sees until a customer complains.
A 256GB iPhone and a 512GB iPhone score above 90 percent on name similarity. They are not the same product. If storage is not checked exactly, a high score links the offer, and the customer who ordered a 512GB unit receives a 256GB one.
So each category names a small set of identity-defining attributes. An offer auto-links only when every one of them is present on both sides and equal, whatever the score says. For Smartphones that set is storage, colour and region: exactly the three variant axes configured for the category.
Why the score alone is dangerous
Two records that a similarity model reads as almost the same thing. One attribute separates them, and it is the one the customer paid for.
A missing identity attribute is treated exactly like a wrong one. A null must never read as "nothing conflicts". A comparison that only guards against mismatches will happily auto-link at 94 percent when the deciding attribute is simply absent. That is the most common shape of this failure, not the rarest.
Move the score, then change what the attributes are doing. The score on its own never decides the outcome.
Recorded, not linked
Move the controls to see how a row is routed.
At 94 percent with one differs, that is the 256GB against the 512GB, and the attribute check overrides the score. Switch the stage to live to see what changes once a category has earned auto-linking.
So the design bias throughout is stated plainly: when in doubt, queue it. A queue is recoverable. A mislinked offer that shipped is not. Precision outranks automation rate, and if precision falls the correct response is to widen the queue, not to protect the automation number.
The service launches scoring everything and linking nothing.
In shadow mode the model scores every row and records what it would have done. LiveOps still decides everything. Each of those pairs, the model's suggestion against the human's actual decision, becomes the evidence that calibrates the thresholds.
Merit exits shadow mode per category, on evidence, not on a date. A clean category is not held back by a noisy one, and a category whose precision drops goes back into shadow without affecting the others.
How a category earns the right to auto-link
The pairs are the whole point. Until the model's suggestion has been compared against enough real reviewer decisions, a confidence threshold is a number somebody guessed.
Two existing documents disagree on the confidence bands: 90 and 60 percent in one, 95 and 70 percent in the other. Neither set is calibrated against real data. A number fixed before shadow mode gains authority it has not earned. So the thresholds stay configurable per category, and the exit bar is measured, never declared.
Two things will make the first weeks harder than the design suggests. Both are known now, and neither is a tuning problem.
Merit holds one real seller catalogue export: a cosmetics file, 11,425 rows, 221 columns. One seller in one vertical, so treat it as a warning and not as a measurement. It still contradicts part of the assumed input.
Field coverage in the one real seller catalogue · 11,425 rows
The four fields the matching design leans on hardest are the four at the bottom. Brand and GTIN are not sparse, they are absent: the columns do not exist in the file. Colour exists only as 44 category-specific fields, and the generic colour column is empty in every row.
Two consequences. Brand cannot be treated as a reliable input, so brand normalisation becomes part of the ingest, not an assumption about it. And an identity-defining attribute may not live in one column, so mapping a seller's layout onto the PIM category schema is work that happens before matching. Nobody has scoped it or owns it.
Normalising all 11,411 distinct product names in that file for word order and punctuation collapsed ten of them.
What word-order normalisation actually recovers
The duplication that costs Merit money is between a seller's file and Merit's catalogue, and between two different sellers. Neither is visible inside a single seller's export, which is why the coverage count matters more than this one does.
The blue 256 against 256 blue example is real and it is a good illustration. It is not where the volume is. The duplication that costs Merit money is between a seller's file and Merit's catalogue, and between different sellers, and this file cannot measure either.
region is one of the three identity attributes for Smartphones, and it does not exist in the PIM schema yet. Until it is added, every Smartphone row is missing an identity attribute, so every Smartphone row queues. Not partly. Completely.
No threshold work moves this, and there is no measurement to take, because the attribute is absent by construction and not by data quality. Smartphones is sequenced behind the schema change, which is due 26 August.
Every example in this document is written against Smartphones, because it is the clearest case. Pick the pilot category on attribute coverage instead, once that has been counted.
model_number is not identity-defining for Smartphones. The part number varies per regional variant by design, so two units that are the same product to a customer carry different part numbers, and gating on it would split one product into three.
It stays mandatory for Electronics data quality, and it stays a strong scoring signal. A missing part number lowers a score. It never blocks a link.
The business requirements are fixed. The PRD, the phasing and the delivery plan are Ruba's.
The delivery skeleton, and where the blocker sits
Phase 2 opens per category, so one clean category can graduate while a noisy one stays in shadow. Smartphones is the category every example is written against, and it is the one that has to wait.
Each one needs the model approach chosen first, so the BRD states what the output must be and stays silent on how to produce it.
PRD for the matching service, and a phasing and delivery plan, both proposed for 3 September. Plus the attribute coverage count, and an interface agreement with Tintash before the plan is final.
Source: BRD, AI Product Matching for Seller Portal Uploads. Written from the 18 August session with Ghaith Fakhouri, the 13 August AI initiative sync with Ghaith and Mawi Shahin, and Rinad Al Tarawneh's comment threads on the Seller Portal Upload PRD. The figures in section 06 come from a seller catalogue sample dated 29 January 2026, one seller in one vertical.