Without a GTIN, two almost identical offers can quietly become two different products.
A 500 ml bottle appears twice in a feed: same brand, same scent, nearly the same title—yet one says “new formula” and the other omits it. Treating them as duplicates may erase a genuine revision; treating them as separate items can split reviews, prices, and stock history.
That small call affects more than a tidy catalogue. Deduplication can hide competing offers, comparison pages can place unlike items side by side, and a colour or size variant may be attached to the wrong parent. A missing GTIN turns matching into an identity decision, not a simple text-similarity check. Stable clues such as manufacturer part numbers, pack count, dimensions, and consistent product attributes deserve more weight than a close-looking name alone.
- A changed pack size or formulation is often a separate sellable product, even when the title barely changes.
A missing GTIN is not one problem
GTIN
A GTIN normally ties a listing to a specific trade item: the exact product and, often, its sellable pack or size. It is strong evidence of identity, not proof that every feed field is correct.
Absent value
The field is empty or omitted. There is no identifier evidence to compare, so a match must rest on other manufacturer-supplied details.
Unusable value
A value may be truncated, fail its check digit, contain placeholder digits, or be copied into the wrong field. Treat it as suspect rather than merely missing.
Conflicting value
A syntactically valid GTIN can still point to a different colour, pack count, or model. A disagreement with trusted product data should block an automatic match.
Alternative identifier
A manufacturer part number is usually the best substitute when it is paired with brand and variant details. Model names and internal SKUs are weaker because they can be shared, renamed, or seller-specific.
An Amazon ASIN identifies a catalogue entry within an Amazon marketplace, not a product everywhere. The same physical item may have different records—or no record at all—across marketplaces. Check Amazon ASINs by country before treating an ASIN as match evidence.
A practical evidence order is: validated GTIN, then brand plus manufacturer part number, then model and tightly matching variant attributes. Seller SKUs, titles, and images can support that case, but should not create a match on their own.
Record why the GTIN is unavailable
-
Check whether the product should have a GTIN
Some handmade goods, bundles, spare parts, and older stock may never have received one. Mark these as “not assigned,” rather than treating the field as an extraction failure.
-
Separate a blank field from a bad value
A missing value, an all-zero placeholder, and a number with the wrong length need different review paths. Preserve the original feed value when it is safe to retain.
-
Test whether the GTIN belongs to another item
A valid-looking code may identify the parent product, a multipack, or a different variant. Compare it with the title, pack count, size, and brand before accepting it.
-
Look for feed or mapping failures
If a supplier has GTINs elsewhere but this feed does not, note “source omitted” or “mapping failed.” That flag helps distinguish a catalog issue from a product-level exception.
-
Store a clear reason code
Useful values include not assigned, unknown, invalid format, conflicting, source omitted, and pending supplier confirmation. Keep the reason beside the matching decision for later review.
“No GTIN supplied” only describes the feed. It does not prove that the product lacks an identifier.
When the cause is uncertain, retain that uncertainty in the record and use a cautious fallback match, such as manufacturer part number plus brand. A later supplier update can then be reprocessed without undoing a misleading classification.
Build identity from brand, MPN, and attributes
A normalized brand + normalized MPN is usually the best non-GTIN identity candidate. Normalization removes presentation noise without changing meaning: trim spaces, standardize case, collapse repeated punctuation, and map known brand aliases to one approved form. For example, ACME Tools, Acme-Tools, and acme tools can resolve to ACME TOOLS; XR-200/BLK and xr 200 blk may resolve to the same comparison key.
That key should not stand alone. Confirm it against specifications that ought to remain stable: model family, capacity, dimensions, material, voltage, or compatible device. Then use variant attributes—such as color, size, pack count, or regional plug—to distinguish sellable versions under the same base model. This also helps categorize products despite inconsistent feed attributes before matching rules become too broad.
A defensible match records the evidence rather than treating text similarity as proof:
- Brand and MPN agree after normalization.
- Core specifications agree within expected formatting differences.
- Variant values either agree or explain a deliberate parent–child relationship.
- Conflicting attributes send the record to review.
Merchant SKUs deserve much less trust. A retailer’s SKU-10482 may be an internal inventory code, recycled after a listing change, or shared by unrelated sellers. It can support a match within the same merchant and feed history, but rarely travels safely across sellers. Keep it as a source reference, not a cross-market identity.
A matching title is a lead, not proof
It may be a family name shared by different sizes, generations, or bundled versions.
Titles often omit the detail that distinguishes a sellable item, such as battery capacity or pack count.
Keep durable tokens: brand, model code, capacity, dimensions, color, and compatible device names.
Discard retailer-added noise such as “new,” “best price,” delivery claims, capitalization, and punctuation.
Fuzzy similarity is best used to assemble a review queue.
Similar wording can connect a base product to a variant, accessory, refill, or multipack. Confirm candidates against MPN and stable attributes.
A simple title cleanup can lowercase text, remove punctuation, and strip phrases such as “limited offer” or “free shipping.” Preserve model-like strings exactly where possible: XK-420, 2.5 L, and 12-pack often carry more identity than the surrounding title words.
When cleaned titles look alike, place the records side by side and check the attributes that define the sellable version before linking them.
Choose the right match level
Before comparing records, decide what a successful match represents. A product family can group all sizes or colours of one model; a sellable variant is one specific size, colour, flavour, or capacity; an exact offer also requires the same pack, condition, and sometimes seller-specific configuration. Treating these as interchangeable is a common source of false merges.
For most catalogue work, match at the sellable-variant level. A red 500 ml bottle and a blue 500 ml bottle may belong to the same family, but they should not become one purchasable item. Likewise, a two-pack is not the same offer as a single unit, even when its title only adds “2 pack.”
Make key attributes non-negotiable
Each category needs a short set of fields that must agree before records can merge. This also helps separate incomplete records from real variants rather than treating every difference as bad data.
- Groceries and household goods: flavour or scent, net quantity, and pack count.
- Apparel: size, colour, fit, and gender or age range where supplied.
- Electronics: model number, storage or capacity, connectivity version, and bundle contents.
- Tools and parts: compatible model, dimensions, thread or fitting type, and voltage.
If a must-agree value is absent, the records may still be linked as possible family members, but should not be merged. Pack count deserves special care: “6 × 330 ml” and “330 ml” describe different sellable units.
Move from candidates to confidence
-
1. Generate a small candidate set
Search within the same brand and product type, using normalized MPN first and stable title tokens second. The aim is a short list of plausible records, not an automatic match.
-
2. Put the evidence in comparable form
Standardize case, spacing, punctuation, units, and common abbreviations before comparing values. Split variant details—such as colour, size, capacity, or pack count—from family-level specifications.
-
3. Score supporting signals separately
Treat an exact normalized MPN as strong evidence; combine it with brand, key specifications, and variant fields. Titles can help rank candidates, but should not outweigh a conflicting part number.
-
4. Reject contradictions early
Discard a candidate when a decisive field disagrees: incompatible pack count, different capacity, another model code, or a clearly different variant. This discipline is a useful part of the wider cross-feed matching workflow, especially when feeds are messy.
-
Assign a review-friendly tier
Mark records as high confidence when identity and variant evidence align, medium when a human check is sensible, and low when only weak clues remain. Keep the evidence and rejection reasons with the decision.
A candidate list that is too broad usually signals weak category filtering or insufficient normalization.
A single score cutoff rarely travels well. An exact MPN may justify a high-confidence match in electronics, while apparel often needs size, colour, and style agreement as well.
Set tiers from known good and known bad examples in the catalog. Then adjust them when false merges are more costly than missed matches—or when the reverse is true.
Let certainty set the automation boundary
A confidence score should trigger an action, not merely describe a candidate. Auto-match only when the evidence is decisive: a normalized brand and MPN agree, required variant attributes align, and no conflicting identifier or specification remains.
Everything else belongs in one of two outcomes:
- Review queue: several plausible candidates, a missing but important variant field, or an unresolved brand/MPN formatting difference.
- Unmatched record: title-only similarity, contradictory attributes, or too little evidence to distinguish the item from its neighbors.
Make the review queue useful
Each queued record should carry the candidate IDs, the evidence that supported each one, the conflict or missing field, and the match level under consideration. A reviewer can then confirm a variant match, reject it, or mark the record unmatched without repeating the initial investigation.
It is tempting to force a weak record into the closest catalog entry to improve coverage. That choice creates false duplicate links, mixed specifications, and misleading comparisons downstream. An unmatched item is visible work; a wrong match is hidden damage that can spread through exports and reporting.
Review outcomes should feed back into the rules. If a particular MPN pattern or pack-count field repeatedly settles decisions, it is a candidate for a future automated check—not a reason to lower today’s threshold.
Treat unmatched as a retained state, not a discard pile. Keep its normalized fields, reason code, and checked candidates so it can be matched later when a better feed arrives.
Run the fallback process as a feedback loop
-
Log the evidence behind every decision
Store the matched record, rule version, supporting fields, conflicts, confidence tier, and whether the result was automated or reviewed. A short reason such as “brand + normalized MPN; color conflict absent” makes later corrections traceable.
-
Sample automated matches regularly
Pull a small random set from each confidence tier and category. Check both false merges and missed matches; a rule can look accurate overall while failing on a common product type.
-
Track errors by rule, not just by record
Tag corrections to the rule or field that caused them: reused MPNs, stripped pack counts, overly broad title tokens, or unreliable brand aliases. Repeated tags reveal where a rule needs narrowing or a category exception.
-
Review rules after feed changes
New suppliers, revised exports, and renamed attributes can quietly change field quality. Compare match rates, review rates, and correction rates before and after a rule update.
-
Retire weak shortcuts
If a signal repeatedly creates costly errors, lower its confidence or remove it from automation. Keeping a record unmatched is often safer than preserving an elegant but brittle rule.
- A correction log is useful only when it points back to a specific rule or missing field.
- Sampling low-volume categories matters; small segments can hide high error rates.
Fallback matching improves through evidence, sampling, and correction, not by making rules steadily broader. Each reviewed error should strengthen a condition, add an exception, or confirm that the record belongs in the unmatched queue.












