mindit.io logo
    • retail
      • / retail

        Retail industry, mindit.io

        In the intricate world of retail, mindit.io stands out as your strategic ally.

    • banking
      • / banking
        Woman using a card payment terminal, retail sector, mindit.io

        We partner with European banks on AI transformation, data-platform modernization, and regulatory-grade integration across Benelux, DACH, and UK.

    • financial services
      • / financial services

        Financial services: mindit.io data and AI

        Custom software and data platforms for asset managers, payments providers, and insurers across DACH and the US.

    • healthcare
      • / healthcare

        Doctor with digital tablet in hospital, healthcare sector, mindit.io

        ML and integration solutions for hospitals, health-tech platforms, and national health systems, including Switzerland’s national health sector.

    • hospitality
      • / hospitality
        Hospitality: mindit.io data and AI

        Operational and guest-experience software for hotels, restaurants, and global travel groups.

    • foodtech
      • / foodtech
        Burger and fresh salad, food and beverage sector, mindit.io

        Product engineering for the food industry: from plant-based configurators to supply-chain analytics.

    • manufacturing
      • / manufacturing
        Aerial view of industrial conveyor belt, manufacturing sector, mindit.io

        Industrial software, IoT integration, and data platforms for manufacturers modernizing operations.

    • publishing
      • / publishing
        Publishing: mindit.io AI solutions

        ML platforms and editorial workflow systems for publishers, including Izzard Ink Publishing.

    • real estate
      • / real estate
        Modern residential buildings at dusk, real estate sector, mindit.io

        Property tech and data analytics for real-estate operators and asset managers.

    • telco
      • / telco
        Person using smartphone, telecom sector, mindit.io

        Telecom-grade software for network optimization, customer-facing apps, and AI-driven insights.

    • / company
    • about us
      • / about us

        mindit.io partners and people
        The partner of choice for data & product engineering to drive business growth & deliver an impact within your organization
    • product engineering
      • / product engineering
        We specialize in Software Product Engineering, transforming your concepts into impactful products.
    • technology
      • / technology
        mindit.io AI Native company
        250+ specialists skilled in software, BI, integration, offering end-to-end services from research to ongoing maintenance.
    • methodology
      • / methodology
        Custom software solutions, mindit.io
        We specialize in software product engineering, transforming your concepts into impactful products.
    • careers
      • / careers
        Careers at mindit.io
        Our team needs one more awesome person, like you. Let’s grow together! Why not give it a try?
    • do good
      • / do good
        mindit.io team member do good ESG initiatives by the sea
        We’re a team devoted to making the world better with small acts. We get involved and always stand for kindness.
    • events
      • / events
        Optimising your Databricks spending. Tips and tricks for a well governed deployment
    • blog
      • / blog
        Building customer matching on Databricks for a luxury brand selling worldwide
        Databricks Bucharest User Group #3: What to Expect from the September 30 Meetup
    • contact us
      • / contact us
        Request a mindit.io webinar
        We would love to hear from you! We have offices and teams in Romania and Switzerland. How can we make your business thrive?
  • / get in touch

helping enterprises become AI-native organizations

Building customer matching on Databricks for a luxury brand selling worldwide

Building customer matching on Databricks for a luxury brand selling worldwide

A client buys two pieces online in March. In April, she is in a store in another country, asking about an alteration to something she bought the year before. She has shopped with the brand since 2019.

The sales associate helping her sees what the CRM holds, and nothing of the rest.

The web order sits in the ecommerce platform, the store purchase in the till system, and the alteration conversation in the CRM. Three systems, each doing its own job well, with nothing that says these records describe the same woman. She is the client whose history the store could not see.

We built the customer matching for this brand on Databricks. The brand set one condition above the rest: two records are joined only when the evidence is conclusive. The hardest cases are the records that look almost identical.

In brief:

  • A luxury brand’s clients are recorded across three source systems: online shop, store tills and CRM. Nothing connects the records that belong to the same person.
  • The matching finds records that could be the same client, scores how strong the case is, and lets independent evidence make the final call. It runs as one Databricks pipeline built with Lakeflow Declarative Pipelines, and what survives becomes one golden record per client.
  • A high score alone never merges two clients. The evidential value of an email or phone number depends on how many people use it.
  • Roughly half of the source records turn out to duplicate another record. Most are joined automatically; the ambiguous ones go to a reviewer.

Four realities a matching system inherits

Each channel captures what its moment allows. A till captures a phone number. A web registration captures a date of birth and a delivery address. Some clients give every detail asked of them and others give only a fragment, so across the whole base no single field is both well populated and able to tell two people apart.

Identifying yourself at a till is voluntary. In Korea, close to nine in ten in-store purchases are anonymous. Japan, six in ten. Europe, just under half. The United States, three in ten. That sets the ceiling on what any matching system can deliver.

The same name is written more than one way, correctly. The brand sells in markets that write in Korean, Japanese and the Latin alphabet. Family name comes first in some and last in others. One client, recorded in Hangul in Seoul, romanised at a web checkout and typed family-name-first in Europe, is three different strings. All three are right.

Contact details belong to more than one person. Households share a mailbox. Colleagues sit behind one office number. An assistant places orders for an executive. A matching system has to expect all of it.

The architecture

Records from the three systems land, are standardised, and flow through a single pipeline built with Lakeflow Declarative Pipelines. The pipeline produces two principal outputs: one golden record per client, which is the customer view the rest of the business reads, and a review queue holding the pairs that require human judgement. Two Databricks dashboards read the pipeline’s datasets under the same governance, one covering the quality of the source data and one covering the matching itself. Unity Catalog carries governance and the audit trail.

The matching pipeline on Databricks

The matching is implemented as a sequence of explicit stages, each represented by its own dataset: standardised records, shared-value counts, eligible participants, scored edges, clusters, and golden-record assignments. Lakeflow Declarative Pipelines manages the dependencies between these stages and processes the data incrementally.

Separating matching into explicit, persisted stages gives the solution four practical benefits:

  • Traceability. Every decision can be followed from the source records through eligibility, pair formation, scoring and clustering to the final golden record. The evidence and intermediate results remain available for investigation and audit.
  • Modularity. Each stage has a clear responsibility and its own dataset. Standardisation, counting, pair formation, scoring and clustering can therefore be changed or extended independently without redesigning the whole matching process.
  • Incremental processing. A run reprocesses only what the changed records can affect, leaving the rest of the customer base untouched.
  • Controlled evolution. Matching rules, configurable thresholds and clustering logic can evolve while the stages remain independently inspectable and previous results remain available. Combined with immutable golden-record identifiers, this lets the matching improve without losing its history or breaking identities already used downstream.

Lakeflow Declarative Pipelines resolved the dependencies between the stages and provided the framework for incremental processing. Spark carried the scale for pair comparison and the custom graph work. Delta held the persistent state and the change tracking. What remained to build was the matching itself: the rules, the identity model and the evidence logic. The solution reached production quickly as a result.

Splink does the comparing. It is an open-source library for probabilistic record linkage built on the Fellegi-Sunter method, and it runs on Spark inside the same pipeline. Splink performs the pairwise comparison and scoring.

The pipeline around Splink decides the rest: which records are eligible to participate, whether a shared email can be trusted as evidence at all, what a score is permitted to decide, how matches are clustered, and which identifier a client keeps from one run to the next.

Splink learns its weights from a representative sample of the brand’s own records: how often each field agrees between two records that are the same person, and how often it agrees between two records that are not. Those two rates turn the five field comparisons into one probability.

Runs are incremental. Change Data Feed carries only changed source rows into the process, so a run scores only the pairs those rows can affect. Each edge is keyed by the two records it joins, so re-scoring an unchanged pair writes nothing, and the same input produces the same clusters regardless of row arrival order. A disputed result can therefore be investigated and reproduced.

The review queue is one of the pipeline’s own datasets, and a reviewer’s decision comes back through the same incremental path as a changed source record. Review is a stage of the pipeline, with its own dataset and its own way back in.

From a raw record to a golden record

Eight steps, every run.

1. Standardise and cleanse. The same client is written differently in every system, and those records cannot be compared as they arrive. A set of Python helper modules rewrites every field into a consistent form: names into a common representation, including romanisation where required; phone numbers into one international format; addresses through one set of abbreviations; countries into one code; and dates of birth checked against the placeholder dates systems use when a real value is unknown.

The same pass identifies values that cannot reliably identify anyone, including disposable email domains and generic mailboxes, so later steps never treat them as evidence. A fan-in pattern brings the three raw feeds into one standardised streaming table, and every downstream step reads one clean, comparable version of every client.

2. Count who shares a value. A value used by many people cannot reliably identify one of them, so before anything is compared the pipeline counts how many distinct people use every email, phone, name-with-birthdate combination and address, once within each source and once across all sources together.

A value that is shared too widely is withdrawn as identifying evidence for that run. Both counts are materialized views over the standardised records.

3. Choose the participants. Standardising a record makes it comparable, but not every record should enter the matching process. The gate applies four conditions: an identifying attribute or a transaction, record validity, record status, and shared identity.

A record participates when:

  • it has an identifying attribute, such as a usable name, email address or phone number, or a transaction
  • it is a valid record
  • its current status permits matching
  • its identifying information does not represent a shared identity

A shared identity is identified through the counts in step two: when more than ten records share the same identifying information, that information is treated as representing a shared identity rather than an individual customer.

A transaction can qualify a record even when its identifying attributes are incomplete. Records that do not meet the gate go no further.

4. Form the pairs. Pairing every record with every other would mean trillions of comparisons. Each participating record instead carries up to four retrieval keys, used by Splink as its blocking rules: two records are paired only if they share a key.

Retrieval key Built from
Email The normalised address
Phone The number in international format
Name with date of birth Sorted name words and the birthdate together
Name with address Sorted name words and the normalised address together

A record’s keys come from the values it still has, so an email or phone number withdrawn by the counting step produces no key. A record can therefore arrive with four keys or with none.

Name and date of birth are only ever used as a pair, because a common name can reach thousands of records and a common birthdate thousands more, while the two together reach a much smaller candidate set.

5. Score each pair and record the edge. Pairing and scoring are one pass in Splink, which settles two questions at once:

  • Which records get compared? The four retrieval keys decide.
  • How strong is the evidence for the pair? The pair is scored against five comparison fields: email, phone, name, address and date of birth. A pair found through a shared email is still scored against all five.
Field Compared by Detail
Email Exact match After normalisation, shared values already withdrawn
Phone Exact match E.164 format, mailing country as a regional hint
Name String similarity (Jaro-Winkler) On romanised, alphabetically sorted words
Address String similarity (Jaro-Winkler) After abbreviation mapping, country excluded
Date of birth Exact, or day/month swapped The most common data-entry transposition still counts

No single field can decide on its own, so all five are weighed together and the pass returns one probability per pair. Exact match describes how a field is compared, never how the decision is made. Agreement can increase the probability and disagreement can decrease it, with the size of each contribution learned from the brand’s data. Two people sharing a mailbox do not therefore score as one person when their names and phone numbers disagree.

A field that is empty on either record takes no part in the comparison. Two records with no birthdate may still reach high confidence from email, phone and name together.

Each scored pair becomes one row in the edge table: the two record identifiers, the probability that they are the same person, and which of the four retrieval keys brought them together. Those rows describe a graph. Each record is a node, and each scored pair becomes an edge. Against the possible combinations, the resulting graph is highly selective.

6. Cluster the matches. The scored edges are resolved into clusters, each holding the records the evidence connects as one client. A client’s records often reach each other only through a third record: A matches B, B matches C, and A and C were never compared.

We use a custom connected-components implementation in Spark to follow every edge that clears the match threshold and collect each set of records joined by a chain into one cluster.

The clustering preserves immutable golden-record identifiers. A record that already carries one keeps it, whatever shape its new cluster takes. The resulting clusters are stored as a dataset of their own, so every subsequent step that needs them reads the same resolved result.

Three records joined by a chain: A and C share no key and are never compared, but both connect through B

7. Assign the golden record. Each cluster represents one client, but the cluster itself is not the identity that downstream systems can safely reference. The pipeline assigns a golden record identifier to the cluster and maintains that identity as the matching evolves.

The identifier is a foreign key. Everything the brand has recorded about a client points at it. That identifier is issued once and never changes, even when new evidence changes the records belonging to the cluster. Replace it and nothing errors. Every row that pointed at the old identifier still finds a client, but it finds the wrong one. That is worse than a broken link, because a broken link announces itself. Each cluster therefore resolves to one golden record: the single customer identity that downstream systems use.

When two existing golden records are subsequently found to describe the same person, one golden record becomes the parent of the other. Neither identifier is deleted or rewritten. Existing references continue to point to the identifiers they were originally given, while the parent relationship provides the path to the current identity.

The parent relationship can itself form a chain. Each golden record therefore resolves to a root: the identifier representing the current identity of the client.

The matching state is held between runs in a streaming table with one row per source record, including the golden-record relationship and the parent link the current root resolves through.

Evidence can change as well. A client changes an email address, or a value is no longer sufficiently unique to support a match, and the two records may no longer be paired. The previous edge remains as historical evidence from the run that produced it, while subsequent runs use the current evidence. An edge that is no longer supported by the current records is kept as history, but takes no part in the current clustering.

A merge already written remains part of the identity history. The parent/root structure preserves how identities were combined while downstream references continue to use the identifiers they already hold.

Two merges across three runs: each absorbed golden record points at its parent, and every identifier in the chain still resolves to the same root

8. Record the matching outcome. Every record leaves the run with a named outcome: matched to an existing golden record, assigned a new golden record, or sent to a reviewer with the evidence for and against side by side.

About 84 percent are matched without human involvement. About 14 percent stand as unique clients. Around 2 percent reach a reviewer, who confirms the match or keeps the records apart, and the next run carries that decision forward.

A record can participate in matching, have every one of its retrieval keys withdrawn by the counting step, and still leave the run with a golden record of its own. Matching decides which records belong together, not which records get an identity.

Among the clients whose records were joined at all, roughly ninety-nine in a hundred were joined across systems rather than within one. The cross-system duplication exists because the three systems do not share a common customer key. That is the gap the matching closes.

Making a name comparable

Step one standardises every field. Names are the hardest, because the same person can be written in several valid ways before the matching can compare them.

The first challenge is the writing system: Korean names romanised with a Korean romaniser, Japanese with Hepburn, Latin accents folded to plain form. The second is word order. Names reduce to alphabetically sorted words, so “Victor Soh” and “Soh Victor” arrive at the same value. The rest is noise. Honorifics and titles are stripped in each language the brand sells in, and punctuation and spacing are regularised.

Sorting the words makes differently ordered names produce the same retrieval key. In markets where the family name comes first or last depending on who typed it, that step prevents word order alone from keeping otherwise comparable records apart.

All of this runs in the first dataset of the pipeline, the standardised source, so every step downstream inherits the same representation of each source record.

The signals people share

An email address can be strong evidence of identity and still be used by more than one person. The value itself does not tell the matching how widely it is shared, yet that context changes how much weight it can carry.

The pipeline therefore measures how widely each identifying value is used. Two datasets do this at different levels: one counts the distinct people behind each email, phone, name-with-birthdate and address within each source system; the other measures the same values across all three systems. These counts become part of the matching evidence, so the weight a value carries follows how widely it is shared, with no list of addresses or numbers to maintain.

How widely a value is shared determines how much weight it can carry

A value used by very few records can provide strong individual evidence. As the number grows, its ability to distinguish one customer from another falls. The production configuration therefore sets a sharing threshold beyond which the value no longer qualifies as individual evidence. It builds no retrieval key and contributes nothing to any score, so it cannot return through a pair that met some other way.

A family mailbox is the hard case. Two or three records may legitimately share the same email, while the same pattern could also represent an assistant ordering for an executive or another shared business arrangement. The count alone cannot tell those cases apart.

The matching therefore looks for independent evidence when a shared value is involved. A phone number, for example, can support a shared email when the two records carry the same phone; a contradiction can keep them apart. Where the available evidence remains inconclusive, the pair goes to review. This is the one place where the counting reaches back into a decision already taken: a pair that has already scored high and already been matched is returned to review when its only strong signal is a shared value and nothing independent supports it.

Evidence beyond the shared email Decision
Phone numbers match Confirmed
Phone numbers contradict Stopped, two people
No phone to compare, names clearly similar Retained on combined evidence
No phone to compare, names differ Held for human review

A bare name never makes a match on its own, but name similarity can provide supporting evidence alongside an identifying contact value. Address stays out of independent confirmation, because it can be shared and because clients move.

A pair in the production data ran into exactly this. Two records share a family email address and an address, and they score 0.9998. One name is written in full, the other is initials, so the names can neither confirm nor contradict. The phone numbers disagree. The merge is declined and the pair goes to a person.

Two pairs, two decisions: a low-scoring pair joined on agreeing evidence, a near-certain pair held on contradicting evidence

The score is not wrong. It reports, faithfully, that the two records are very similar. But a score measures how similar two records are, not whether they are the same person.

One score, two thresholds

Grouping records and merging identities are different acts with different costs, so one threshold cannot carry both jobs. The implementation uses separate thresholds for clustering and golden-record merging, and both are configurable. The production values are 0.95 for clustering and 0.999 for merging, and the merge threshold is where the brand’s condition becomes a number: two records are joined only when the evidence is conclusive.

The four zones of the match score

The gap between the two is deliberate, and it is where the review queue lives. A pair that clears the clustering threshold but not the merge threshold can belong to the same cluster while both golden-record identifiers remain separate. Nothing is absorbed until the stronger threshold is reached or a person decides. The longer review queue is the accepted cost of that bar.

One agreeing field never clears the clustering threshold by itself. Email agreement on its own does not carry a match, and an explicit rule holds that line regardless of the probability assigned by the model. That rule remains in place independently of later threshold changes or model retraining.

The exception is not a second field but independent corroboration. When the same identifying value appears independently in two of the three systems, and the value is not classified as shared, that cross-system agreement provides evidence in its own right: two systems recorded the same value separately, without knowing about each other’s record. It can carry a pair that would otherwise have waited for a reviewer.

What changed for the brand

When she is next in a store, the associate sees the web order, the alteration and the six years of history, assembled from all three systems into one client. That golden record becomes part of the customer 360 view the rest of the business builds on.

The value extends beyond the consolidated identity:

  • Better data quality at source. Two quality dashboards show where identifying information is missing, inconsistent or shared across records, and where the matching produces ambiguous cases. Those findings can feed back into registration, checkout and client interactions, improving collection at source.
  • A stronger customer signal across the business. Segmentation can span the complete relationship, clienteling can see online and store behaviour together, marketing can use one consistent identity for audience selection and frequency decisions, and analytics can calculate lifetime value, retention and channel behaviour per client.
  • An explainable customer view. The matching history records the evidence behind each link, so a disputed relationship can be explained from the record. The same lineage provides the starting point for downstream processes such as a right-to-erasure workflow.

Three consequences reach the client directly. Her spend is no longer divided across identities: what she bought online, in Seoul and in Paris counts once against one client, so her tier reflects her complete relationship with the brand rather than the largest of three fragments. A consent decision can be attached to the consolidated customer identity. And she receives one message, instead of the same offer three times from three systems.

If you are building one of these

Five things transfer to any platform and any client base.

  • Standardise before you match. Put names into a common representation, phones into one format and addresses into one shape before forming pairs. A name written as “Victor Soh” and “Soh Victor” should not become a missed match simply because the words were entered in a different order.
  • Decide who takes part. Records that do not meet the matching gate should be excluded before any pair is scored, and they remain available for their original purpose and audit trail.
  • A score measures record similarity, never personal identity. A pair can score 0.9998 and still be two people, so clustering and golden-record merging should not rely on the same threshold.
  • Treat shared values as evidence with a limit. A value used by many records should not continue to carry the same evidential value as an individual identifier. Where a smaller shared group needs to be resolved, use independent evidence rather than letting the shared value confirm itself.
  • Keep the matching history. A disputed link should show the evidence that created it, and later operations such as rebuilding a cluster or supporting a right-to-erasure workflow should be able to start from that same history.

This solution runs in production on the Databricks Data Intelligence Platform, using Lakeflow Declarative Pipelines, Spark, Delta capabilities, Unity Catalog, Databricks dashboards and Splink. For more on Databricks in retail and consumer goods, see the retail industry solutions.

Distribute:

/turn your vision into reality

The best way to start a long-term collaboration is with a Pilot project. Let’s talk.