Konabayev
DataCRMData QualityMarketing Operations

CRM Deduplication Test Dataset: 24 Fictional Record Pairs

·7 min read
Last updated on
Fictional CRM record pairs routed to candidate, review, or separate outcomes

Direct Answer

This CRM deduplication test dataset provides 24 fictional record pairs with authored decisions: link_candidate, review, or keep_separate. It is designed to reveal unsafe matching shortcuts before you apply a deduplication workflow to real customer records.

There are no real contacts in the dataset. Email addresses use test domains, names are fictional labels, and phone-shaped values are deliberately invalid test tokens. No CRM product has been measured against the cases. The labels describe our conservative acceptance policy, not verified real-world identity.

Download: JSON with rules and dictionary, CSV, JSONL, and canonical generator.

Release fieldScope
Version1.0.0; downloadable source released alongside the data
Inputs24 authored synthetic cases; no representative sample
ResultsExpected labels only; no vendor performance measured
ComparatorExact pass, mismatch or missing; diagnosis is manual

Download and use these files locally for your own acceptance tests, and cite the dataset and version. No open redistribution license is asserted for this fixture release. Keep the versioned source download with the generated files when reproducing this release. Before running it, compare the downloaded file’s SHA-256 with reproduction.source_sha256 in the version 1.0.0 JSON dataset. For example, shasum -a 256 generator.mjs prints the local file hash on macOS. Stop if the hashes differ, and retain the matching source and dataset together.

What the Three Labels Mean

A matching signal can justify review without justifying a destructive merge. Every case sets automatic_merge_allowed to false.

LabelMeaning within this datasetWhat it does not establish
link_candidateA strong identifier agrees and the fixture has no conflicting populated identifierPermission to delete a record or certainty that two records represent one person
reviewA conflict, ambiguity or weak matching signal needs examinationThat the records are definitely duplicates
keep_separateThis contract supplies no adequate linking signalProof that the records belong to different people

The distinction matters because duplicate detection, identity resolution and field survivorship are different operations. A matching engine may nominate a pair correctly while a merge rule still overwrites the wrong source, loses a suppression flag or discards necessary provenance.

Treat these files as a review exercise for your CRM implementation, not a replacement for source-system documentation or an approved merge policy.

The Matching Contract

Each pair is an independent test within a fictional source namespace. Do not concatenate the pairs into a customer table: reused tokens such as F001 have meaning only inside their own fixture. A nonempty shared source ID can nominate a link. An exact, syntactically valid, verified non-role email on both records can also nominate a link, provided another populated identifier does not conflict.

Email normalization trims surrounding whitespace and lowercases the domain. It does not lowercase the local part, remove dots or strip a plus suffix. Those transformations can be appropriate under a separately verified mailbox policy, but this dataset does not assume that every domain implements them identically.

Different populated source IDs, emails or phone identifiers require review when another field proposes a match. Blank values never count as shared identifiers. A matching name and company are review signals only. A shared role mailbox is not treated as a personal identifier.

The explicit weak-review triggers are: an identical email when either verification flag is false; local-part case, plus-suffix or dot-placement differences; an identical role mailbox; an identical phone token alone; an identical name even when the company differs; NFC/NFD-equivalent names; accent-only name differences; and an identical malformed email despite verification flags. Different email domains or different names alone, different IDs alone, and empty pairs remain keep_separate under this contract.

These are author-selected rules, not a claim about default behavior in HubSpot, Salesforce or another product. Before adapting the exercise, write down the actual source namespace, identifier authority, verification process and conflict precedence. A copied ID from a different CRM is not automatically part of the same identity namespace.

HubSpot’s import documentation, read on October 4, 2026, says that a mapped HubSpot Record ID supersedes other mapped unique identifiers. Our conflict-to-review rule is deliberately a separate policy; it does not predict that importer. The Arcs & Curves migration checklist covers field mapping, relationships and acceptance decisions. This download supplies a narrower executable pair-level exercise; it cannot replace that wider migration review.

Cases That Catch Dangerous Shortcuts

The 24 pairs vary the decision, not merely the spelling of a contact. They include both positive candidates and explicit counterexamples.

FixtureInput issueExpected decision
CRM-001Same populated source ID, no identifier conflictlink_candidate
CRM-002Same ID but two different populated email addressesreview
CRM-006Verified address with a domain-case differencelink_candidate
CRM-008Plus-suffixed address versus unsuffixed addressreview
CRM-012Two verified copies of a role mailboxreview
CRM-016Two entirely empty recordskeep_separate
CRM-018Matching verified email but conflicting phone tokensreview
CRM-024Identical malformed email with verification flags setreview

CRM-016 is a useful control against a naive equality implementation. Two records can have identical empty fields without carrying any identity evidence. CRM-024 checks a different failure: trusting a verification flag while ignoring a malformed value.

The Unicode name cases preserve the distinction between equivalent character encoding and identity. CRM-020 uses different NFC/NFD code units that become equal after NFC normalization. CRM-021 compares an accented name with an unaccented name; those remain different after NFC normalization. Both are review, not identity proof. Likewise, a changed name attached to the same nonconflicting source ID can remain a linking candidate without prescribing which name should survive.

Reproduce and Inspect the Exports

The generator reproduces the authored corpus; it does not deduplicate a database. Download the version 1.0.0 source linked above and save it as generator.mjs. With Node.js 22 or later, run the commands below; no repository access or package installation is required:

node ./generator.mjs --out ./fixture-data
node ./generator.mjs --out ./fixture-data --check

That check compares local files with this source contract. To verify the published downloads instead, put all nine dataset files in a separate ./downloaded folder and run node ./generator.mjs --out ./downloaded --check. Expect “Verified 9 files”; this checks published byte parity, not a CRM product’s identity decisions.

The JSON contains the complete rule definitions, data dictionary, provenance and all cases. JSONL has one case per line. CSV has the same records, with input, expected and rule_ids serialized as JSON in quoted cells. Parse CSV properly before decoding those fields.

The field named phone_test_token holds only intentionally invalid +999 test tokens in this release. They test equality and conflict handling, not telephone validity or reachability. Do not load them into a dialer. The email_verified flag is also an authored input, not evidence that a verification service contacted a mailbox.

Retain the downloaded version 1.0.0 source with any result; do not substitute a later release when reproducing the baseline. Keep test data separate from production sync jobs, campaigns and enrichment tools. Nothing in reproducing these files requires uploading private CRM exports to an external reviewer.

Run a Matcher and Preserve the Review Trail

Test a read-only adapter before permitting writes. Feed each pair to the matcher or rule evaluator, then map its result to the three labels and the explicit no-automatic-merge flag.

Submit an array of { fixture_id, actual } records to the generator module’s exported compareResults function. Missing fixture results remain visible as missing; duplicate or unknown IDs reject the run. A candidate labelled review is a mismatch with this exact contract, although it may be an acceptable conservative choice for your separately documented process.

Retain raw output, normalized output, matcher settings, version and the reason for each disagreement. Do not silently change a normalization rule halfway through a run. If the vendor cannot expose a rule explanation, record that limitation rather than inventing an explanation from its label.

For an operational rollout, preserve an export, stable IDs and an audit trail. Define field ownership, source timestamps, association handling and reversal before any approved merge. The HubSpot and Salesforce integration guide provides related questions for cross-system workflows; this fixture contract intentionally covers only one namespace.

What a Passing Report Can Claim

Passing means agreement with 24 authored expected decisions under one versioned policy. It does not estimate a production false-merge rate, duplicate prevalence or time saved.

This dataset has no representative sample of your CRM, no hidden evaluation split and no ground-truth biographies. Tuning directly against all its expected outputs can establish regression consistency but cannot demonstrate generalization. Add independently reviewed, permitted cases from your workload before drawing an operational conclusion.

Keep disagreements grouped by risk. Accidentally proposing a merge for two empty records is different from sending a plausible candidate to manual review. Report both instead of hiding them behind a single pass percentage. The marketing analytics framework explains the broader need to retain definitions and uncertainty when reporting business outcomes.

FAQ

Are these records anonymized customer data?

No. They were authored as fictional test inputs. There is no underlying private customer export or real identity record to recover.

Why does the dataset never authorize automatic merging?

Pair matching does not determine field survivorship, permissions, suppression or linked-object handling. The dataset deliberately tests candidate decisions while leaving destructive operations outside scope. An operational merge requires a separate reviewed contract.

Can I compare two CRM products with these cases?

You can report a carefully documented test of these cases if both adapters implement the same policy and disclose unsupported inputs. Do not call the result a general accuracy benchmark or imply that untested products fail. This release itself contains no vendor measurements.

Last checked: October 4, 2026. Authored fixture labels, export parity and comparison safeguards reviewed. Current third-party matching behavior has not been verified by this dataset.

AI Marketing Automation

Custom AI workflows, dashboards, and internal marketing tools.

View service details

Google Preferred Sources

See more of my research in Google

Add Konabayev.com as a preferred source to find more fresh marketing and AI research in Google Search.

Have a relevant product or documented use case? View sponsorship formats and editorial requirements.