Most CRM projects get sold on the interface, the automations, the dashboards. Most of the actual project time goes somewhere less visible: getting the data that already exists, in whatever state it is currently in, into a shape the new system can use without lying to whoever looks at it next.
This is the mechanical version of that process, drawn from HubSpot and Zoho migrations we have run, not the argument for why cleanup matters (that case is made elsewhere) but the actual sequence: what gets checked, in what order, and what breaks when a step gets skipped.
Step 1: audit before anything moves
Before a single record is touched, the source data gets pulled and profiled. That means counting records per object, checking which fields are actually populated versus technically present, and finding out how many records are duplicates before dedup logic gets designed rather than after.
| Check | What it catches |
|---|---|
| Field fill rate | Fields that look standard but are empty on 80% of records, not worth mapping |
| Duplicate rate by matching key | How much of the contact list is the same person saved two or three times |
| Orphan records | Contacts with no associated company, deals with no associated contact |
| Date sanity | Created dates in the future, close dates decades in the past, common in manually entered spreadsheets |
A common source is an Office 365 or Outlook contact export headed into HubSpot. Outlook contacts are built for one person's address book, not a shared CRM: no company-contact relationship, no lifecycle stage, no deal history, and years of the same person saved multiple times under slightly different names or emails. That audit step is where you find out you are not migrating 4,000 contacts, you are migrating maybe 2,400 unique people, and that number changes the whole plan.
Step 2: dedupe against a matching key, not a gut check
Deduplication needs a defined matching key before it needs a tool. Email is the obvious first key, but it is not sufficient on its own: the same person can have a work email and a personal one saved as two contacts, and two different people can share a generic info@ address that should never be treated as a duplicate.
- Primary key: exact email match, case-insensitive, trimmed of whitespace.
- Secondary pass: name plus phone number match, for records with no email or a placeholder one.
- Company dedup separately from contact dedup: "Acme LLC," "Acme," and "ACME L.L.C." are the same company under three spellings, and merging them wrong creates a single company record with the wrong contacts attached.
- Manual review queue for anything the matching rules flag as ambiguous rather than confident, instead of auto-merging on a low-confidence match.
Which record survives a merge matters as much as catching the duplicate. The rule we use by default: most recently updated record wins on conflicting fields, but any field the losing record has populated and the winning record has empty gets kept rather than dropped.
Step 3: map fields to where they actually belong
Source fields rarely map one-to-one to the destination CRM's schema, especially when the destination has custom objects the source system never had. This is where a Zoho CRM setup for a specific business process, a business brokerage listing pipeline is a real example, diverges from a standard migration: the standard Contacts and Deals modules do not represent "listing," "buyer," "seller" and "NDA status" as distinct entities, so custom modules get built before the data has anywhere correct to land.
| Source | Destination | Transformation needed |
|---|---|---|
| Free-text "Status" column | Standardised deal stage picklist | Every unique text value gets mapped to one of the new fixed stages, including catching typos and near-duplicates |
| Single "Notes" field with years of entries | Timeline of dated activity records | Parsed and split by date pattern where one exists, otherwise attached as a single historical note rather than lost |
| Separate spreadsheet of company data | Linked company object | Matched to contacts by domain or company name before import, so contacts don't land as orphans |
Custom fields and modules get built and tested with sample records first. Importing 6,000 records against a field structure that turns out to be wrong means redoing the mapping and reimporting, which costs more time than getting the structure right before the bulk import runs.
Step 4: import in batches, not all at once
A full import in one pass makes errors expensive to find, because they are buried in thousands of rows. Batching by object type and by a few hundred records at a time means a mapping error shows up in batch two instead of after everything is already in the system.
- 1Companies import first, since contacts and deals will need to associate against them.
- 2Contacts import second, matched to companies by domain or explicit ID.
- 3Deals and activity history import last, once the contacts they reference already exist.
- 4Each batch gets spot-checked against the source before the next batch runs.
Step 5: validate twice
Validation happens before the migration, on the source, to catch what needs fixing ahead of time. It happens again after the migration, on the destination, to confirm nothing broke in transit. Skipping the second pass is how a mapping error someone approved in testing quietly makes it into production and stays wrong until a salesperson notices six weeks later.
- Record counts by object match the audited source count, accounting for merged duplicates.
- A sample of records, not just the first ten, gets checked field by field against the source.
- Required fields for automations and workflows (the ones a workflow's trigger depends on) are confirmed populated, not just present.
- Associations, contact to company, deal to contact, resolve correctly rather than pointing at the wrong record after a merge.
The pattern across every migration we run is the same regardless of source or destination platform: the failures that matter are never in the import itself. They are decisions made, or skipped, in the audit and mapping steps before any record moves.
Common questions
It depends far more on data quality than on record count. A clean 5,000-contact export can move faster than a messy 1,000-contact one. The audit step is what reveals which situation you are in before a timeline gets promised.
Technically yes, HubSpot accepts a CSV export directly, but a direct import carries over every duplicate, orphaned entry and inconsistent field HubSpot's own dedup tools will only partially catch. A dedup and mapping pass first produces a cleaner result than importing and cleaning up afterward.
No. Standard businesses fit the default Contacts, Accounts and Deals modules well. Niche processes, brokerage listings, membership renewals, multi-step approvals, usually need at least one custom module or layout to represent the workflow accurately.
They migrate as duplicates. Reports double-count, automations fire twice for the same person, and a sales team starts double-checking the CRM against their own notes, which is the point most teams quietly stop trusting the system.