DEV Community

Cover image for I Built a Shopify HubSpot Integration for a Client With 200K Contacts
Elsie Rainee for WPWeb Infotech

Posted on

I Built a Shopify HubSpot Integration for a Client With 200K Contacts

If you have a Shopify store and a large HubSpot database, the real problem is rarely connecting the two platforms. The difficult part is making sure customer, order, marketing, and lifecycle data move between them without creating duplicate contacts, overwriting useful information, or turning your CRM into a mess.

I ran into exactly that challenge while building a Shopify HubSpot integration for a client with 200K contacts. The experience taught me that a successful integration is less about connecting APIs and more about deciding what data should move, when it should move, and which system should own each field.

Why a 200K-Contact Integration Is Different

A small Shopify store can often get away with a basic connector. A large database is another story, and it's where a Shopify HubSpot integration service has to do more than switch on a sync.

With 200,000 HubSpot contacts, even a small synchronization mistake can have a significant impact. A workflow that accidentally creates duplicate contacts, repeatedly updates the same records, or sends unnecessary events can quickly become difficult to clean up.

Before writing any integration code, I mapped the client's data flow. The main systems were:

  • Shopify: customers, orders, products, discounts, and purchase activity
  • HubSpot: contacts, lifecycle stages, marketing data, segmentation, and workflows
  • Integration layer: transfers and transforms data between the platforms

The first question was not, "How do we connect Shopify to HubSpot?" It was, "What information does HubSpot actually need from Shopify to support the client's workflows?"

That distinction saved a lot of unnecessary complexity.

The Data Mapping Came First

One of the biggest mistakes in ecommerce integrations is starting with API calls before defining the data model. I created a field-by-field mapping before building any synchronization logic:

Shopify Field HubSpot Destination Action
Customer email Contact email Identify contact
First name First name Update
Last name Last name Update
Phone Phone Update when available
Order count Custom property Recalculate/update
Total spend Custom property Update
Last order date Custom property Update
Customer tags Contact property Transform before sync
Order details Ecommerce/order data Sync based on requirements

This exercise also exposed an important issue: not every Shopify field belongs in HubSpot. Sending everything sounds convenient, but it creates unnecessary data, more synchronization events, and a harder-to-maintain CRM.

The goal was to transfer useful business information, not to reproduce Shopify inside HubSpot.

Contact Matching Was the Critical Part

With 200K contacts, duplicate prevention became one of the most important technical requirements. The integration needed a reliable way to answer one question: "Does this Shopify customer already exist in HubSpot?"

Email was the primary matching value for the client's setup, but the integration still had to account for incomplete or changing customer information. A simplified synchronization process looked like this:

  1. Receive or retrieve Shopify customer data.
  2. Normalize the relevant fields.
  3. Check whether the contact already exists in HubSpot.
  4. Update the existing record when a match is found.
  5. Create a new contact only when no valid match exists.
  6. Record the synchronization result.

Handle failures separately instead of silently ignoring them.
That last step matters more than it sounds. An integration should not simply "try again" indefinitely when something fails. You need to know what failed, why it failed, and whether retrying is safe.

Handling 200K Contacts Required More Than a One-Time Import

The initial migration was only one part of the project, and a bulk migration of 200,000 contacts needs to be treated differently from everyday synchronization.

Instead of processing the entire database as one operation, I approached the migration in smaller batches. The basic pattern was:

Extract → Transform → Validate → Sync → Log → Verify

Batch processing made it easier to monitor progress and identify problems without restarting the entire migration. It also helped with API limits. Both Shopify and HubSpot have rules around API usage, so an integration handling large volumes needs controlled request rates, pagination, retries, and appropriate error handling.

This is where a technically simple integration can become operationally complex.

Shopify and HubSpot Should Not Automatically "Own" Everything

Another important decision was determining which platform should be the source of truth for specific information. Shopify may be authoritative for ecommerce activity, while HubSpot may control marketing-related properties or lifecycle information.

Without clear ownership, synchronization can become a loop: Shopify updates HubSpot, HubSpot updates Shopify, and Shopify updates HubSpot again. That causes unnecessary API requests and unexpected data changes.

So I defined ownership rules for key fields before enabling ongoing synchronization. For each field, we needed to know:

  • Where does the value originate?
  • Which platform can modify it?
  • When should it synchronize?
  • What happens if both systems contain different values?
  • Should an empty value overwrite an existing value?

These questions are easy to overlook when you're first planning the integration. They become painful once real customer data is involved.

Error Handling Was Part of the Integration, Not an Extra Feature

One lesson from this project was to design error handling from the beginning. The failures we planned for included:

  • API rate limits
  • Temporary network failures
  • Invalid customer information
  • Missing required fields
  • Duplicate records
  • Unexpected API responses
  • Authentication problems
  • Partial synchronization

For recoverable failures, retry logic can help, but retries need limits and appropriate delays. For permanent failures, the system should log enough information to investigate the affected record.

I also separated successful synchronization from failed synchronization, rather than treating the entire batch as successful because most records worked. That made verification considerably easier.

Testing With Realistic Data Matters

Testing with 10 sample contacts can make an integration look perfect. Testing with data that resembles a 200K-contact database is much more revealing. I focused on scenarios such as:

  • Existing Shopify customer already in HubSpot
  • New Shopify customer
  • Customer with incomplete information
  • Customer placing multiple orders
  • Updated customer information
  • Duplicate email situations
  • API rate-limit responses
  • Failed synchronization followed by retry
  • Large batch processing
  • Re-running previously processed records

One particularly important test was idempotency. If the same event or customer record is processed twice, the integration should not create another contact or corrupt existing data. That principle becomes extremely important when working with webhooks, retries, queues, and large migrations.

Testing made one thing clear: most of the problems weren't in the code. They came from decisions made before we wrote the code. That's why the lessons below matter more than any specific API call.

Shopify Tips for a Cleaner HubSpot Integration

If I were starting this project again, these are the Shopify tips I'd follow from day one, whether the database has a few thousand contacts or a few hundred thousand.

  • Start with the data model, not the API documentation: Work out which Shopify data your HubSpot workflows actually need before writing any code. Once the business workflow is clear, the technical decisions get much easier.
  • Define field ownership early: Every important field needs a clear source of truth. Shopify usually owns ecommerce activity, while HubSpot often owns marketing and lifecycle data. Without that line, you risk sync loops and unexpected overwrites.
  • Design for failure from the start: APIs fail, requests time out, and data shows up in unexpected formats. Retries with sensible limits and clear logs for permanent failures keep one bad record from becoming a bigger problem.
  • Keep the migration separate from ongoing sync: Moving 200,000 contacts in batches is a different job from handling one new customer at checkout. Treating them separately makes both easier to monitor and fix.
  • Build in observability: Sync status, error records, and basic reporting help you see what went wrong when something breaks.

Above all, build the integration around your client's real workflow, not around whatever the APIs happen to make easy.

Conclusion

Building a Shopify HubSpot integration for a client with 200K contacts showed me the hardest part isn't connecting two platforms. The real work is creating a dependable data flow between them. Contact matching, field ownership, batch processing, API limits, error handling, idempotency, and verification all matter when the database is large.

A good integration should make customer data more useful without creating another operational problem for the team managing it.

Frequently Asked Questions

1. What is a Shopify HubSpot integration?

A Shopify HubSpot integration connects ecommerce data from Shopify with HubSpot so you can use customer, order, and related information in CRM and marketing workflows. The exact data synchronized depends on the business requirements and integration setup.

2. Can Shopify integrate with HubSpot with 200K contacts?

Yes. A Shopify HubSpot integration can handle a large contact database, but 200K contacts require careful batch processing, contact matching, API-limit management, error handling, and migration planning rather than treating the database as a small import.

3. How do you prevent duplicate contacts between Shopify and HubSpot?

Duplicate prevention starts with a reliable matching strategy, commonly using a unique customer identifier such as email where appropriate. The integration should check for an existing HubSpot record before creating a new contact and should also handle normalization, missing data, and repeat events.

4. What data should Shopify send to HubSpot?

Common data includes customer identity information, purchase history, order activity, total spend, order frequency, and other fields required for CRM segmentation or automation. Not every Shopify field needs to sync, so map the data to your business requirements.

5. What is the biggest challenge when integrating Shopify with HubSpot?

For a large database, the biggest challenge is maintaining reliable synchronization, not just establishing the connection. Data ownership, duplicate prevention, API limits, failed requests, retries, and ongoing monitoring all need to be addressed to keep the integration dependable.

Top comments (4)

Collapse
 
mayur-upadhyay profile image
Mayur Upadhyay •

This is one of the more practical writeups on Shopify-HubSpot sync I've read. Most tutorials jump straight to API calls, but you're right that the field ownership question matters more, deciding whether Shopify or HubSpot controls a value is what actually prevents those sync loops later.

The idempotency testing point is underrated too. A lot of integrations look fine in a demo with 10 records and then fall apart the first time a webhook fires twice or a batch job gets re-run after a partial failure.

Curious about one thing: for fields where both systems could technically modify the same value (like customer tags), did you go with strict last-write-wins, or did you build in some kind of conflict flagging for manual review? That seems like the hardest edge case to get right at 200K scale.

Collapse
 
elsie-rainee profile image
Elsie Rainee WPWeb Infotech •

Great question. We ended up going with field level ownership rather than a blanket last-write-wins rule. For something like tags, Shopify was the source of truth since that's where the behavior actually originates, so a Shopify update would win by default. But for properties that HubSpot marketing owned, like lifecycle stage, the integration would skip the Shopify value entirely rather than let it overwrite. Pure last-write-wins sounds simpler but it gets messy fast once two systems can both legitimately claim a field. Conflict flagging for manual review was something we considered but decided against for this client, mainly because the volume would have made manual review impractical.

Collapse
 
rafidbottler profile image
Rafid Bottler •

I've seen a few integration writeups that gloss over the migration vs ongoing sync distinction, but you called it out clearly. Treating a one time 200K record migration the same way as day to day webhook syncing is where a lot of teams get into trouble with rate limits and partial failures.

The point about not letting empty values overwrite existing data is something I wish more people thought about before go live. Did you end up building a dead letter queue type setup for the permanent failures, or just structured logs that someone reviews manually?

Collapse
 
elsie-rainee profile image
Elsie Rainee WPWeb Infotech •

Yes, this ended up being one of the more important lessons. We built structured logs for every failed record rather than a full dead letter queue setup, partly to keep the infrastructure simpler for the client's team to maintain long term. Each failure got logged with the reason, the payload, and whether it was retried, so someone could query and reprocess specific records without touching the whole batch. A proper dead letter queue would have been nice at even larger scale, but for 200K it felt like more infrastructure than the problem needed. The main thing was just making sure failures were visible instead of silently swallowed.