A shoe retailer came to us with 11,000 products and a Magento to Shopify migration in front of them.
A shoe retailer came to us with 11,000 products and a Magento to Shopify migration in front of them.
The instinct in that situation is to treat it as a moving job. Export everything from the old system, wrestle it into shape, import it into the new one, and count a successful launch as everything arriving intact.
That instinct is what produces a new website with the old website's problems, rendered in a nicer theme. A replatform is the one moment you get to change the model underneath, and almost nobody uses it.
Nobody asks whether the new site will actually be better
When a business relaunches on a new platform, the attention goes to features. The layouts, the checkout, the colour of the buttons.
The question that rarely gets asked is whether the new site will be better than the old one in the way that matters commercially: whether customers can find things, filter them, and get enough information to buy.
It sounds like a silly question. It almost never gets answered.
There is also a technical reason a straight copy was never on the table. Shopify and Magento have fundamentally different product data models. They structure things differently, they manage them differently, and even the terminology means different things. Something has to be redesigned in the move, so the only real decision is whether you redesign deliberately or by accident.
The standard approach carries the problem across
Here is what usually happens instead.
Somebody runs an export out of the legacy system, whether that is Magento or the ERP behind it. Then they spend roughly a month inside a spreadsheet, running macros, trying to force the output into a shape the new platform will accept. Sometimes there is tooling involved. Often there is not.
In parallel, the team recreates the existing structure in the new platform, because that is the version everybody recognises and it feels like the safe option.
Two things go wrong with that. The new platform has different limitations, so the old structure does not fit cleanly. And more often than anybody wants to admit, the old structure was wrong in the first place. Lifting and shifting reproduces it faithfully, at cost, on a brand new platform.
You do not have 11,000 products
One of the first questions we ask is how many products a business has. The answer is almost always a version of I think it is about this many.
The follow-up is the useful part. Are those products, or variants? Sellable SKUs, or parent records? Because the answer changes the size of the job by an order of magnitude.
Shoes make the point better than any other category. A single style comes in a run of sizes, several colours, sometimes multiple widths for running styles. One shoe can appear as 40 or 50 rows of data.
So 11,000 rows is not 11,000 products. It is roughly 1,500 shoes, with the variation sitting underneath them. The moment you see it that way, cleaning, migrating and enriching the catalogue stops looking like a year of work.
Getting the terminology straight is part of this and it trips up nearly everybody. An item or product in an ERP is almost always the sellable thing, the variant. A product in a PIM or an ecommerce platform is usually the parent. Same word, opposite end of the hierarchy. Settle that before anybody starts counting.
Profile what you have, then design what you want
Two things happen next, and they run in parallel rather than in sequence.
The first is profiling, which is a grand word for looking properly at what came out of the old system. You are hunting for missing values, duplicates, and the same product listed twice under different records. You will also find things that are not products at all. Packaging is a common one. Services are another. Businesses store all sorts in the back end of an ERP, and none of it belongs on the new site.
The output is a clear view of what is valid, what is junk, and what is absent.
The second is designing the target model, meaning both the categorisation structure and the way products are modelled for variation. Your existing data is an input to that design, not a template for it.
In practice that looks like this. Today the hierarchy might be two levels deep, shoes and then men's shoes. Tomorrow there is real granularity underneath: running, casual, and so on. Customers get a faster route to what they want, and search engines get something they can categorise.
The same applies to attributes. If men's shoes currently carries three filters, ask what the category actually warrants. Material. Sole type. Heel height, depending on the shoe. The result is a more granular, more customer centric model per category, and yes, you will not have data for all of it yet. That is the point of knowing.
Moving from source to target
Now you have profiled source data on one side and a target model on the other, and two questions left. How does the valid old data move into the new structure, and how do you fill what is missing?
Mapping used to be the painful part. Sitting with two schemas, matching one attribute to another, one row at a time. AI models have genuinely changed that, and most of the mapping work no longer needs to be done by hand.
Filling gaps has more than one route. Often the old platform held good descriptions but never held granular attributes, in which case the existing content is the source: sole type and material can be extracted from the copy you already own. Where that is not enough, you go back to suppliers, which is its own discipline and its own episode.
The discipline that saves the most work is deciding what level each attribute belongs at. Descriptions for a shoe live at the colour level or above, not at the size level, because nobody needs bespoke copy explaining that this one is a size 42. Size and colour specific values sit on the variant.
Which is why the enrichment brief is not 11,000 descriptions. It is closer to 1,500. Roughly a tenth of the work, for the same result.
The takeaway
None of this makes the import itself easier. Loading Shopify still means running imports and mapping attributes in, and that job is the same size it always was.
What changes is what lands on the other side. A clean, well structured data set drives better search, better filtering, better SEO and better visibility in generative engines, from day one rather than eighteen months into a remediation project nobody budgeted for.
So if a replatform is on your roadmap, treat the product data as the project rather than the payload. Count what you actually have before you count rows. Profile the old data honestly. Design the model you want rather than the one you inherited.
You get one clean opportunity to change the structure. It arrives on the day you change platform, and it does not come round again for years.
Listen to the full episode
Episode 18 of Product Data Weekly is available now. For more episodes and the weekly newsletter on operational issues inside product data and ecommerce teams, visit productdataweekly.com.
See SKULaunch in action
Watch how we handle AI enrichment, supplier onboarding, and catalogue scale in a live 30-minute demo.
.avif)