PIM and ecommerce replatform projects run late and over budget with remarkable consistency.
PIM and ecommerce replatform projects run late and over budget with remarkable consistency. PIM in particular has a terrible track record. Write-offs, overspends, implementations that quietly get shelved.
The easy explanation is the technology, or the vendor. It is almost never either.
What actually goes wrong is that the product data was not ready, and nobody checked before the contract was signed. Preparation is almost always cheaper than remediation. That holds in more or less every product data project we have worked on.
So here are the four things to do before you commit to the software. None of them need a vendor. All of them are cheaper now than they will be later.
The two versions of this failure
The first is a business putting in an ecommerce platform for the first time. Everything on that new site depends on complete product data. The filtering, the navigation, the product pages themselves. If the data is not ready, the platform goes live with product pages that have nothing on them.
The second is a business that already has an ecommerce platform and is now implementing a PIM. What usually happens is that they lift the old data out of the existing platform, which is rarely any good, push it into the PIM, and get no value from it. New system, same data, same problems.
Neither of those is a technology failure. Both are preparation failures, and preparation is the part nobody budgets for.
1. Audit where your data actually lives
You cannot start a project like this without knowing what you have and where it comes from. That sounds obvious. It gets skipped more often than any other step.
Go and find every place product data lives. The ERP. The ecommerce platform. Shared folders. The spreadsheet somebody has been maintaining for seven years. The database nobody quite knows the build of and nobody touches.
Three things tend to surprise people. How many sources there are. How much data sits in them. And how much of it contradicts itself from one source to the next.
For each source you want four facts: what data it holds, where it lives, who owns it, and how often it changes. If you cannot answer those, stop. Do not go further into the project until you can.
The alternative is well documented. You commit to the board that the new platform goes live in six or ten weeks, then discover you need six months of data migration readiness work first. You pay for it either way. The only question is whether you pay before the deadline is set or after.
2. Map the channels before you build the data model
Step one was about where data comes from. This one is about where it goes.
You are buying a PIM or a platform to provision data to channels, whether that is B2C, B2B or both. So the data model has to be guided by what those channels need. Not by what your sources happen to contain.
Most projects get this the wrong way round. Teams take the sources they have, build a model that fits them, and end up with something that cannot support filtering, search or SEO, because the sources were never complete enough for any of it.
Build the list from the other end. Every destination the data goes to, every attribute that destination requires, the format it needs, whether values need localising, whether they need contextualising. For a replatform that means everything driving the site: product pages, filtering, navigation, search indexing, SEO, plus feed management for Google Shopping, Amazon and eBay.
The commonly forgotten ones are compliance and market specific requirements, localisation, and back office consumers. A PIM often ends up serving the logistics team too, and nobody asked logistics what they needed.
3. Define what good data looks like
This one sounds too obvious to write down. It gets skipped constantly.
Good is not a single standard. A B2B portal and a consumer site need the data to do different jobs, so define it per channel and per job.
Three decisions make up the definition. Which attributes are mandatory. What format each one takes, down to how dimensions are presented and what a long description actually looks like. And a publish ready sample you can hold up and say: this is what good data means here.
The document this lives in is a data dictionary. One large spreadsheet holding every channel, every attribute ID and name, the format, the type and the rules, filterable so you can check you have everything for every channel. People are put off by the term, because it sounds like something a data governance function produces in a large enterprise. It is not. We have built these with businesses running catalogues of a thousand products and with some of the largest catalogues you will find. The principle does not change.
What does not work is having it written for you. The business has to own the definition, with input from every team that touches product data and one single accountable owner. We have started calling that role the product data owner rather than the product owner, because they are not the same job.
4. Fill the gaps before you migrate
The most obvious step, and the one most likely to be deferred. By this point you have your sources, your target and your definition of good. What is left is the difference between them.
Most businesses plan to fix that difference in the new system. It does not work like that.
A PIM does not fix your data. It gives you a new shell. It will not fill your gaps, and it is frequently sold as though it will, which is a large part of why the disappointment is so consistent.
So run enrichment and migration as a project alongside the implementation, not after it. It is genuinely big work. Enriching or migrating 10,000 SKUs is a serious undertaking, and pretending otherwise is how timelines slip.
Migrate incomplete data and the gaps do not stay hidden. They surface across every downstream channel at once, and you rectify them later, under more pressure and in public.
Whenever a new ecommerce platform launch gets announced on LinkedIn, the first thing we do is go and look at the product pages. You can guess how that usually goes.
The steps only work in order
These four are a sequence, not a menu.
You cannot fill gaps until you can say what a gap is, which means you need the definition of good. You cannot write that definition sensibly until you know what the channels require. And you cannot plan any of it until you know what data you hold and where it sits.
Run them out of order and you get the familiar version of this project. A model built from whatever the sources happened to contain. A launch date agreed before anybody counted the work. And a remediation phase that nobody was willing to call a remediation phase.
All four steps are preparation, and preparation is the cheap part. It stays cheap right up until the contract is signed.
The one hour version
If four steps sounds like a project in itself, start with the hour that matters most.
Open a blank document. List every place your product data lives today. Write down who owns each one. That is roughly an hour of work, and almost everybody who does it finds considerably more sources than the official systems list contains.
That single page gives you the footprint: what data exists, where it sits, and in whose hands. It is also the point at which a vendor conversation becomes useful rather than theoretical, because the questions they will ask are these questions. Where does your data come from. Where does it need to go. What does good look like.
Arrive with the answers and you are negotiating. Arrive without them and you are buying a shell.
Listen to the full episode
Episode 9 of Product Data Weekly is available now. For more episodes and the weekly newsletter on operational issues inside product data and ecommerce teams, visit productdataweekly.com.
See SKULaunch in action
Watch how we handle AI enrichment, supplier onboarding, and catalogue scale in a live 30-minute demo.
.avif)