Most data platform projects are scoped as 'migrate to X', and the vendor comparisons that follow are about X. The expensive mistakes happen earlier: in what gets migrated, who owns it afterwards and what the first useful output is. Settle those and the platform choice gets easier, and so does the vendor choice.
Decide what you're not migrating
Every warehouse has tables nobody has queried in a year and pipelines nobody can explain. Migrating them costs the same as migrating the ones that matter. Start with usage logs: the tables, reports and models queried in the last ninety days are the scope; the rest is archived, not migrated. This single step usually halves the estimate.
Pick the first consumer
A platform with no consumer is a cost centre. Choose one business team whose reports or models will move first, with a named owner who'll say when they're satisfied. That gives the project a definition of done, gives the vendor something to show in eight weeks, and tells you whether the platform works for your people before you've committed everything to it.
Settle ownership
Who owns the pipelines after go-live: your team, the vendor on a managed service, or a mix? The answer changes the build. A platform your team will run should be built in the tools your team knows, with the vendor training them as it goes. A platform the vendor will run can use whatever the vendor is fastest in. Decide this before the RFP, because vendors will assume whichever answer suits them.
Then compare platforms on your workload
The three platforms are converging, and the honest differences are about your estate, not their feature lists.
| Platform | Strongest when | Watch the |
|---|---|---|
| Snowflake | SQL-first analytics, data sharing, many business users | Consumption pricing without governance; compute runs when nobody's looking |
| Databricks | Engineering-heavy pipelines, ML and AI, open table formats | Skills: you need engineers, not just analysts |
| Microsoft Fabric | Microsoft estates, Power BI users, capacity-based pricing | Maturity of the newer workloads; capacity sizing |
Questions for the partner
- How many migrations from our source platform have you done, and can we talk to two of those clients?
- What's your approach to the tables we don't migrate?
- Who on your team will train ours, and how many hours are in the plan for it?
- What does the platform cost to run in month twelve, on our volumes, and how do you know?
- What would you build first, and what would we see at week eight?
- data platform
- Snowflake
- Databricks
- Microsoft Fabric
- data engineering
- migration
The Scopelyst team
Written by the people who run the marketplace, from what buyers and vendors do on it. No guest posts, no sponsored articles.
