Why a PE-backed platform needs a data lake to make the numbers tie out
A platform acquiring companies on different systems can't get consolidated numbers it trusts until every entity's data lands in one place. That place is a data lake, and it comes before the dashboards and the AI.
Picture a platform that’s bought five companies in two years. One runs on a cloud field-service tool, two on different versions of an on-prem ERP, one on QuickBooks, and the newest still lives mostly in spreadsheets. The sponsor wants consolidated KPIs two weeks after the latest close. Someone spends a week exporting reports and pasting them into a master workbook, and the revenue figure still doesn’t match what two of the operating companies think it is.
That isn’t really a reporting problem. It’s a foundation problem, and it’s the clearest sign a business needs a data lake.
What a data lake actually is
Strip away the vendor language and the idea is simple. A data lake is one place where data from every system you run lands in its raw form, before anyone decides what to do with it. Financials, sales records, payroll, work orders, spreadsheets, all of it flows into a single store and stays as-is, with structure applied later when you turn it into a report or a metric.
That last part is what people miss. A data lake doesn’t ask you to clean and model everything up front. It captures first and shapes later, which is exactly what you want when new sources keep arriving and you can’t predict every question you’ll eventually need to ask of the data.
Data lake, data warehouse, lakehouse: what an operator needs to know
You’ll hear three terms, and the distinction matters less than the vendors selling them suggest. A data lake holds everything raw. A data warehouse holds structured, modeled data that’s ready for reporting. A lakehouse is the newer pattern that folds both into one tier.
For an operator, the practical version is this: the lake is the landing zone every entity flows into, and the warehouse layer is where that raw data becomes a number the board can trust. You’ll end up using both. The lake just has to come first, because you can’t model what you haven’t captured.
| Data lake | Data warehouse | Lakehouse | |
|---|---|---|---|
| Holds | All data, raw, any format | Structured, modeled data | Both, in one tier |
| Structure applied | Later, when you query it | Up front, before loading | Flexible |
| Best at | Landing everything cheaply | Fast, trusted reporting | Collapsing the two |
| Role for an operator | Where every entity lands | Where “one number” lives | How the two get built together |
The short way to hold it in your head: the lake is where every acquisition’s data lands, and the warehouse layer is where it becomes a number the board can act on.
Why a multi-entity business needs one more than anyone
A single company with a small, stable systems footprint can sometimes postpone building a shared data foundation. A platform growing by acquisition does not have that luxury, and the reason is structural rather than a matter of effort.
Every company you buy arrives speaking its own dialect. Its own systems, its own account names, its own idea of what counts as revenue or a finished job. The same vendor shows up three different ways across three entities, and nothing shares a common key. Until all of that lands in one place and gets reconciled to a single definition, “consolidated revenue” is a guess dressed up as a number.
A data lake is what makes that reconciliation possible, and repeatable. Once the landing pattern exists, a new acquisition stops being a from-scratch project and becomes a defined onboarding: connect the sources, land the data, map it to definitions you’ve already built. The work compounds across deals instead of resetting with each one.
What it costs to skip it
The bill for a missing foundation shows up in a few predictable places. Finance rebuilds the same consolidation by hand every month. Operating teams spend the review reconciling definitions. Analysts repair one-off exports that cannot be traced back to the source. Leadership loses confidence because two reports answer the same question differently.
A weak foundation also limits AI. An agent connected to only part of the business can produce a polished answer from incomplete or unreconciled evidence. The model is not a substitute for shared definitions, lineage, permissions, and quality checks. It is why we say AI maturity is capped by data maturity.
”Can’t we just buy a dashboard or connect the ERP?”
This is the question that kills more data projects than any budget line, and the instinct behind it is fair. You already own a BI tool, or your ERP has a reporting module, so why stand up something new.
Because a dashboard inherits whatever sits beneath it. Point it at five disconnected systems and you get five versions of the truth rendered in nicer charts. Connect AI to that same partial picture and it answers confidently from data that was never reconciled, which is more dangerous than no answer because people believe it. The tool on top is only ever as trustworthy as the layer underneath, and for a multi-entity operator that layer doesn’t exist until someone builds it.
What good looks like
Done right, the data lake stops being a project and becomes the thing nobody thinks about because it just works. Every entity’s data, whatever system it came from, flows into one governed store. Each metric is defined exactly once, so revenue and job margin mean the same thing in every operating company, every dashboard, every report, and eventually every AI answer. A new acquisition is live in days, not quarters.
The part that matters most for a PE-backed platform is ownership. The lake, the models, the infrastructure, and the definitions are yours, not rented from a vendor who can change the terms or retire the product. Vendor-agnostic by design, no black box, nothing locked away. When the foundation is solid, the board gets numbers that tie out, the operating companies get reporting they trust, and the platform gets a base that makes every future deal easier rather than harder.
That’s the real case for a data lake. Not the technology for its own sake, but the one place your data has to live before any of the things you actually want from it become possible.
Frequently asked
Why does a company with multiple subsidiaries need a data lake?
Because each subsidiary usually runs its own systems, its own chart of accounts, and its own definitions. There is no shared key across them, so the same vendor or metric shows up several different ways. A data lake is the one place all of that lands and gets reconciled to a single definition, which is what makes consolidated reporting across entities possible.
What is the difference between a data lake and a data warehouse?
A data lake stores all of your data in raw form, applying structure later when you need it. A data warehouse stores structured, modeled data that is ready for reporting. In practice you use both: the lake is the landing zone every source flows into, and the warehouse layer is where that data becomes a number the board can trust.
How do private equity firms consolidate data across portfolio companies?
The reliable pattern is to land every portfolio company's data in one governed store, then model it to shared definitions so each metric means the same thing everywhere. Doing it by hand in spreadsheets does not survive the acquisition cadence. A data lake plus a defined metric layer turns each new deal into a repeatable onboarding.
Can't we just use a BI tool or connect our ERP instead?
A dashboard only reflects what sits beneath it. Pointed at several disconnected systems, it produces several versions of the truth in nicer charts. The reporting layer is only as trustworthy as the foundation under it, and for a multi-entity operator that foundation does not exist until it is built.
Is a data lake only useful if we are doing AI?
No. It pays off immediately in consolidated reporting and reliable KPIs. But it is also the prerequisite for AI, because models built on a partial or unreconciled view of the business return answers that are confidently wrong. The foundation serves both.
Building on a foundation that isn't there yet?
That's the gap we close. We stand up the warehouse, run it, and layer AI on once the base is solid. Built by us, owned by you.
Get in touch →