← Back to list

Stop Building Data Lakes. Start Building Data Products.

The data lake holds fifty petabytes. The CEO asks for a customer churn prediction. The team quotes six weeks. This is not a technology…

Rotimi Ademola · 2026-04-18 13:32 · 5 claps · 5.3 min read
#data-lake #product-data #data-value #ai-agent #data-mesh
Open on Medium ↗
Wiki topics: AGT · AI Agents BIZ · Business Strategy GRW · Growth & Analytics 🔧 · Data Engineering

Stop Building Data Lakes. Start Building Data Products.

The data lake holds fifty petabytes. The CEO asks for a customer churn prediction. The team quotes six weeks. This is not a technology problem — it is a mental model problem, and it is the single largest constraint on data value in most enterprises today.

The lake metaphor has run its course

Data lakes were sold as the single source of truth. In practice, they became warehouses of unstructured liability — repositories where data goes to be forgotten. Gartner warned of the data swamp as early as 2014, noting that a lake accepts any data without oversight or governance, and that without descriptive metadata, every subsequent use of data means analysts start from scratch. More than a decade on, the warning reads as prophecy.

The mental model is the issue. The lake metaphor encourages hoarding — store everything now, extract meaning later. But meaning does not emerge from volume. It emerges from intent. A data lake treats data as inventory. A data product treats data as inventory with a purpose, a named owner, and a contractual promise to consumers. Those distinctions change everything downstream.

Why lake thinking erodes credibility

Ingestion is the optimised dimension of a lake. Petabytes stream in, storage scales elastically, and the capacity charts look impressive. Then the business asks a simple question — “How many active customers did we have last quarter?” — and the day disappears into reconciliation.

Three tables contain “customer.” Four definitions of “active” circulate across domains. Two upstream systems have not synchronised in weeks. A data scientist, three months in post, has no idea which source to trust. The lake solved the storage problem. It created a discovery-and-trust catastrophe.

This is not hypothetical. It is the daily reality of most enterprise data teams, and the reputational cost is severe. Business stakeholders do not care about ingestion throughput. They care about getting the correct answer today. Every “let me check” and every “that depends on which table you use” compounds into a conclusion: the data team is a cost centre that slows things down. For data leaders, that reputation is career-limiting, and the projects that might reverse it rarely secure funding.

The data product alternative

A data product is not a dataset with a README. It is a contractual commitment — an interface, an owner, and a guarantee.

The shift from storing everything to serving specific needs with guarantees transforms the relationship between the data team and the business. A team with products is no longer the team that has data — it becomes the team that delivers reliable answers.

Gartner’s 2025 Hype Cycle for Data Management positions data products as critical for data and analytics success precisely because they provide integrated and prepared data that is easily findable, trusted, self-contained, and certified for reuse. The vocabulary is the vocabulary of software engineering, and that is not an accident.

Why the transition is urgent now

Three converging pressures have tipped the economics decisively.

AI at production scale. Models require curated, trustworthy inputs. A churn model cannot be trained on three conflicting customer tables and produce useful predictions. Gartner predicts that through 2026, organisations will abandon sixty per cent of AI projects unsupported by AI-ready data, framing AI readiness as a data contract problem rather than a modelling problem. Data products supply the bounded, contract-governed inputs that production AI demands. Lakes do not.

Regulatory pressure. GDPR, CCPA, DORA, and the EU AI Act all converge on the same requirement set: documented lineage, quality attestation, demonstrable accountability, and auditable retention. In financial services, BCBS 239 has required risk data lineage for more than a decade. A lake creates liability — vast, undocumented, and unevenly governed. A data product creates governance artefacts by construction, because the contract is the artefact. The regulatory direction of travel is not ambiguous.

Platform economics. Modern cloud data platforms — Snowflake, Databricks, BigQuery — have made storage cheap and complexity expensive. The cost of finding and fixing issues across a sprawling lake now exceeds the cost of building properly governed products upfront. Gartner predicts that by 2027, eighty per cent of data and analytics governance initiatives will fail for lack of a real or manufactured crisis. In lake-heavy estates, the crisis is already well established — it is simply not named.

How to start without draining the lake

The lake does not need to be dismantled overnight. Products can be built from it incrementally in five steps.

Identify high-value domains. Which business questions cause the most friction? Which datasets support three or more downstream teams? That is the product backlog. Early candidates in financial services typically include customer master, product reference, positions, trades, and regulatory reporting aggregates.

Assign ownership. Each product needs an accountable owner for schema, quality, and consumer satisfaction. Not a committee — a named individual. The ownership question is where most product programmes stall, and it is rarely solved by an organisation chart alone.

Define the contract. Schema, freshness SLI, quality rules, versioning policy, and deprecation terms — documented, version-controlled, and machine-readable. The Open Data Contract Standard (ODCS), maintained by the Bitol project under the Linux Foundation AI & Data umbrella and originally contributed by PayPal, provides a mature YAML specification that removes most of the argument about format.

Build the interface. A table, a view, an API, a semantic layer — the physical shape is secondary. What matters is that consumers know how to access it, how to interpret it, and what to trust.

Iterate on usage. Instrumented consumption tells the team which products matter. Products no one uses become inventory again, and inventory should be deprecated. The discipline is cultural, not technical.

The PayPal reference point

PayPal is the most frequently cited public reference for data products at scale. Jean-Georges Perrin, who led the mesh implementation, has documented the programme on PayPal’s technology blog and described on the Data Engineering Podcast how the team designed the platform for thousands of data quanta from inception rather than a handful. Domain teams own their products, contracts define cross-team interfaces, consumers discover products in a catalogue rather than hunting through storage, and quality incidents route to product owners rather than a central queue.

The compliance and auditability drivers at PayPal transfer directly to most financial services estates. Where a lake-based estate requires reconciliation to answer a regulator, a product-based estate already has the lineage, ownership, and quality attestation in its contract metadata. The reconciliation does not disappear — it becomes a first-class artefact rather than a quarterly emergency.

The mental model is the lever

The data lake was the correct architecture for 2015. Centralised storage solved a real problem at a moment when storage economics still dominated the design conversation. But the mental model the lake installed — hoard now, structure later — is now the primary obstacle to data value.

In 2026, credibility comes from delivery, not storage. Boards do not fund lakes. Boards fund outcomes. Data products are the architectural bridge between infrastructure and outcomes, with ownership, contracts, and the trust that follows from both.

Data leaders’ career trajectories are diverging along this line. The teams building products are being pulled into strategic conversations. The teams maintaining lakes are being positioned as legacy support. Which side of the transition a team lands on is a choice about the mental model more than about technology.

Where to begin

Start small. Pick one high-value domain. Name one owner. Write one contract. Prove the model against one real consumer with one real SLA. Then expand.

The lake is not the enemy. The model that built it is. Evolve the model, and everything downstream improves — the ML pipelines, the regulatory posture, the stakeholder trust, and the team’s place in the business.


메타데이터
post_id
a9305dbb5741
slug
stop-building-data-lakes-start-building-data-products-a9305dbb5741
url
https://medium.com/@arrufus/stop-building-data-lakes-start-building-data-products-a9305dbb5741
canonical_url
https://medium.com/@arrufus/stop-building-data-lakes-start-building-data-products-a9305dbb5741
author_url
https://medium.com/@arrufus
status
ok
fetched_at
2026-06-12 22:02:08