← Back to list

Databricks CustomerLake: Gravity wins

During Databricks’ Data + AI Summit, the company officially announced CustomerLake, its “Agentic CDP” offering. The move was bold. The…

David Chan · 2026-06-18 13:39 · 0 claps · 8.8 min read
#cdp #databricks #customerlake #composable-architecture #agentic-ai
Open on Medium ↗
Wiki topics: AGT · AI Agents 🔧 · Data Engineering 🏛️ · Architecture

Databricks CustomerLake: Gravity wins

Databricks CustomerLake

Databricks CustomerLake

During Databricks’ Data + AI Summit, the company officially announced CustomerLake, its “Agentic CDP” offering. The move was bold. The implications are far-reaching. And for people hearing about it for the first time, many are still trying to process what it all means.

Let’s take a look at the actual press release titled “Introducing CustomerLake: The Agentic CDP embedded in Databricks”.

Today at Data + AI Summit, we’re announcing Databricks CustomerLake, a new Agentic Customer Data Platform (CDP) natively embedded in Databricks. CustomerLake brings core CDP capabilities, including Customer 360, identity resolution, audience building, campaign automation, activation, and personalization, directly into the lakehouse where customer data, AI models, and governance already reside.

With CustomerLake, marketing and data teams work together on a shared, governed foundation to turn customer data into always-on, 1:1 customer experiences. Instead of relying on manual campaign work and disconnected systems, marketers can deploy agents that continuously analyze behavior, decide, and act, delivering intelligent engagement at enterprise scale without creating new silos, duplicating sensitive data, or adding martech complexity.

That sounds simple enough. But a lot had to go into Databricks’ decision to enter the CDP space.

  1. First, the CDP market is crowded. According to the Customer Data Platform Institute, there were approximately 200 CDP product vendors as of July 2025.
  2. Second, this move will likely create some angst across Databricks’ partner ecosystem, especially among partners that already compete in the CDP space.
  3. Finally, it cuts against the grain of the typical cloud business model: a consumption-based platform supported by a healthy ecosystem of partners identifying use cases that drive more consumption.

That last point is worth unpacking.

For those less familiar with the typical cloud business model, cloud companies make money based on how much customers use their products. The best signal for usage is consumption. Cloud companies typically avoid going too far downstream into application development because app-level bets carry more risk. It is much easier to provide the platform and let a community of builders take on the risk of developing applications on top of it.

The easiest way to think about this symbiotic relationship is to frame Databricks as the operating system (OS) and its ecosystem partners as the develops of applications in the App Store. The more valuable the apps are, the more valuable the OS becomes, and vice versa.

There is risk on both sides — the OS and the apps — but rarely (though not never) do we see the OS venture this deeply into the app side with conviction.

So, why now?

This is probably the most common question I’ve heard discussed on this announcement. I don’t know the answer, but I have a theory.

It is not just that Databricks chose to enter the CDP space. It is also the positioning: Agentic CDP. The launch announcement is as much about riding the Composable and Agentic CDP wave as it is about promoting AI-ready data that can power agents to automate campaigns at infinite scale. So why didn’t Databricks wait patiently on the sidelines like its peers and let the composable ecosystem continue trying to solve this? My guess is that Databricks looked at the market and saw a lot of architectural agreement, but not enough operational breakthrough.

I think it is worth revisiting the three decision points above.

  1. Yes, the CDP category is crowded. But despite several years of waiting for the dust to settle, there are still no clear winners. Just more pivots, reframing, and repackaging of narratives.

Data is messy. And despite a free market that has supported approximately 200 CDP vendors — vendors that have collectively invested an estimated $1.2 billion to $1.5 billion in product R&D over the last five years — one could argue we are only marginally closer to unraveling the mystery that is customer data. Databricks likely sees that scenario as an opportunity to slice through Gordian Knot where others have failed. 2. Yes, Databricks’ ecosystem partners may not be thrilled about the company crossing the line and entering the martech zone. But let’s do the math on impacted partners. Databricks has over 6,000 partners. Approximately 500 are ISVs. Of those, roughly 20 have a CDP offering. That translates to about 4%.

Four percent is not insignificant, but it brings the conversation back down to earth. Does this new Agentic CDP generate more revenue than what those 4% of partners could drive? My guess is someone at Databricks crunched those numbers and scenario-planned the expected tradeoffs and greenlit it. 3. Yes, Databricks as an OS probably should not be investing its time in app development. But sometimes you watch the market struggle to make progress, and you might decide you need to pave your own path to generate that AI demand.

I once met a successful businessman who offered frozen yogurt supplies to a dozen or so froyo shops. He made good money doing it. He explained that he originally got into the business by opening standalone shops. But as he expanded, it became a headache to manage the finance and operations of multiple locations. Eventually, he realized he could make more money — and deal with fewer headaches — by supplying the raw materials to other shops. But he couldn’t rely on the individual success of the shops to grow demand. Therefore he would continuously run one to two shops himself to demonstrate business success, then sell those businesses, and in doing so, generate his own demand for frozen yogurt supplies.

For me, that last point is the key takeaway. I don’t think Databricks actually wants to be in the CDP business. I think CDP is a means to an end. And that end is AI.

Databricks’ heritage has always centered on enabling the Spark engineer driving compute-heavy analytical workloads. For Databricks to sell more apps on its OS and more frozen yogurt supplies — it needed someone to solve the CDP / Customer 360 / trusted-and-governed-data puzzle so it could unlock more compute-heavy agentic workloads.

And sometimes when you need something done right, you just have to do it yourself…

So what has Databricks done? It brought in basically the core team from ActionIQ, led by former CEO Tasso Argyros, to finish what they started. ActionIQ pioneered one of the first iterations of what we now call composable / zero-copy CDP through “Hybrid Compute.” It was a fascinating concept at the time, and ActionIQ was one of the few vendors actively promoting that model in the market. But I think, at times, even they might admit the solution suffered from scalability challenges.

Those challenges basically go away when you no longer need two different systems seamlessly to interoperate with each other. Now the two systems have collapsed into one: Databricks, as the main data gravity.

Three takeaways and questions

  1. This is good for brands buying these capabilities — Implementing a CDP is hard. There are no ifs, ands, or buts about it. Tech complexity. Organizational alignment. Process adherence. Ownership. There is a reason that every couple of weeks someone publishes another article about why CDP projects fail or questioning what the value of CDP is.

It is hard because integrating different systems, often with different owners, requires teams to rationalize different opinions, different processes, and different definitions of truth.

Databricks’ CustomerLake compresses this. Maybe it does not collapse it entirely, but it removes some of the friction and people-based middleware from the equation. And that is something we should celebrate. 2. What will Databricks’ competitors do? — Not to be a conspiracy theorist, but I’ll ask the question: was an unspoken rule broken? Since I started tracking the CDP space about 10 years ago and conducting my own due diligence in the category, I never fully understood why the cloud hyperscalers did not simply enter the space directly. The adjacency always seemed logical. I know some of them conducted their own market research and analysis. But almost all of them refrained from officially launching a CDP offering, with Microsoft being the exception.

Even then, Microsoft’s offering was not marketed aggressively. And it always felt more like CRM + Insights on steroids. While the rest of the CDP category was evolving and recalibrating, Microsoft stayed the course and did their own thing.

So the open question is whether hyperscalers will fast-follow to reach parity, similar to how airline loyalty programs respond to competitor changes. My feeling is that most will take a wait-and-see approach to understand how the market reacts, even if some stealth initiatives get kicked off incognito. 3. The Agentic CDP features Databricks launched with are telling — For years, people have debated what is a CDP and what is not. Databricks had the opportunity to choose which MVP CDP features would be announced at launch. It had to find the right balance. Too many, and it would take too long to get to GA to show value. Too few, and it would look like a poor man’s CDP and fail to be compelling. Looking at what Databricks ultimately launched with gives us insight into which capabilities it considers most valuable and most necessary for a CDP.

And this is even more interesting because it is coming through the lens of Tasso and his experience at ActionIQ. He is getting a chance few get: to reboot the vision, but with the benefit of lived experience. Edge of Tomorrow-style. Yes, I know I could have gone with Groundhog Day.

CustomerLake launches with Customer 360, Identity resolution, Audience building, Campaign management (automation), Measurement and insights, and a robust catalog of Agents.

On that last point, let’s discuss what was was and wasn’t surprising.

CDP capabilities included at launch that are NOT surprising

  • Agent catalog — This is an Agentic CDP after all.
  • Audience builder — In my opinion, this is a key capability that enables business self-service and has become table stakes for a CDP solution. It is now widely accepted as a core CDP feature and one of the first capabilities that traditionally lived in the Martech zone but now gravity has pulled it closer to the data.
  • Customer 360 — This is what you would expect to exist in the gold layer of an enterprise data platform like Databricks
  • Measurement and insights — This is what you would expect to exist in the gold layer and/or metrics layer of an enterprise data platform like Databricks

CDP capabilities included at launch that ARE surprising, from least to most

  • Reverse ETL — Databricks appears to intend to productize a downstream destinations catalog that integrates with martech, adtech, and other systems. This is not just dropping files into a cloud storage bucket or exposing an API. This is a pre-built integration library!
  • Campaign management automation — This is huge. If this is a campaign workflow canvas that enables automated campaign orchestration, then it may be one of the first, if not the first, of its kind. It would be unique because it would be a native feature embedded directly in Databricks. For all the talk about embedded CDP, I think embedded campaign management and orchestration would be a clear runner-up.
  • Identity resolution — This is very interesting. I consider myself pretty knowledgeable about the identity space, and for years cloud hyperscalers have mostly left this area alone, relying on MDM solutions or third-party identity vendors to fill the gap. AWS was the only one recently to break that mold by introducing Entity Resolution as a native service. But as you can tell from the intentional branding of entity vs identity. Entity means it looks to solve all domains.

Databricks, however, is going right at identity resolution, meaning it is targeting individuals. People. Customers.

It is choosing to invest in solving what I believe is one of the most important problems in customer data.Yes, data is messy. And yes, trying to clean it and govern it is important. Yes, yes, and yes…But if you cannot establish a primary identifer to link and relate records together so you know who an individual actually is, then you are still stuck in the same place. Identity resolution is one of the most important capabilities needed to unlock Customer 360 and make the CDP valuable. Similar to AWS Entity Resolution, I am not expecting too much out of the gate. But if Databricks can help companies solve this natively in-platform, that would be pretty awesome.

FINAL THOUGHTS

I would be lying if I told you I knew whether Databricks made the right move. But one thing is clear:

The CDP is not dead.

If anything, the last few years of the emergence of composable / integrated / embedded CDP proved the opposite. It was never a sign that CDP did not matter. It was a sign that the old architecture was under pressure. Gravity wins because the single view of the customer problem keeps pulling more functionality back toward the enterprise data platform. The AI and Agentic problem keeps pulling more conversations back to whether the data is clean, trusted, permissioned, and actually usable. Real-time personalization keeps pulling more activation, decisioning, and orchestration closer to the single source of truth. At some point, the CDP stops looking like a category waiting to be 86’d and starts looking more like a capability so foundational that the fact multiple vendors are trying to embed it into existing systems is not a signal about commoditizing a simple feature — it is about acknowledging that the the CDP is so essential that AI, agentic, and real-time personalization story does not actually work without it.

Long live the CDP.


메타데이터
post_id
1d4fbb19d867
slug
databricks-customerlake-gravity-wins-1d4fbb19d867
url
https://medium.com/@iamdavidchan/databricks-customerlake-gravity-wins-1d4fbb19d867
canonical_url
https://medium.com/@iamdavidchan/databricks-customerlake-gravity-wins-1d4fbb19d867
author_url
https://medium.com/@iamdavidchan
status
ok
fetched_at
2026-06-20 20:29:01