The Open Data Platform, They Say
We are the Gen Z of data practitioners. We joined the field when there was only BIG data — no small, no medium. Hadoop was something from…
The Open Data Platform, They Say
We are the Gen Z of data practitioners. We joined the field when there was only BIG data — no small, no medium. Hadoop was something from the past, like video cassettes or dial-up internet. We were ready to charge forward, eager to embrace the future of data.
For us, the early days of the data lake represented a new frontier. It was the ultimate mix-and-match era — use any file type, spin up a compute engine, and congrats — your data lake was coming together. That used to be the bare minimum, back when the world was naive and simple.
“Gen Z of data practitioners” by ChatGPTs prompt suggestions.
Soon enough, we realized this was hardly the beginning. There were many moving parts to integrate, and many options to choose from — metastores, pipeline orchestrators, data ingestion tools, and many more. But worst of all, these odd compositions had to work in unison, which was not always trivial. It was essential to build infrastructure to tie everything together in order to serve the different users and use cases.
When open table formats (OTF) emerged, new horizons opened. We then realized that many of the limitations of data lakes were bound to diminish. Things were not seamless yet, but we could definitely see the light.
OTFs were the first step in making the data-lake closer to the database-like-experience. After crossing that bridge, more categories swept into the domain, like Access Management, Data Quality, Lineage, and Data Catalogs. All paving the way to what we now know as the data lakehouse - a prominent modern architecture for data platforms.
This post was not generated with genAI, but all the art was.
During these times, Databricks has grown to be the dominant platform for data lakes, with Snowflake as its strong alter-ego. Customers opting for Snowflake may have traded some aspects of the openness of their data platforms, but have gained the simplicity and robustness of this new generation data warehouse.
Over time, we witnessed how each of these two rivals grew to serve as full-blown data platforms. They aspired to fulfill every need in a one-stop shop — different compute setups, internal accelerations and built-in governance. Entire categories have collapsed into these two parallel worlds, hoping to sweep us all away from the chaotic world of self managed platforms.
And then, all this Iceberg madness :) With Delta Lake being mostly contributed by Databricks, Iceberg was a whole different story, and the adoption was wild.
Driven by the community, these major players were urged to unlock their gates. We saw intent from both directions, but for a while it was still hard to figure out where it was going. Recently, this intent accelerated into two dramatic announcements. First Snowflake’s release of Polaris, their open source Iceberg catalog, and no later than 24 hours, Databricks announced their acquisition of Tabular. What’s mind blowing about this, is that both products’ releases were full of mentions for multi-engine support, integrations with other vendors and interoperability. There’s still a lot to uncover, but what’s certain is that this is the next step in our relationship with those two giants. Now that’s real madness.
These moves may lead to an alignment around OTFs and the catalog spec, but is that enough? Will this make our data platforms interoperable and easier to maintain? Well, in a sense. Remember that with Polaris on one end, and Unity Catalog + Tabular on the other — we now have the open source foundations (Delta Lake & Iceberg) controlled by commercialized giants with specific agendas on top of the openness of our platforms.
These moves can play out in many ways — Should we expect a change of licensing? Any proprietary features? Would Delta Lake and Iceberg consolidate? What will be the future of Unity Catalog and Tabular? Will Polaris adhere to the OTF spec or will it serve Snowflake alone?
What is certain now is that “multi-engine” is not only real, but essential.
Data Platforms must satisfy many unique requirements and polarities — data scientists vs. analysts vs. data engineers, internal vs. customer facing analytics, ad-hoc querying vs. dashboards, model training vs. ETLs. Basically, every query is a snowflake, right? Like a real snowflake I mean. There’s different context, different SLA, different data scanned, different operations we wish to perform and different cost/performance balance requirements. Going all in with a single solution might feel right, but also raises many existential questions. Eventually, we all want simplicity, but also a price we can easily manage.
But let’s face it, managing a multi-engine stack is hard. These challenges span from technical, to organizational and sometimes even political hardships. You need people, you need expertise, there are many dependencies and opinionated users and stakeholders to keep happy. All this while not interfering with production or dev velocity. There’s a reason why changes in this area can take anywhere from months to years.
So perhaps, it’s the sad story of the multi-engine? Can’t live with it, can’t live without it.
True interoperability lies in our ability to operate multi-engine platforms seamlessly, in a way that infra migrations would seize being complicated events that require adjustments and preparations across entire engineering organizations.
When this elasticity and interoperability become truly workable, our data lakes would expose new possibilities for a better data experience and greater innovation.
It’s time for THIS to become easier, and truly manageable. To me this feels like the beginning, or the rebirth of the modern data stack.
If multi-engine data lakes is something you’re excited about, I’d love to hear from you! DM me on LinkedIn.
메타데이터
- post_id
- 5d3bbab66271
- slug
- the-open-data-platform-they-say-5d3bbab66271
- url
- https://medium.com/@doron_72556/the-open-data-platform-they-say-5d3bbab66271
- canonical_url
- https://medium.com/@doron_72556/the-open-data-platform-they-say-5d3bbab66271
- author_url
- https://medium.com/@doron_72556
- status
- ok
- fetched_at
- 2026-06-25 07:00:49