← Back to list

The 5 challenges with Data Lakes..

If you ask any CIO or Business manager what is top of their mind right now they will all tell you it is about defining the Digital agenda…

SGK · 2016-09-27 15:07 · 153 claps · 5.2 min read
#big-data #data-science #bdaas #digital #analytics
Open on Medium ↗
Wiki topics: ML · Machine Learning CRY · Crypto & Web3 GRW · Growth & Analytics 🔧 · Data Engineering 🔬 · Science · General

The 5 challenges with Data Lakes..

If you ask any CIO or Business manager what is top of their mind right now they will all tell you it is about defining the Digital agenda for their enterprise.

But what good is any Digital strategy without the Data?

Its all about the Data

Your Digital strategy will require you generate insights and analytics so that you can develop that that 6th sense about your customer. For that you will need to pull data from all your back end business applications and operation data stores. But even that is not enough — because consumers are now always connected on their phones, responsive and real-time visibility is the norm. For that you will need to ensure that your systems are exposed as APIs providing real-time data whether that is an IOT device like a water injection pump on an Oil rig or the on-board telematics inside your truck distribution network. Finally there is all the data that you don’t own, public data sets, social media streams and all that unstructured data that you will need to somehow collect to truly build your digital enterprise.

What is the idea behind the Data Lake?

And so the Data Lake was born. Metaphorically speaking it was an idea that was very attractive to the business since they finally had this idea of all the data in one place — you should needed a single Big Data Platform in which to store it.

What we will look at are 5 different approaches on how you can deploy a Data Lake, but first all what is the purpose of the Data Lake:

“A Data Lake is a place to store lots of data until you decide what to do with it…”

Throw all the data into the Lake #1:

For: The basic principle was to be very liberal with all enterprise data — structured and unstructured data put it into the data lake avoiding the silo constraints of typical EDWH. By putting into the HDFS file-system in its original format, there was no need to spend time doing any transformation. By having all your enterprise data in a single data lake would result in increased information usage. Similarly you would benefit from not investing time on pre-determined schemas, unlike a EDWH, this is called late-binding on query execution or schema-on-read.

Against: But most enterprises are still struggling with the immaturity in how its used, how its governed and even how its defined. I should add this is a problem for any enterprise platform service, but especially acute in a technology stack that continues to change quite a bit. The other problem is a lack of certainty about data quality, or a metadata repository describing the raw data or what analysts found interesting about that data. *When all your data is accessible on the data lake, you need to factor concerns about security and access control on the data (more detail needed here)

Hybrid Data Lake as an Agility Layer #2:

In the hybrid data lake approach — we take a little more of a pragmatic view.

For: What this means is that the Data Lake is treated as an Agility layer. The idea is to enable a shorter time to discovery. The Data Lake acts just as a sandbox. Once the business has found something interesting about the data it then moved more formally into your EDW, so the Data Lake is your staging area. All the data governance and cleansing happens as part of the move to the Data Lake. you can use a technology like Apache Falcon as the data management during the data ingestion and discovery phase. * Each Business Unit can have their own Data Lake without worry about issues around security, governance or operations.

Against: Its a pretty expensive sandbox. What is the incentive for the business to promote the data into the EDW? *

Prepare before pre-ingestion into the Data Lake #3:

In the pre-ingestion model, you need to do some preparatory work on the data before ingestion.

**For: **What that means is ensuring that the data follows a standardised schema (e.g. AVRO, ORC, Parquet, JSON). This also means some effort needs to be put to profile the data, refine it, enrich it before its moved to the Data Lake. This requires establishing governance to review any changes to the schema. security concerns can also be addressed by adding privacy meta-data beforehand.

Against: *Its hard to fault this model. There are still some challenges around governance of the Lake, but for the most part are addressed with the security meta-data.

A multi-step Data Lake #4:

We basically build upon the Hybrid Data Lake — except its not a Hybrid, the final storage governed and managed Data Lake.

For: Each Business Unit can have their own Data Lake sandbox without worry about issues around security, governance or operations. What this means is that the Data Lake is treated as an Agility layer. The idea is to enable a shorter time to discovery. Once the business has found something interesting about the data it then moved into a formally managed Data Lake.

Against: What is the incentive for the business to promote the data into the EDW?

So what is the 5th reason why businesses don’t care?

The 5th problem actually holds true everything that has been discuss so far, and that is one related to technology. Data Lakes are based on Big Data technologies like Map Reduce / HDFS / Yarn / Hive / Stinger / Spark / Spark Streaming / Shark / SQL / Pig / Tez / Storm / Falcon / Kafka / GraphX / Mesos / Tachyon.

You get the point? opensource technologies are part of the Big Data distributions with each one betting on slightly different technology stacks and through an almost Darwinian process de-facto standards develop over time.

That is a lot of headache for a business to handle, its rarely the case that you can just deploy your Data Lake platform and assume its mature, its already out of date by the time you have deployed it, so you will need to continue to tinker and update and upgrade it and ensure all your jobs still work and nothing is broken.

So enterprises what to be abstracted away from that problem as much as possible.

What the Business actually cares about is Outcomes.

That could be technology outcomes like a Big Data as a Service (BDaaS) or a business outcome like Analytics as a Service (AaaS).

If you can solve the technical data platform, the data ingestion problem, the transformation problems, the data governance and master data issues and can enable the semantic layer, discovery and search — then the customer will bite your hand off. And all you need to provide is self-service analytics.


메타데이터
post_id
806e2397b30c
slug
5-reasons-why-business-doesnt-care-about-data-lakes-806e2397b30c
url
https://medium.com/@captaink99/5-reasons-why-business-doesnt-care-about-data-lakes-806e2397b30c
canonical_url
https://medium.com/@captaink99/5-reasons-why-business-doesnt-care-about-data-lakes-806e2397b30c
author_url
https://medium.com/@captaink99
status
ok
fetched_at
2026-07-30 16:59:00