Vendia | Understanding real-time data lakes

Understanding real-time data lakes

Understanding real-time data lakes

Too many articles on data lakes cover abstractions, and the stories they tell can leave you thinking it will solve all your data problems (and world hunger, while you’re at it). They may also leave you paddling around in water puns. (Warning: We couldn’t resist it, either).

So, let’s take a fresh dip into data lakes by revisiting their foundational characteristics and answering some key questions:

Let’s start with some core concepts and work our way out from data itself, to storing it, to using it.

What is a data lake?

Data is a hot topic, so much so that it’s been compared to oil as the means for innovation and personalization. Companies that strategically leverage data will capitalize on their potential, and those that don’t are set to fall behind. And any time you are moving or consolidating data, it takes time. In the data world, faster is better than slower. What’s the best form of fast? Real-time.

Any data architecture, operational and analytical, has four core steps it must perform:

  1. Get your data
  2. Know your data
  3. Use your data
  4. Share your data

Data comes from…

And it can be…

…But that data always needs to be accessible for both driving the next step in an operational workflow and building analytical insights.

Naturally, to access our data, we must store it in an accessible way. (It’s a bit circular, but you get the point. Storing our data lets us get to our data.)

A data lake fits comfortably in Step 1 of a data architecture, “Get your data.” It brings all your data, across all forms and functions, without size limits, into one place.

NOTE: In some architectures, “getting” your data can be logical as it involves identifying the collection of sources and data stores (we see you virtualization purists and advanced data mesh architectures).

Data lakes and real-time data streaming

A data lake pools all your data into a single place. It removes the barriers, or dam stop gates, for getting data into a single location, a single store. And it does so by no longer needing to shape the data into the storage method’s limiting structure (e.g., rows and columns).

The flexibility of a data lake means you can load first, think later. Any data can be loaded into your data store. This flexibility removes the burden of needing to know your data or the questions you want to ask of your data before you acquire it. Data stores that require transformation prior to loading will cause information loss, either in data dropped from errors or information contained in the raw data that’s cleansed out.

By allowing all data forms, there’s no need to transform it before storing it; therefore, it can be consumed upon creation, fast. Real-time fast.

What type of data benefits from this flexibility?

Data lake architecture

The architecture of a data lake is what you make it. When storing heterogeneous data, it can be as organized or unorganized as you choose (even if your choice is inadvertent).

In theory, you can throw all your data into a lake, without thought (we don’t advocate for this).

To demonstrate the value of a data lake architecture for enabling use of our data, let’s leave the data world for a minute and think about our homes.

Imagine a home where all things—electronics, clothes, kitchenware, appliances, papers, and sports equipment—are piled into a single heap in a large room. Utter. Chaos. Good luck finding your passport.

This is representative of data lake architectures that toss all data, without thought, into your data lake.

Now, let’s imagine a slightly more organized world in which you organize these things by purchase date. Better, but still inefficient. You may have a toaster in the closet and a new shirt hanging in the fridge. So, every time you want to toast a bagel, you have to walk from the kitchen to the bedroom and back again to grab the knife to spread the cream cheese. Oh, and there’s your passport; tucked under the salad forks.

This is representative of data lake architectures organized only by update dates.

What we actually do is organize our homes by function (ignoring our messy days). We keep the toaster by the food, the shirts with the clothes, and our papers (mostly) together. We keep the frequently used items out and the infrequently used items (hello, fondue set) tucked away. This feels better. Much more efficient.

This is representative of data lake architectures organized by function (or source) and update dates.

What does this look like in a data lake? It’s a folder structure that organizes data according to the SaaS application which produced it, e.g., the CRM, ERP, CDP, web applications, IoT systems, etc. This structure can be further partitioned by the year, month, and day the data was created or updated.

To recap:

Data lake vs. data warehouse

A data lake, consuming data without limitation, requires a schema-on-read approach to add definition when using the data. (Ex: Our home, organized by function and frequency).

A data warehouse, requiring data to be structured, requires a schema-on-write approach to add definition upon storing it. (Ex: The use of items across our homes into outfits).

Data lake use cases

A data lake is a widely accepted staple in any complete data architecture. Any industry that has more than an inch of data can benefit from incorporating data lakes into the overall data architecture given the value for access to real-time heterogeneous data. For example:

All these use cases leverage semi-structured data that would rely on data lake technologies. All of which can be stored, and shared real-time, on Vendia.

The benefits of data lakes

Data lakes have notable benefits worth reiterating:

Challenges for data lakes

In addition to not solving the world hunger problems, there are some challenges with traditional data lake architectures: