Vendia | Understanding real-time data lakes
Understanding real-time data lakes
- August 3, 2022
Understanding real-time data lakes
Too many articles on data lakes cover abstractions, and the stories they tell can leave you thinking it will solve all your data problems (and world hunger, while you’re at it). They may also leave you paddling around in water puns. (Warning: We couldn’t resist it, either).
So, let’s take a fresh dip into data lakes by revisiting their foundational characteristics and answering some key questions:
- What exactly is a data lake and how can you use it?
- How does it fit among the other data-things, such as a data warehouse and virtualization?
- Is a data lake relevant for your use case?
- What are the challenges and benefits of a data lake?
- Vendia’s fit in it all
Let’s start with some core concepts and work our way out from data itself, to storing it, to using it.
What is a data lake?
Data is a hot topic, so much so that it’s been compared to oil as the means for innovation and personalization. Companies that strategically leverage data will capitalize on their potential, and those that don’t are set to fall behind. And any time you are moving or consolidating data, it takes time. In the data world, faster is better than slower. What’s the best form of fast? Real-time.
Any data architecture, operational and analytical, has four core steps it must perform:
- Get your data
- Know your data
- Use your data
- Share your data
Data comes from…
- single event-based record (e.g., closing a deal or receiving payment)
- recurring batch
- stream of messages
- bulk load from other data stores or SaaS applications
And it can be…
- Structured, in rows and columns
- Unstructured collection of emails or images
- Semi-structured messages, receipts
…But that data always needs to be accessible for both driving the next step in an operational workflow and building analytical insights.
Naturally, to access our data, we must store it in an accessible way. (It’s a bit circular, but you get the point. Storing our data lets us get to our data.)
A data lake fits comfortably in Step 1 of a data architecture, “Get your data.” It brings all your data, across all forms and functions, without size limits, into one place.
NOTE: In some architectures, “getting” your data can be logical as it involves identifying the collection of sources and data stores (we see you virtualization purists and advanced data mesh architectures).
Data lakes and real-time data streaming
A data lake pools all your data into a single place. It removes the barriers, or dam stop gates, for getting data into a single location, a single store. And it does so by no longer needing to shape the data into the storage method’s limiting structure (e.g., rows and columns).
The flexibility of a data lake means you can load first, think later. Any data can be loaded into your data store. This flexibility removes the burden of needing to know your data or the questions you want to ask of your data before you acquire it. Data stores that require transformation prior to loading will cause information loss, either in data dropped from errors or information contained in the raw data that’s cleansed out.
By allowing all data forms, there’s no need to transform it before storing it; therefore, it can be consumed upon creation, fast. Real-time fast.
What type of data benefits from this flexibility?
- Streaming data is particularly unknown at its genesis (e.g., data from IoT sensors, logs, web clicks, and other event-generating messages). The meat of the data is packed in nested JSON attributes, and the systems that produce this data do so continuously, in high volumes and velocity, i.e., emulating streams. This streaming data is fast, large, and undefined.
Data lake architecture
The architecture of a data lake is what you make it. When storing heterogeneous data, it can be as organized or unorganized as you choose (even if your choice is inadvertent).
In theory, you can throw all your data into a lake, without thought (we don’t advocate for this).
- Yes: You want to put all types (forms and functions) of data into a lake.
- No: You don’t want to do so without some definition.
To demonstrate the value of a data lake architecture for enabling use of our data, let’s leave the data world for a minute and think about our homes.
Imagine a home where all things—electronics, clothes, kitchenware, appliances, papers, and sports equipment—are piled into a single heap in a large room. Utter. Chaos. Good luck finding your passport.
This is representative of data lake architectures that toss all data, without thought, into your data lake.
Now, let’s imagine a slightly more organized world in which you organize these things by purchase date. Better, but still inefficient. You may have a toaster in the closet and a new shirt hanging in the fridge. So, every time you want to toast a bagel, you have to walk from the kitchen to the bedroom and back again to grab the knife to spread the cream cheese. Oh, and there’s your passport; tucked under the salad forks.
This is representative of data lake architectures organized only by update dates.
What we actually do is organize our homes by function (ignoring our messy days). We keep the toaster by the food, the shirts with the clothes, and our papers (mostly) together. We keep the frequently used items out and the infrequently used items (hello, fondue set) tucked away. This feels better. Much more efficient.
This is representative of data lake architectures organized by function (or source) and update dates.
What does this look like in a data lake? It’s a folder structure that organizes data according to the SaaS application which produced it, e.g., the CRM, ERP, CDP, web applications, IoT systems, etc. This structure can be further partitioned by the year, month, and day the data was created or updated.
To recap:
- A data lake brings your data together, but it doesn’t define your data.
- Data architectures that organize our data increase the usability of data in our data lake.
- Having data be generated from a shared schema (hint: a schema defined in a Vendia Share Uni) enables us to better use our data.
Data lake vs. data warehouse
A data lake, consuming data without limitation, requires a schema-on-read approach to add definition when using the data. (Ex: Our home, organized by function and frequency).
A data warehouse, requiring data to be structured, requires a schema-on-write approach to add definition upon storing it. (Ex: The use of items across our homes into outfits).
Data lake use cases
A data lake is a widely accepted staple in any complete data architecture. Any industry that has more than an inch of data can benefit from incorporating data lakes into the overall data architecture given the value for access to real-time heterogeneous data. For example:
- Airlines with semi-structured ticket, passenger, and flight information in nested JSON objects
- Financial service loan servicing with recurring payments and mortgage file uploads
- Patient medical records with physician notes, lab images, and prescriptions
All these use cases leverage semi-structured data that would rely on data lake technologies. All of which can be stored, and shared real-time, on Vendia.
The benefits of data lakes
Data lakes have notable benefits worth reiterating:
- Flexibility in various data formats, including scalar or files
- Scalability in the size, amount, and type of data without restriction
- Speed to consume and store with the removal of transformation barriers
- Real-time access to data in its original format
Challenges for data lakes
In addition to not solving the world hunger problems, there are some challenges with traditional data lake architectures:
- Undefined data with risk for being unusable
- Vendia’s use of a schema definition across a uni resolves this by creating an aligned understanding of what the data contains.
- Security and access to the data store
- Vendia creates a secure GraphQL API interface with the availability to secure who has access to what data with role-based and fine-grained access controls.
- Consistency of data across different companies’ data lakes
- Vendia manages distributed data stores that are fully synchronized in not only the data but the history of changes over time.