Lambda Architecture: Details, Examples and Purpose

Lecture



Lambda architecture is a data-processing architecture designed to handle massive quantities of data by taking advantage of both batch and stream-processing methods. This approach to architecture attempts to balance latency , throughput and fault tolerance by using batch processing to provide comprehensive and accurate views of batch data, while simultaneously using real-time stream processing to provide views of online data. The two view outputs may be joined before presentation. The rise of lambda architecture is correlated with the growth of big data , real-time analytics, and the drive to mitigate the latencies of map-reduce.

Lambda architecture depends on a data model with an append-only, immutable data source that serves as a system of record. It is intended for ingesting and processing timestamped events that are appended to existing events rather than overwriting them. State is determined from the natural time-based ordering of the data.

Computerized batch processing is the running of "jobs that can run without end-user interaction, or can be scheduled to run as resources permit."

Stream processing is a computer programming paradigm,equivalent to dataflow programming , event stream processing , and reactive programming , that allows some applications to more easily exploit a limited form of parallel processing . Such applications can use multiple computational units, such as a unit for floating-point arithmetic on a graphics processing unit or field-programmable gate arrays (FPGAs) , without explicitly managing allocation, synchronization, or communication among those units.

The stream processing paradigm simplifies parallel software and hardware by restricting the parallel computation that can be performed. Given a sequence of data ( a stream ), each element of the stream is subjected to a series of operations ( kernel functions )

Lambda Architecture: Details, Examples and Purpose

Lambda Architecture: Details, Examples and Purpose

Data flow through the processing and serving layers of a typical lambda architecture

Overview

Lambda architecture describes a system consisting of three layers: batch processing, speed (or real-time) processing, and a serving layer for responding to queries. The processing layers ingest from an immutable master copy of the entire data set. This paradigm was first described by Nathan Marz in a blog post titled "How to beat the CAP theorem", in which he originally termed it a "batch/realtime architecture".

Batch layer: the "cold path"

The batch layer precomputes results using a distributed processing system that can handle very large quantities of data. The batch layer aims at perfect accuracy by being able to process all available data when generating views. This means it can fix any errors by recomputing based on the complete data set, then updating existing views. Output is typically stored in a read-only database, with updates completely replacing existing precomputed views. : 18

As of 2014, Apache Hadoop was considered the leading batch-processing system. Later, other relational databases such as Snowflake , Redshift, Synapse and Big Query were also used in this role.

Speed layer: the "hot path"

Lambda Architecture: Details, Examples and Purpose
A diagram showing the flow of data through the processing and serving layers of a lambda architecture. Example named components are shown.

The speed layer processes data streams in real time and without the requirements of fix-ups or completeness. This layer sacrifices throughput as it aims to minimize latency by providing real-time views into the most recent data. Essentially, the speed layer is responsible for filling the "gap" caused by the batch layer's lag in providing views based on the most recent data. This layer's views may not be as accurate or complete as the ones eventually produced by the batch layer, but they are available almost immediately after data is received, and can be replaced when the batch layer's views for the same data become available.

Stream-processing technologies typically used in this layer include Apache Storm , SQLstream , Apache Samza , Apache Spark , Azure Stream Analytics . Output is typically stored on fast NoSQL databases.

Serving layer

Lambda Architecture: Details, Examples and Purpose
A diagram showing a lambda architecture with a Druid data store.

Output from the batch and speed layers is stored in the serving layer, which responds to ad-hoc queries by returning precomputed views or building views from the processed data.

Examples of technologies used in the serving layer include Druid , which provides a single cluster to handle output from both layers. Dedicated stores used in the serving layer include Apache Cassandra , Apache HBase , Azure Cosmos DB , MongoDB , VoltDB or Elasticsearch for speed-layer output, and Elephant DB , Apache Impala , SAP HANA or Apache Hive for batch-layer output.

Optimizations

To optimize the data set and improve query efficiency, various rollup and aggregation techniques are applied to the raw data , while estimation techniques are employed to further reduce computation costs. And while expensive full recomputation is required for fault tolerance, incremental computation algorithms may be selectively added to increase efficiency, and techniques such as partial computation and resource-usage optimizations can effectively help lower latency.

Lambda architecture in use

Metamarkets, which provides analytics for companies in the programmatic advertising space, employs a version of the lambda architecture that uses Druid for storing and serving both the streamed and batch-processed data.

Yahoo took a similar approach to the analytics of its ad data warehouse , also using Apache Storm , Apache Hadoop and Druid .

The Netflix Suro project has separate processing paths for data, but does not strictly follow lambda architecture since the paths may be intended to serve different purposes and not necessarily to provide the same type of views. Nevertheless, the overall idea is to make selected real-time event data available to queries with very low latency, while the entire data set is also processed via a batch pipeline. The latter is intended for applications that are less sensitive to latency and require map-reduce processing.

Lambda architecture is one in which data processing is split into two paths

  • The "cold path" is the batch layer, where all incoming data is stored in raw form and processed in batch mode. Typically, the batch layer is represented by big data lakes (Data Lakes) based on Apache Hadoop, which hold "raw" information that is not modified or updated but only supplemented with new data. This is where the so-called master data, or reference data, resides: business-critical information about customers, products, services, personnel, technologies, materials and other domain knowledge that changes fairly rarely and is not transactional. In the example considered, Machine Learning algorithms use batch data to segment customers and build predictive models by analyzing the stored history of user behavior. It is also important that this layer is the one used to process data on a schedule, i.e. to form batch jobs.
  • The "hot path" is the speed layer (or acceleration layer), where data is analyzed in real time. This layer provides minimal processing latency at the cost of some loss of accuracy. It is a collection of data stores in which information with a short life cycle is added to separate real-time views in order to avoid data duplication. Big Data stream-processing frameworks are usually used to implement the speed layer: Apache Spark, Storm or Flink.

ADVANTAGES AND DISADVANTAGES OF THE ARCHITECTURE

So, lambda architecture provides the following advantages for a Big Data system :

  • high durability of historical data, with a low probability of errors and failures;
  • a balance of speed and reliability;
  • scalability.

However, the λ approach has the following disadvantages :

  • the data analysis strategy cannot be changed "on the fly"
  • no BI tools
  • complexity

Criticism

Criticism of lambda architecture has focused on its inherent complexity and its limiting influence. The batch and streaming sides each require a different code base that must be maintained and kept in sync so that processed data produces the same result from both paths. However, attempting to abstract the code bases into a single framework puts many of the specialized tools in the batch and real-time ecosystems out of reach.

In a technical discussion of the merits of using a pure streaming approach, it was noted that using a flexible streaming framework such as Apache Samza could provide some of the same benefits as batch processing without the latency. Such a streaming framework could allow gathering and processing arbitrarily large windows of data, accommodate blocking, and handle state.

See also

  • Event stream processing
  • AWS Lambda
  • Kappa architecture
  • Data warehouses

Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Web site or software design"

Terms: Web site or software design