A dataflow view describes a computation through operations and the data that move between them. An operation can begin when its required inputs are ready; independent branches can run in parallel. This makes the arrangement of data part of the program’s structure rather than an incidental detail.

For machine learning, that structure extends from preparation to training and evaluation. Filtering examples, joining records, building features and combining statistics can demand as much planning as the model update itself. MapReduce, Hadoop and Spark offer related ways to organise this work across machines, with different abstractions for repeated computation and shared data.

Branching blue river channels flowing through green wetlands into open water.
Map, shuffle and reduce stagesMapShuffleReduce
Data is processed, grouped by key, and combined.

Map, group and reduce

The Google Research publication MapReduce: Simplified Data Processing on Large Clusters described a map function that transforms input key/value pairs and a reduce function that combines intermediate values sharing a key. Between these stages, intermediate records must reach the appropriate destination. This grouping step is often called the shuffle. It connects the logical grouping requested by the program to physical movement across workers.

The model separates a task’s data transformations from several responsibilities of the runtime. The publication described automatic partitioning, scheduling, inter-machine communication and handling machine failures. A program can therefore express local transformations and grouped aggregation while the system arranges their distributed execution. Counts and summary statistics fit naturally into this pattern because separate pieces of data can contribute to a combined result.

Communication still matters. If many records belong to one key, the associated work can concentrate on a small part of the cluster. If an operation produces a large intermediate collection, moving that collection may dominate the arithmetic. The dataflow description makes these dependencies explicit, but it does not make data skew or network costs disappear.

Hadoop connects storage and processing

Apache Hadoop’s project description presents a framework for distributed processing across computer clusters and places failure handling at the application layer. Its components have different responsibilities: the distributed file system provides access to stored data, resource management arranges cluster work, and MapReduce supplies a parallel processing model. Distinguishing these roles helps explain why Hadoop is more than the map and reduce functions.

Data placement affects the work that follows. A worker can process a local portion of a collection and send a smaller summary elsewhere, or a job can exchange records to bring related information together. Repeated movement carries a cost. Feature preparation that depends on joins or grouping may therefore have a different communication pattern from a training update that combines a compact set of model statistics.

Repeated computation and data in memory

Many learning procedures revisit the same data. A training algorithm may refine parameters over several iterations, while exploratory analysis may repeat different queries over a shared collection. Repeating the complete path from storage to transformation can make the organisation of intermediate data significant. An abstraction suited to a single batch task needs further thought when successive operations reuse earlier results.

The USENIX paper Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing introduced RDDs as a distributed-memory abstraction for fault-tolerant computation. It identified iterative algorithms and interactive data mining as motivating applications, and described coarse-grained transformations rather than fine-grained changes to shared state. The paper reported implementing the abstraction in Spark. The AMPLab research programme included Spark among its data-analytics projects.

An RDD represents a distributed collection on which transformations can be applied. The important distinction is between describing a collection’s transformation and treating every element as freely mutable shared memory. This restriction gives the execution system information about the work it can schedule. It also illustrates a wider systems trade-off: limiting the programming model can make coordination and fault tolerance more tractable.

From batches to event streams

Apache Spark’s project site describes an engine for data engineering, data science and machine learning on individual machines or clusters, including batch and streaming data. These activities can share transformations, but their timing differs. A batch operation works on a defined collection; stream processing uses a sequence of events as an ongoing input.

A stream raises questions about how much state remains active and when a result is complete. A time window, for example, gives an operation a bounded portion of the event sequence to consider. Event order and arrival order need not serve the same purpose. A learning process consuming a stream must also distinguish the production of features from the moment a label becomes available for an update.

Following the bottleneck through the graph

Dataflow planning asks where independent work is available and where dependencies force a meeting point. A wide branch of local transformations can still end in a costly global aggregation. Keeping intermediate data in memory can help only while the relevant working set fits the available resources. Fault tolerance adds another requirement: a failed stage needs a defined restart path whose result remains consistent with the intended calculation.

The graph-parallel approach exposes dependencies in the model’s own relationships, while learning inside databases moves attention toward managed tables and queries. These perspectives complement dataflow. Each asks how the representation of a computation helps the runtime organise work, limit communication and produce a useful result.