What “big learning” means
Big learning asks how machine learning changes when data, models or repeated computation exceed the practical capacity of a familiar sequential approach. A problem can be large because it contains many examples, because each example has many features, or because the model requires expensive inference. Scale is therefore a relationship between a task and the resources available to complete it.
The useful question is how much learning can be achieved within a computing budget. More storage makes additional observations available, but using them still requires reading, transforming and analysing them. A method that works on a small collection can face a different balance of costs when it operates on terabytes or petabytes.

When computing time becomes the limit
Léon Bottou’s 2010 paper Large-Scale Machine Learning with Stochastic Gradient Descent described data sizes growing faster than processor speed and argued that computing time could become the limiting factor for statistical learning. That changes the comparison between algorithms. A method that reaches a precise solution through expensive repeated passes may be less useful under a fixed time budget than one that makes cheaper, approximate progress.
Small-scale reasoning often starts with a fixed dataset and asks how accurately an optimisation problem can be solved. Large-scale reasoning also asks how much data can be processed before the budget is exhausted. Statistical uncertainty, model choice and numerical error all remain relevant, but they interact with the cost of the algorithm. Spending additional effort on optimisation has little value if the underlying observations cannot support the apparent precision.
Large data, large models and difficult structure
The NeurIPS description of the 2011 Big Learning workshop named bioinformatics, astronomy, recommendation systems, social networks, computer vision, web search and online advertising as application domains. It also made room for computationally intensive models and structured prediction. The workshop’s central question was what changed when learning faced terabytes or petabytes of data, rather than merely how to store those data.
Those domains expose different pressures. A recommendation task can involve a wide space of relationships; an image task can require substantial numerical computation; a stream can keep arriving while an existing model is still being updated. These are task characteristics, not a promise that one system fits every setting. The number of observations alone says little about the arrangement of data, the dependencies within the model or the timing of useful answers.
What scalability promises
Wikipedia’s scalability article defines scalability around handling a growing amount of work, sometimes by adding resources. For learning, the growing workload might be a larger training collection, a wider feature representation or additional experiments. The extra resources might be memory, processor cores, storage capacity or networked machines. These additions address different limits and introduce different costs.
Vertical scaling adds capacity to an existing machine. Horizontal scaling spreads work across more machines. More memory can allow data or parameters to remain close to computation; additional machines require an arrangement for distributing those data and coordinating results. Neither approach removes the need to understand the algorithm. A task with tightly connected updates may spend substantial effort communicating even when there is ample arithmetic capacity.
Parallel work and distributed work
Parallel computing carries out multiple calculations at the same time. Distributed computing places cooperating components on networked computers. The concepts overlap: a cluster can execute many operations in parallel while communicating through messages. A multicore machine can also execute parallel work while its processors share memory. The distinction helps explain why a method designed around shared memory cannot automatically be transferred unchanged to a cluster.
Wikipedia’s distributed-computing article identifies coordination, the absence of a global clock and independent component failure as challenges. Learning systems must decide which information moves between workers, when it moves and what happens if one worker stops. Fault tolerance and restart behaviour belong to the learning design because a long computation must still produce a meaningful result when its execution environment changes.
The whole route from observations to evaluation
Evaluation also places a boundary around useful scale. A larger training collection does not itself establish that a model generalises to new observations. Holding data aside for assessment separates the fitting process from the question of predictive behaviour. Repeating that assessment across candidate settings adds computation, so the evaluation plan belongs in the budget from the beginning rather than being treated as a negligible final step.
Training sits inside a larger process. Data preparation selects and transforms observations; the representation determines how features occupy memory; optimisation changes model parameters; evaluation checks behaviour on data outside training. A narrow focus on the training calculation can overlook repeated data movement or the expense of producing features. It can also hide the fact that evaluating several candidate models repeats much of the same work.
The topic pages follow these connections. Stochastic optimisation explains inexpensive updates, while dataflow systems explains how operations follow data dependencies. Parameter servers organise shared model state, and hashing and sketching examines compact representations. Together, these ideas make scale a set of explicit choices about computation, communication and statistical goals.