The vocabulary of big learning joins statistical methods to computing systems. The definitions below distinguish algorithms, execution models, data representations and presentation formats, with an internal link from each term to its context.

Wikipedia describes scalability through growing work. The large-scale optimisation survey on arXiv reviews learning algorithms, while the MapReduce paper describes a distributed execution model. These are different layers of the field, alongside the meeting formats in Wikipedia's academic conference overview.

Learning and statistical methods

Asynchronous SGD
Workers make gradient updates without waiting for every other worker to complete the same step, linking execution timing to learning.
Batch learning
A model is fitted using an available training collection considered together, contrasting with updates made as individual observations arrive.
Big learning
Learning algorithms and computing systems are studied together when datasets or model calculations demand large-scale computation.
Bootstrap
Repeated samples drawn from data support estimates of how a statistic varies; sampling with replacement is one common form.
Convolutional neural network
A neural network learns filters to process data patterns, with shared filter weights connecting its representation to input structure.
Cross-validation
Different portions of data are used for fitting and testing a model to examine behaviour outside the fitting portion.
Deep learning
Multiple layers of computation learn representations from data, making a layered model central to this machine-learning approach.
Dirichlet process
A probability distribution over probability distributions expresses prior beliefs about an unknown distribution in a Bayesian model.
Gradient descent
Variables are adjusted repeatedly using derivatives of an objective, with updates directed towards decreasing that objective.
k-means clustering
Observations are assigned to groups by distance from group means, with assignments and means revised through an iterative procedure.
Latent Dirichlet allocation
A probabilistic topic model represents documents as mixtures of hidden topics and topics as distributions over words.
Learning rate
The step size sets the scale of an optimisation update, making it a parameter of the procedure rather than the dataset.
Machine learning
Statistical algorithms learn patterns from observed data and aim to apply those patterns to observations outside the training collection.
Matrix factorisation
Smaller matrices represent a larger matrix; recommendation research uses such factors to describe latent relationships between users and items.
Mini-batch
Several training examples are processed together for one gradient update, between using a single example and the whole collection.
Neural network
Connected computational units have weighted relationships, with training adjusting model parameters to fit the task represented by the data.
Online learning
A model is revised as observations arrive in sequence, making the relationship between data arrival and updates part of learning.
Principal component analysis
A linear transformation identifies directions of variation, with selected directions providing a lower-dimensional representation of the data.
Stochastic gradient descent
Model updates use gradients estimated from sampled training examples, so each update may differ from the full-data gradient.
Variational inference
A tractable distribution approximates a difficult posterior, with optimisation selecting an approximation from an available family of distributions.

Data and distributed execution

AllReduce
A shared operation combines local worker values and distributes the resulting value back to every participating worker.
Big data
The size or complexity of a dataset makes it difficult to handle with traditional processing tools and data workflows.
Bulk-synchronous parallel
Local computation and communication form supersteps, with a barrier coordinating the end of each step across participating workers.
Cloud computing
Network access provides a shared pool of physical or virtual computing resources, including provisioning and resource management.
Count–min sketch
Hash functions place stream events into a compact frequency summary, where collisions can cause estimates to exceed actual frequencies.
Crowdsourcing
Dispersed participants contribute pieces of work to a common task, which in learning research can include producing data labels.
Data parallelism
The same operation processes different portions of data concurrently, dividing work by examples rather than distinct parts of a model.
Dataflow
Operations connect through the data they receive and produce, with those connections expressing the dependencies of the computation.
Distributed computing
Components on networked computers coordinate through messages, making communication and independent failures part of the system design.
Fault tolerance
A system continues functioning when components fail or behave incorrectly, with its design containing the consequences of faults.
Feature hashing
A hash function maps features directly to positions in a numeric representation, where distinct features can share a position.
Graph partition
Vertices are divided into separate groups, with edges between groups linking the resulting parts of the graph.
Graph-parallel computation
Computation follows graph-structured data and relationships, exposing dependencies that a parallel execution system needs to handle.
MapReduce
Map operations produce intermediate key–value pairs, and reduce operations combine values sharing a key in a structured processing model.
Model parallelism
Different parts of a model are placed on different workers, dividing model computation rather than only dividing training examples.
Parameter server
Workers handle local data and computation while server nodes maintain shared model parameters exchanged through communication.
Relational database
Tables represent data through rows and columns, while relational operations and queries provide ways to manipulate that representation.
Resilient distributed dataset
An RDD distributes a memory-based dataset across a cluster, with restricted transformations supporting fault tolerance during computation.
Scalability
A system handles increasing work, potentially through additional resources, so its behaviour is considered in relation to workload growth.
Sketch
A compact summary retains selected information about a larger dataset or stream, trading detail for a smaller representation.

Hardware and systems

FPGA
A field-programmable gate array contains configurable logic blocks and connections that can be programmed after manufacture.
GPU
General-purpose computation uses a graphics processing unit beyond graphics, applying its parallel structure to other computational workloads.
GraphLab
A research programming abstraction supports iterative learning with sparse dependencies, addressing parallel execution and consistency of shared data.
Hadoop
An open-source framework processes large datasets across clusters, with handling of computer failures included at the application layer.
Multicore processor
Multiple processing cores occupy one processor package, allowing parallel software to carry out work concurrently on those cores.
Spark
An engine supports data processing and learning on one machine or a cluster, including research on in-memory distributed datasets.

Workshop presentation formats

Contributed talk
Accepted workshop work is presented in a scheduled speaking slot, a programme category distinct from an invited presentation.
Extended abstract
A compact written contribution provides more detail than a brief summary; Big Learning used this format for workshop submissions.
Poster session
Contributions are discussed around displayed posters during a shared period, allowing smaller conversations to take place alongside one another.
Spotlight
A short presentation draws attention to a poster contribution, as in the spotlight blocks within the 2011 programme.