Glossary
The vocabulary of big learning joins statistical methods to computing systems. The definitions below distinguish algorithms, execution models, data representations and presentation formats, with an internal link from each term to its context.
Wikipedia describes scalability through growing work. The large-scale optimisation survey on arXiv reviews learning algorithms, while the MapReduce paper describes a distributed execution model. These are different layers of the field, alongside the meeting formats in Wikipedia's academic conference overview.
Learning and statistical methods
- Asynchronous SGD
- Workers make gradient updates without waiting for every other worker to complete the same step, linking execution timing to learning.
- Batch learning
- A model is fitted using an available training collection considered together, contrasting with updates made as individual observations arrive.
- Big learning
- Learning algorithms and computing systems are studied together when datasets or model calculations demand large-scale computation.
- Bootstrap
- Repeated samples drawn from data support estimates of how a statistic varies; sampling with replacement is one common form.
- Convolutional neural network
- A neural network learns filters to process data patterns, with shared filter weights connecting its representation to input structure.
- Cross-validation
- Different portions of data are used for fitting and testing a model to examine behaviour outside the fitting portion.
- Deep learning
- Multiple layers of computation learn representations from data, making a layered model central to this machine-learning approach.
- Dirichlet process
- A probability distribution over probability distributions expresses prior beliefs about an unknown distribution in a Bayesian model.
- Gradient descent
- Variables are adjusted repeatedly using derivatives of an objective, with updates directed towards decreasing that objective.
- k-means clustering
- Observations are assigned to groups by distance from group means, with assignments and means revised through an iterative procedure.
- Latent Dirichlet allocation
- A probabilistic topic model represents documents as mixtures of hidden topics and topics as distributions over words.
- Learning rate
- The step size sets the scale of an optimisation update, making it a parameter of the procedure rather than the dataset.
- Machine learning
- Statistical algorithms learn patterns from observed data and aim to apply those patterns to observations outside the training collection.
- Matrix factorisation
- Smaller matrices represent a larger matrix; recommendation research uses such factors to describe latent relationships between users and items.
- Mini-batch
- Several training examples are processed together for one gradient update, between using a single example and the whole collection.
- Neural network
- Connected computational units have weighted relationships, with training adjusting model parameters to fit the task represented by the data.
- Online learning
- A model is revised as observations arrive in sequence, making the relationship between data arrival and updates part of learning.
- Principal component analysis
- A linear transformation identifies directions of variation, with selected directions providing a lower-dimensional representation of the data.
- Stochastic gradient descent
- Model updates use gradients estimated from sampled training examples, so each update may differ from the full-data gradient.
- Variational inference
- A tractable distribution approximates a difficult posterior, with optimisation selecting an approximation from an available family of distributions.
Data and distributed execution
- AllReduce
- A shared operation combines local worker values and distributes the resulting value back to every participating worker.
- Big data
- The size or complexity of a dataset makes it difficult to handle with traditional processing tools and data workflows.
- Bulk-synchronous parallel
- Local computation and communication form supersteps, with a barrier coordinating the end of each step across participating workers.
- Cloud computing
- Network access provides a shared pool of physical or virtual computing resources, including provisioning and resource management.
- Count–min sketch
- Hash functions place stream events into a compact frequency summary, where collisions can cause estimates to exceed actual frequencies.
- Crowdsourcing
- Dispersed participants contribute pieces of work to a common task, which in learning research can include producing data labels.
- Data parallelism
- The same operation processes different portions of data concurrently, dividing work by examples rather than distinct parts of a model.
- Dataflow
- Operations connect through the data they receive and produce, with those connections expressing the dependencies of the computation.
- Distributed computing
- Components on networked computers coordinate through messages, making communication and independent failures part of the system design.
- Fault tolerance
- A system continues functioning when components fail or behave incorrectly, with its design containing the consequences of faults.
- Feature hashing
- A hash function maps features directly to positions in a numeric representation, where distinct features can share a position.
- Graph partition
- Vertices are divided into separate groups, with edges between groups linking the resulting parts of the graph.
- Graph-parallel computation
- Computation follows graph-structured data and relationships, exposing dependencies that a parallel execution system needs to handle.
- MapReduce
- Map operations produce intermediate key–value pairs, and reduce operations combine values sharing a key in a structured processing model.
- Model parallelism
- Different parts of a model are placed on different workers, dividing model computation rather than only dividing training examples.
- Parameter server
- Workers handle local data and computation while server nodes maintain shared model parameters exchanged through communication.
- Relational database
- Tables represent data through rows and columns, while relational operations and queries provide ways to manipulate that representation.
- Resilient distributed dataset
- An RDD distributes a memory-based dataset across a cluster, with restricted transformations supporting fault tolerance during computation.
- Scalability
- A system handles increasing work, potentially through additional resources, so its behaviour is considered in relation to workload growth.
- Sketch
- A compact summary retains selected information about a larger dataset or stream, trading detail for a smaller representation.
Hardware and systems
- FPGA
- A field-programmable gate array contains configurable logic blocks and connections that can be programmed after manufacture.
- GPU
- General-purpose computation uses a graphics processing unit beyond graphics, applying its parallel structure to other computational workloads.
- GraphLab
- A research programming abstraction supports iterative learning with sparse dependencies, addressing parallel execution and consistency of shared data.
- Hadoop
- An open-source framework processes large datasets across clusters, with handling of computer failures included at the application layer.
- Multicore processor
- Multiple processing cores occupy one processor package, allowing parallel software to carry out work concurrently on those cores.
- Spark
- An engine supports data processing and learning on one machine or a cluster, including research on in-memory distributed datasets.
Workshop presentation formats
- Contributed talk
- Accepted workshop work is presented in a scheduled speaking slot, a programme category distinct from an invited presentation.
- Extended abstract
- A compact written contribution provides more detail than a brief summary; Big Learning used this format for workshop submissions.
- Poster session
- Contributions are discussed around displayed posters during a shared period, allowing smaller conversations to take place alongside one another.
- Spotlight
- A short presentation draws attention to a poster contribution, as in the spotlight blocks within the 2011 programme.