Big Learning 2012
Big Learning: Algorithms, Systems, and Tools met on Saturday 8 December 2012 at Lake Tahoe, Nevada. Harveys Emerald Bay A held the workshop’s morning and afternoon sessions. The edition continued the series’ interest in learning problems whose data, models or computation demanded parallel and distributed approaches, while giving equal space to practical applications and the software needed to carry out the work.
Its programme brought together three invited talks, four contributed talks and poster sessions. The invited-talk page lists the speakers and titles; the accepted-paper page gives all contributed and poster titles with their author names. Together they show several ways in which large-scale learning connected statistics with computational systems.

Four areas of interest
The 2012 workshop description at NeurIPS grouped its interests into big data, models and algorithms, applications, and tools, software and systems. These areas were related but distinct. Data questions concerned what arrived and how it could be understood. Algorithm questions concerned the learning and inference calculations. Application questions concerned the setting in which those calculations had to operate. System questions concerned the languages, libraries, storage and hardware that made an implementation possible.
Big data as an organisation problem
The data area covered large, unstructured and streaming collections, together with cleaning, visualisation, interactive exploration, sketching and summarisation. This emphasis made the handling of data part of the research question. A model depends on the representation and quality of its inputs, while a computational system depends on their size and arrival pattern. A dataset can be difficult because it is wide, because its observations arrive continuously, or because its relationships resist a simple table.
That area connects with hashing and sketching, where compact representations offer ways to work with very wide data. It also connects with dataflow systems, where the form and movement of the input shape the computation. These are conceptual connections between topic areas, rather than claims about the results of an individual accepted paper.
Models and algorithms under computational constraints
The algorithm area included parallel and distributed methods, online algorithms, accelerator architectures, theoretical analysis and implementation studies. Fault tolerance was among the stated interests. The accepted titles reflected this breadth through Markov chain Monte Carlo, graph partitioning, conditional-gradient methods, Gaussian processes, structured prediction and topic models. Their shared setting was scale, while their mathematical objects and computational requirements differed. The full lists retain the distinction between contributed talks and posters.
For probabilistic vocabulary, Wikipedia’s introduction to the Dirichlet process explains a probability distribution over probability distributions, used in Bayesian inference. That background helps orient readers to a contributed title involving Dirichlet-process mixtures. It does not establish how the workshop paper’s algorithm behaved. The scalable-inference page places this term beside variational methods, topic models and Monte Carlo computation.
Applications and the people using systems
The application area asked for practical studies and the challenges of building real systems. It included end-user needs, stream and batch characteristics, and alternatives for obtaining labels. These subjects connect an abstract learning objective with the conditions around a task. A collection assembled in advance has different operational demands from a continuous stream; carefully curated labels raise different questions from labels obtained through many contributors. The workshop treated these differences as part of the learning problem.
Scientific data also appeared among the accepted titles. Molecular biophysics and genomic variant calling stood beside web-oriented query understanding, topic modelling and computer vision. Those title-level connections illustrate the variety of domains in the programme. They do not justify claims about the papers’ outcomes or about uses of the methods outside the meeting. The biology topic page develops the broader relationship between molecular data and machine learning.
Tools, storage and programming models
The tools area concerned languages and libraries for parallel or distributed learning, with cloud computation, scalable storage and specialised hardware among its interests. The programme’s titles included online expectation maximisation with Spark Streaming, Hadoop-based graph construction and several learning frameworks. A programming model gives a method a way to express its work; a storage model determines how the needed data can be accessed. These choices can shape an implementation even when the statistical objective remains the same.
Wikipedia’s stream-processing account describes computation organised around sequences of events in time. That perspective supplies background for the programme’s streaming and temporal subjects. It highlights the relationship between an arriving input and the operations applied to it. The database-learning page offers a complementary view through data management and the placement of analytical work.
The workshop day
The morning began with poster setup and opening remarks, followed by the invited topic-model talk. A contributed Monte Carlo talk, a poster session, graph partitioning and conditional-gradient presentations followed. The midday interval included lunch and posters. The afternoon began with temporal analytics, continued with the molecular-biophysics contributed talk and posters, and ended with an invited subject on randomized sampling in exploration seismology. Closing remarks completed the programme.
The day’s mixture can be read alongside the 2011 edition and the 2013 edition. The earlier meeting had a two-day programme with spotlights and tutorials. The later meeting concentrated on the interface between machine learning and database management. The key dates and organisers page provide the edition-level calendar and roles, while the talk and paper pages retain the programme’s named contributions.