The Big Learning 2012 invited programme brought three subjects to the Lake Tahoe meeting on 8 December: scalable probabilistic modelling, randomized sampling in exploration seismology, and temporal analytics. The talks took place in Harveys Emerald Bay A, within the workshop’s morning and afternoon sessions.

The titles and speaker names below identify the invited entries. Their accompanying context develops the general statistical and computational vocabulary involved. The 2012 edition introduction explains the four topic areas of the meeting, and the accepted-paper page supplies the contributed-talk and poster titles that appeared beside these subjects.

Rows of dark folding seats in an empty tiered lecture theatre.

Probability, time and the shape of data

The 2012 workshop description at NeurIPS connected algorithms and systems with practical learning domains, including the sciences. Its invited subjects approached that relationship from different directions. Probabilistic models concern uncertainty and the structure represented by data. Temporal analytics concerns the order and timing of observations. Randomized sampling concerns the selection of observations for a calculation. Their combination made the programme a useful meeting point between statistical choices and computational organisation.

The separate paper Online Learning for Latent Dirichlet Allocation connected online variational inference with stochastic optimisation. The big-data account explains how volume, variety and velocity can complicate sampling. The stream-processing explanation places sequences of events in time at the centre of computation. These references develop the model, sampling and temporal vocabulary that connects the three invited subjects.

The distinctions are worth carrying through the programme. A large document collection can be represented through latent topics. An event stream can be organised by time. A scientific dataset can be approached through the relationship between the available observations and the calculation performed on them. These are different ways to define a task before deciding how parallel resources should execute it. The big-learning introduction places this connection between representation and computation in a broader setting.

Morning Session

  1. Big Data and Bayes: Stochastic Variational Inference and Scalable Topic ModelsDavid Blei

Latent Dirichlet allocation describes documents through mixtures of unobserved topics, with topics represented by distributions over words. The NIPS paper Online Learning for Latent Dirichlet Allocation described an online variational-Bayes algorithm using stochastic optimisation. This supplies a separate reference for the online and variational vocabulary in the title. It connects the subject of probabilistic modelling with computation that processes observations through repeated updates.

Afternoon Session

  1. Randomized sampling in exploration seismologyFelix J. Herrmann
  2. Temporal Analytics on Big DataJonathan Goldstein

The title placed randomized sampling within a scientific application domain. Wikipedia’s big-data account identifies analysis, storage, transfer and visualisation among the challenges of large datasets, and notes that scale can complicate sampling. These general concerns supply context for thinking about the relationship between an observation collection and an analytical task. The scientific setting named in the title adds a different application perspective to the programme’s probabilistic and temporal subjects.

Wikipedia’s stream-processing explanation identifies sequences of events in time as central inputs and outputs of computation. That gives a background vocabulary for temporal analytics, where the ordering of observations is part of how data is understood. Stream processing connects these event sequences with programming models, distribution and scheduling. The title therefore sits naturally beside the workshop’s interest in both data characteristics and the systems used to analyse them.

Inference as a computational task

A probabilistic model expresses relationships among quantities that are not all directly observed. Inference connects those relationships with data. Topic modelling supplies one example: documents and words are observed, while topic mixtures provide a statistical representation. Computing that representation requires a method as well as a model. The scalable-inference page distinguishes topic models, variational methods and Monte Carlo approaches.

Online and batch methods organise that work differently. A batch procedure uses a collection as a fixed input to a calculation. An online procedure updates its state while processing observations. The 2010 online-LDA publication is useful background because it explicitly connected variational inference with stochastic optimisation. The optimisation page explains the update-oriented vocabulary without treating every inference task as the same algorithm.

Temporal data and scientific collections

Time can be part of the data’s meaning as well as its arrival pattern. A fixed collection may contain timestamps, while a stream presents observations in an ongoing sequence. These are related but distinct properties. Temporal analysis therefore reaches into representation and interpretation, while stream processing reaches into the arrangement of computational operations. The dataflow-systems page develops the connection between input sequences, operations and distributed execution.

Scientific collections extend the same questions into another setting. Data volume can affect the ability to store, transfer and examine observations, while the structure of the analytical task affects what must be retained or represented. The workshop’s accepted titles included molecular biophysics and genomic data alongside more general algorithmic subjects. The biology page follows that application connection through molecular data, learning models and public research infrastructure.

From the invited subjects to the wider day

The morning included the topic-model invited talk, contributed talks on Monte Carlo methods, graph partitioning and conditional-gradient algorithms, and a poster session. The afternoon paired temporal analytics with a molecular-biophysics contribution, further posters and the randomized-sampling invited subject. This arrangement put general methods beside application data and systems questions. The key dates and edition introduction provide the calendar around the programme.

The 2011 invited talks offer a broader two-day range of hardware, applications and tools. The 2013 invited talks then draw attention to data management, dataflow and the computational–statistical interface. Reading the titles together shows several ways to frame learning at scale: by the statistical model, by the time structure of the input, by the application, or by the system that organises the work.