The invited programme at Big Learning 2013 connected machine learning with the systems that held, transformed and coordinated its data. Morning and afternoon talks surrounded a midday tutorial. Four invited talks took place; the closing invited talk was cancelled. The workshop met on 9 December at Lake Tahoe under the title Big Learning: Advances in Algorithms and Data Management.

The NeurIPS description of the 2013 workshop framed the meeting as an exchange between database research and machine learning. The titles below connect data work, execution models and the relation between computation and statistical reasoning. The day's programme also includes the contributed talks and poster sessions that accompanied the invited presentations.

Two unmarked coffee cups on a wooden table beside a window overlooking snowy pine trees.

Morning Session

  1. The Thorn in the Side of Big Data: Too Few Artists

    Chris Ré
  2. Scale-out Beyond Map-Reduce

    Raghu Ramakrishnan

Large-scale learning includes the preparation and organisation of data as well as the fitting of a model. The workshop's stated data-management focus placed implementation studies, system properties and practical limitations alongside algorithms. That framing gives a useful setting for a title concerned with people doing data work: the learning task depends on how data is represented, handled and made available to computation. It also leaves room for discussion of engineering choices that a model description alone does not settle.

The MapReduce paper published through Google Research describes a model in which a map function produces intermediate key–value pairs and a reduce function combines values sharing a key. Its runtime takes responsibility for such tasks as partitioning inputs, scheduling work and communication between machines. A title about computation beyond that model raises the broader question of which execution structure matches a learning workload. Iterative learning, data dependencies and repeated use of intermediate results are relevant subjects in the dataflow systems guide.

Afternoon Session

  1. Timely dataflow in Naiad

    Derek Murray
  2. On the Computational and Statistical Interface and Big Data

    Michael Jordan

Dataflow expresses computation through operations connected by the data passed between them. The dependencies help explain which operations can proceed and which must wait for inputs. This way of describing a programme relates naturally to learning pipelines that transform data before, during and after model fitting. The general idea is broader than a particular system, and the workshop's interest in data management made such programming structures part of the conversation about scale.

The optimisation survey on arXiv reviews numerical optimisation for machine learning and identifies stochastic gradient methods as central to the large-scale setting. Statistical considerations determine what is estimated, while computational considerations determine how an estimation procedure is executed with finite time and memory. The survey's discussion of noisy directions and second-order approximations connects statistical variation to the choice of numerical technique. The optimisation guide follows those relationships through online updates and parallel training.

Midday tutorial

  1. Tutorial on Vowpal Wabbit

    John Langford

The Vowpal Wabbit project describes online learning techniques for supervised, reinforcement and interactive learning. Online learning processes examples through a sequence of updates, making the order and flow of data part of the implementation question. A workshop tutorial on that subject connected algorithmic ideas to a concrete learning library and its working abstractions. Feature representation and compact summaries provide another connection, developed in the hashing and sketching guide.

Afternoon Session

  1. Big Eympathy: Growing Up

    Joe Hellerstein

Joe Hellerstein’s keynote listed above was canceled.

The title belongs with the planned programme, with that status stated explicitly. The wider workshop theme nevertheless supplies a setting for discussion of large-scale learning as joint work across research communities: its goals joined parallel algorithms to database implementations and system-design considerations. Those goals concerned the relationship among different kinds of expertise, rather than making a claim about an individual talk's content.

Reading across the subjects

The two morning presentations preceded the contributed-talk block; the two completed afternoon presentations followed the tutorial interval. This organisation put shared framing sessions around more specific research contributions and practical instruction. Poster periods added another way to take part, with discussion centred on individual papers rather than a single presentation to the whole room. The pattern connected the technical programme to the workshop's aim of exchange across communities.

The titles range from the practical work of handling data to the relation between statistical reasoning and computation. That range reflects an important feature of large-scale learning: no single layer explains the whole process. A learning method, a data representation and an execution system each impose requirements on the others. The workshop placed them in a shared programme so that algorithm and database questions could be discussed together.

The database guide examines this relationship from the storage and query side. The parameter-server guide takes up the coordination of model state across workers, while the introduction to big learning explains why scale affects both data and computation. Together, these topics show how data layout, model updates and execution order belong to the same scale question.

The 2011 invited programme and the 2012 invited programme show other combinations of hardware, systems and statistical questions. Together, the three editions make clear how broad the phrase “big learning” was: it covered computationally demanding models, large collections of examples and the tools that linked learning procedures to parallel or distributed machines.