The Big Learning 2011 invited programme spanned 16–17 December at the Montebajo theater in Sierra Nevada, Spain. Its ten listed talk entries ranged from accelerator hardware and application data to sketches, cluster computation and graph-oriented software. Two tutorials added practical sessions on Vowpal Wabbit and GraphLab 2.0.

The day and session divisions below follow the programme, with the subject labels attached to the invited entries. Context beside each title explains its broader computational subject. The 2011 edition introduction describes the two-day meeting, while the accepted-paper page gives its contributed talks and spotlights.

Rows of dark folding seats in an empty tiered lecture theatre.

Hardware, applications and abstractions

The 2011 workshop description at NeurIPS placed tools, software and systems alongside parallel learning algorithms. That relationship runs through the invited subjects. Hardware determines which operations can run together; an abstraction determines how a program expresses its work; an application determines what the calculation needs to represent. The programme moved among those layers rather than treating large-scale learning as a single model family. The big-learning introduction provides a broader vocabulary for that relationship.

The GPGPU introduction explains graphics processors performing calculations traditionally handled by CPUs. The 2012 USENIX paper on resilient distributed datasets supplies a separate account of fault-tolerant distributed-memory computation. The 2010 GraphLab paper addresses asynchronous iterative algorithms, sparse dependencies and data consistency. Together, these references connect hardware choices with the abstractions used to express parallel learning.

Day 1 - December 16th, 2011

Morning Session

Hardware Accelerated Learning

  1. GPU Metaprogramming: A Case Study in Large-Scale Convolutional Neural NetworksNicolas Pinto
  2. NeuFlow: A Runtime Reconfigurable Dataflow Processor for VisionYann LeCun and Clement Farabet

General-purpose GPU computing uses graphics processors for work beyond graphics. Wikipedia’s GPGPU introduction explains that role in relation to computation traditionally assigned to CPUs. Convolutional neural networks add the subject of learned filters, connecting a model’s operations with the parallel resources that execute them.

Reconfigurable hardware allows the arrangement of computational logic to change after manufacturing. An FPGA expresses that idea through programmable logic blocks and interconnections. This provides background for the title’s pairing of dataflow and vision, where the organisation of operations matters alongside the representation of an image.

Afternoon Session

Applications

  1. Towards Human Behavior Understanding from Pervasive Data: Opportunities and Challenges AheadNuria Oliver

Pervasive data brings questions about the variety and arrival pattern of observations. Large collections can contain different types of information, while a continuous stream raises different organisational questions from a fixed batch. These are general data-management issues around the behavioural subject named in the title.

Tools and Software

  1. Big Machine Learning Made EasyMiguel Araujo and Charles Parker

A learning tool provides an interface between a mathematical task and its execution. Its representation of inputs, models and operations shapes how a calculation can be expressed. This places software design within the workshop’s larger discussion of algorithms, systems and practical learning tasks.

Applications

  1. Machine Learning's Role in the Search for Fundamental ParticlesDaniel Whiteson

Scientific learning problems can involve large collections of observations and complicated models. The 2011 workshop named astronomy among its application domains, and its interests extended to other computational sciences. The title placed fundamental-particle research within that broader discussion of data and learning.

Day 2 - December 17th, 2011

Morning Session

Models and Algorithms

  1. Hazy: Making Data-driven Statistical Applications Easier to build and MaintainChris Re
  2. Real Time Data SketchesAlex Smola

A statistical application needs both a learning calculation and a way to organise its input. Relational data management supplies one familiar vocabulary through rows, columns and queries. The title connected these data-oriented concerns with the construction and maintenance of statistical applications.

A sketch stores a compact summary rather than every observation in its original form. Frequency sketches, for example, estimate counts from a stream using a data structure built around hash functions. This gives background for the title’s combination of sketches and real-time data.

Afternoon Session

Tools and Software

  1. Spark: In-Memory Cluster Computing for Iterative and Interactive ApplicationsMatei Zaharia
  2. Machine Learning and Apache HadoopJeff Hammerbacher

The 2012 USENIX paper on resilient distributed datasets described a distributed-memory abstraction for fault-tolerant computation. Its motivating subjects included iterative algorithms and interactive data exploration, and it identified Spark as an implementation. This supplies a separate research reference for the cluster and in-memory vocabulary in the invited title.

MapReduce divides a calculation into operations over input and operations over grouped results. A distributed implementation also coordinates tasks, data transfers and fault tolerance. Hadoop is one setting for this programming model, connecting the title with the workshop’s systems vocabulary.

Josh Wills gave the talk listed for Jeff Hammerbacher.

Tools and Systems

  1. GraphLab 2: The Challenges of Large Scale Computation on Natural GraphsCarlos Guestrin

The 2010 GraphLab paper described an abstraction for asynchronous iterative algorithms with sparse dependencies. It also addressed the consistency of data used by those computations. This supplies context for graph-oriented learning, where the relationships among values affect which work can proceed together.

Tutorials

The tutorial slots took place in the early afternoon, before the longer afternoon sessions. They gave the programme a practical component alongside invited subjects, contributed research and poster discussion. Their two tool choices also represented different computational perspectives: online learning and graph-oriented iterative work.

  1. Tutorial: Vowpal WabbitJohn Langford
  2. Tutorial: GraphLab 2.0Joseph Gonzalez and Yucheng Low

Vowpal Wabbit is an online and interactive machine-learning system. Online learning concerns updates as observations are processed, rather than treating every task as a single fixed batch calculation. The tutorial title connects this computational pattern with a practical tool, while the optimisation page explains its general vocabulary.

Graph-oriented computation makes relationships among data part of the programming model. Local changes can depend on neighbouring values, so consistency and the scheduling of updates become useful concepts. The graph-parallel page explains those ideas alongside the GraphLab research line and other graph computation models.

Reading across the programme

The invited subjects can be followed through several topic pages. Hardware for learning explains the accelerator terms in the first session. Hashing and sketching develops compact representations, and dataflow systems considers the organisation of distributed work. These connections place the talk titles beside general concepts, while the speaker names identify their roles in the 2011 meeting.