Large-scale machine learning

Big Learning

Large-scale machine learning and the NIPS workshops, 2011–2013

Big learning asks what changes when a machine-learning problem exceeds the practical reach of a single sequential computation. The pressure can come from terabytes or petabytes of observations, from a model with many interacting parts, or from data that keeps arriving while an analysis runs. Algorithms, storage, communication and hardware then become connected design choices. A statistical method has to be understood alongside the system that carries out its work.

The Big Learning workshops at NIPS in 2011, 2012 and 2013 brought those questions into a shared programme. Their subjects ranged from accelerator hardware and probabilistic inference to distributed software and database management. The three editions provide useful starting points for understanding how learning at scale became a meeting place for several research communities.

Rows of dark server cabinets beside a dim aisle with green indicator lights.

Scale changes the whole task

A large dataset does not define the problem by itself. Its shape, arrival pattern and relationship to the model also matter. A collection of independent examples creates different opportunities from a graph whose neighbouring values influence each other. A stream calls for different handling from a fixed batch. A method that revisits the same observations has different storage needs from one that processes them once. These distinctions explain why big learning reaches beyond a simple count of rows.

Wikipedia’s account of big data describes datasets that exceed the reach of traditional processing software, and connects the difficulty with volume, variety and velocity. Learning adds further questions about uncertainty, representation and evaluation. Distributing a calculation creates work around data placement and communication; changing an algorithm creates work around its statistical assumptions. The useful question is how those choices fit together for a particular task.

Three workshops, three points of emphasis

2011: algorithms, systems and tools for learning at scale

The first edition met on 16–17 December 2011 in the Montebajo theater at Sierra Nevada, Spain. It combined invited talks, contributed presentations, spotlights, posters and tutorials. Its NeurIPS workshop description named application domains including bioinformatics, astronomy, recommendation systems, social networks, computer vision, web search and online advertising. Hardware and software appeared alongside statistical methods because the meeting asked how parallel computation changed the learning task. The 2011 edition page introduces the two-day programme and connects its talks and accepted papers.

2012: a shared vocabulary for data and computation

The second edition met on Saturday 8 December 2012 at Lake Tahoe, Nevada, in Harveys Emerald Bay A. The 2012 programme description at NeurIPS organised its interests into big data, models and algorithms, applications, and tools, software and systems. Its talks and accepted papers placed probabilistic inference, graph methods, streaming computation and scientific data within the same discussion. The 2012 edition page explains the four topic areas and the programme’s morning and afternoon sessions.

2013: learning meets data management

The third edition met on Monday 9 December 2013 at Lake Tahoe, in Harveys Emerald Bay B. The 2013 workshop description at NeurIPS made the relationship between machine learning and database research explicit. Distributed and multicore algorithms remained central, while availability, consistency and the design of scalable data systems received particular attention. The 2013 edition page connects that theme with the day’s invited talks, contributed papers, poster sessions and Vowpal Wabbit tutorial.

Routes through the subject

Readers approaching the field for the first time can start with what big learning means and use the glossary alongside it. The topic pages then separate recurring ideas: how optimisation proceeds, how data moves through a system, how graph dependencies affect scheduling, and how shared parameters are coordinated across workers. Hardware, probabilistic inference and database systems offer further views of the same practical problem.

Readers interested in the meetings can begin with an edition, then compare its paper titles with its invited-talk subjects. The schedule gives the order of the 2013 day; the dates and author-guidelines pages explain the workshops’ preparation and presentation formats. Before and after Big Learning places the series among related meetings. The biology page follows one application domain into research on proteins, peptides and genomes, keeping its focus on data, models and computation.

These routes can be read together. A topic page supplies vocabulary for a title in a paper list, while an edition page shows why that topic appeared beside other questions. An algorithmic idea can be considered at several levels: the learning objective, the way observations are represented, the coordination of repeated updates, and the physical resources required to carry them out. That relationship between mathematical choices and practical systems remains the thread connecting the workshop programmes.

Explore Big Learning

The workshops

Learning at scale

Reference