The 2013 Big Learning call asked how parallel learning algorithms and data-management systems could be studied together. It concerned both online and batch learning, work shared across multiple machines, and work divided among cores within a machine. Theory and implementation studies appeared alongside database systems, with attention to the properties and limitations of the systems supporting a learning task.

The workshop took place on 9 December 2013 at Lake Tahoe. No call is open. The NeurIPS description of the 2013 workshop explains its focus on bringing machine-learning and database research into conversation. The accepted-papers list shows the titles and authors selected for the day, and the author-guidelines page explains the format used for that process.

A stack of plain paper and a closed unmarked laptop on a wooden table.

Distributed online and batch learning

The call included distributed algorithms for both online and batch learning. This distinction concerned how a learning procedure encountered its examples. An online procedure worked through a sequence of observations and updates; a batch procedure treated an available collection as the basis for fitting or revising a model. Dividing either form across machines introduced questions about which data and model values each worker needed, and when information passed between workers.

The emphasis was therefore broader than a request for a large dataset. A distributed implementation had to connect the statistical procedure to communication and coordination. The optimisation guide discusses incremental updates, while the parameter-server guide considers shared model state. These subjects explain why “online” and “distributed” described different aspects of the same potential learning system.

Parallel work within a machine

Parallel multicore algorithms were another focal point. The call distinguished them from algorithms distributed across machines, keeping the organisation of computation visible. Multiple cores could work concurrently within one computer, but a learning procedure still had to define which operations were independent and how updates interacted. A method with sequential dependencies did not become a parallel method simply because additional processors were available.

The hardware guide extends this distinction to GPUs, FPGAs and multicore processors. Hardware and algorithm questions were linked throughout the workshop series, although the 2013 title gave data management particular prominence. The call's scope allowed the implementation model to be discussed as part of a learning method rather than treated as incidental machinery outside the research problem.

Theory and implementation studies

Theoretical analysis of parallel and distributed algorithms was explicitly included. Such analysis concerns the assumptions under which an algorithm behaves as claimed, including dependencies between computations or the information available during an update. In the call, it sat next to implementation studies, so formal reasoning and practical experience were both relevant ways to examine a system. Neither category replaced the other.

Implementation studies were asked to address the challenges and lessons of large-scale distributed inference and learning. That wording admitted work about difficulties as well as proposed methods. Data layout, coordination, failure handling and programming choices could affect whether an algorithm's assumptions matched its execution. The graph-parallel guide and the inference guide develop two areas where dependency structure is central.

Database systems as part of learning

The call's systems category included learning models and algorithms implemented in database management systems. It also asked about availability, consistency and scalability, together with strengths and limitations. These properties describe how a data system behaves, not just which model it can run. The AMPLab announcement of the 2013 call set out that systems category alongside data, algorithms and applications.

The MADlib paper on arXiv supplies related background: it described an open-source library of analytic methods that executed inside a database engine. This is one example of the relationship the call asked researchers to consider, rather than a claim about any accepted paper. A database could be part of the execution of learning itself, with the handling of stored data and the statistical computation considered together.

How the earlier calls differed

The 2011 call gave hardware acceleration its own category. It also included application studies, tools and systems, and models and algorithms. The Computing Community Consortium's announcement of the 2011 workshop described its interest in practical studies, demonstrations, benchmarks and lessons from implementation. The two-day 2011 edition accordingly placed hardware and systems topics beside applications and statistical methods.

The 2012 edition grouped its interests as Big Data, Models & Algorithms, Applications of Big Learning, and Tools, Software & Systems. Data cleaning, streams, summaries and interpretation sat beside parallel methods and scalable storage. Compared with those broad categories, the 2013 framing made the exchange between database and machine-learning research especially explicit. The editions nevertheless shared a concern with the relationship among algorithms, data characteristics and actual system building.

Read together, the calls treated scale as a collection of research questions. Large data volumes mattered, but so did structured models, the availability of labels, the flow of observations and the abstractions used to execute learning. The big-learning introduction follows those connections, while the calendar gives the past milestones associated with the three workshop calls.