Big Learning 2013
Big Learning: Advances in Algorithms and Data Management met on Monday 9 December 2013 at Lake Tahoe, Nevada, in Harveys Emerald Bay B. This edition placed the relationship between machine learning and database research at the centre of its programme. It examined how parallel learning algorithms and scalable data systems could inform one another, particularly when a model’s dependencies complicated the division of work.
The workshop combined invited talks, contributed talks, poster sessions and a midday Vowpal Wabbit tutorial. Its schedule, accepted papers and invited talks provide complementary routes through the day. A planned closing invited talk was cancelled.

Machine learning and data management
The 2013 workshop description at NeurIPS presented the meeting as an exchange between two communities. Machine-learning researchers faced the challenge of parallelising models with serial dependencies. Database researchers brought experience with concurrent access and correctness, and with systems designed for analytical work. The workshop asked both communities to consider the other’s problems, connecting the shape of a learning calculation with the principles used to organise and coordinate its data.
This was a change of emphasis within the series. The 2011 edition had ranged across hardware, tools, algorithms and applications. The 2012 edition had grouped its interests into four broad areas. The 2013 programme retained parallel learning while drawing particular attention to database concepts. A model’s mathematical objective and a system’s data-management guarantees could now be examined within one discussion.
Five focal questions
Distributed algorithms for online and batch learning formed one focal area. The distinction between online and batch methods concerns how observations enter the calculation and how updates are organised. Distribution adds questions about the location of observations, communication among workers and the coordination of shared state. The optimisation topic page explains why an update rule and its execution arrangement need to be considered together.
Parallel multicore algorithms formed a second area. Multiple cores in one machine create a different setting from a cluster of machines connected through a network. Memory access and scheduling still matter, but the path by which components exchange information is different. The workshop placed both settings within scope, making the distinction between parallel and distributed execution a useful part of its vocabulary.
Theory was a third focal area. A computation divided across resources still has to address the assumptions of its algorithm. The order of operations, the dependencies among values and the way results are combined can matter to that reasoning. The workshop’s interest in theoretical analysis sat beside its practical interest in systems, encouraging discussion of what a method required as well as how it was implemented.
Implementation studies formed a fourth area, including the challenges and lessons of large-scale inference and learning. Such studies connect a mathematical procedure with memory, data movement, scheduling and fault tolerance. A system can spend time on tasks other than the central calculation, including preparing input or coordinating workers. The parameter-server page develops one model for arranging work around shared parameters.
Database systems for Big Learning formed the fifth area. The workshop asked about implemented models and algorithms, system properties, strengths and limitations. Availability, consistency and scalability were named concerns. These terms focus attention on whether data can be reached, which view of it a computation sees, and how the system responds as its workload grows. They give readers a vocabulary for examining data-management choices alongside learning objectives.
From tables to shared parameters
Wikipedia’s introduction to relational databases describes management systems that organise data in rows and columns. That supplies one background model for the database side of the meeting. Large-scale learning can also involve graphs, streams and parameter vectors, so the useful data representation depends on the work being carried out. The learning-inside-databases page explores analytical computation in relation to relational and other storage models.
A separate 2014 USENIX parameter-server paper described a design in which workers held data and workloads while server nodes maintained shared parameters. Its description included asynchronous communication, consistency choices and fault tolerance. Those are subjects that help explain the systems vocabulary surrounding the 2013 accepted-paper title on parameter servers. The later publication provides its own research account; the workshop list identifies the earlier contribution by title and authors.
A day of talks, posters and a tutorial
The morning began with two invited talks, followed by posters and coffee. Three contributed talks then addressed parameter servers, distributed PCA and clustering, and parallel inference for the Dirichlet process. A midday interval included posters and lunch. In the afternoon, invited subjects concerned timely dataflow and the computational–statistical interface. Further posters and coffee preceded the cancelled closing slot. This sequence placed algorithmic and systems subjects next to opportunities for discussion.
The Vowpal Wabbit tutorial occupied the early afternoon. John Langford’s account of the NIPS tutorials identified its time as 1:30–3:00 pm in Harveys Emerald Bay B. A tutorial is a different kind of programme element from a contributed research talk: its place in the day supports learning about a tool and its vocabulary. The invited-talk page includes its exact title and presenter alongside the other listed subjects.
The 2013 call topics connect these programme choices with the meeting’s stated interests. The author guidelines explain the extended-abstract format and presentation arrangements, while the calendar places preparation before the workshop day. Together these pages show the edition as a meeting around learning algorithms and data management, with distinct roles for calls, accepted contributions, invited subjects and practical instruction.