Before and after Big Learning
Big Learning belonged to a wider discussion about the relationship between machine learning and computing systems. Earlier NIPS workshops had addressed parallel implementations, massive datasets, low-rank methods and learning across cores, clusters and clouds. Related meetings elsewhere addressed graph computation and algorithms for large datasets. The titles and years below identify those events without treating their subjects as interchangeable.
The three Big Learning editions then put algorithms, systems, tools and data management into a shared programme. The NeurIPS description of the 2011 workshop asked what changed when learning faced terabytes or petabytes of data. Later ML Systems workshops and the separate MLSys conference continued to address connections between learning and systems. They were distinct events, with their own scopes and programmes.
Related workshops previously at NIPS
The earlier titles show several ways of approaching the scale question. “Parallel Implementations of Learning Algorithms” put implementation in the foreground, while “Large-Scale Machine Learning: Parallelism and Massive Datasets” connected the execution model to the amount of data. The two 2010 titles addressed computing settings and low-rank methods respectively. These are descriptions signalled by the titles, rather than accounts of individual presentations or findings.
The 2007 title, “Statistical Learning Techniques for Solving Systems Problems,” approached the relationship from another direction. It named learning as a means of studying systems problems. Read alongside the later titles, it shows that the intersection could be discussed in both directions: systems for learning and learning applied to systems. Big Learning's later emphasis on data management added another explicit connection within that broad field.
- Big Learning Workshop 2012
- Big Learning Workshop 2011
- Learning on Cores, Clusters, and Clouds 2010
- Low-rank methods for Large-scale Machine Learning 2010
- Large-Scale Machine Learning: Parallelism and Massive Datasets 2009
- Parallel Implementations of Learning Algorithms 2008
- Statistical Learning Techniques for Solving Systems Problems 2007
“Cores, Clusters, and Clouds” also named distinct computing settings. Cores can share one processor package, clusters connect multiple computers, and cloud computing provides access to shared resources over a network. The setting affects how a learning procedure obtains data and coordinates work. That is why titles about computing platforms sit naturally beside titles about mathematical methods: each addresses part of the relationship between a statistical task and its execution.
Related events elsewhere
The related-event list also included Graphlab workshops, a tutorial on scaling up machine learning, and meetings on algorithms for modern massive datasets. A workshop, a tutorial and an algorithms meeting offered different settings for the subject. Their inclusion alongside the NIPS workshops reflected the breadth of the technical discussion, rather than placing every event in one administrative series.
The graph-workshop titles provide a natural connection to graph-parallel computation, while the algorithms titles connect to optimisation and compact data representations. These topic guides explain the general ideas. An event title alone does not establish which algorithms were presented, how they were evaluated or what results were reported.
- 2nd Graphlab Workshop 2013
- Graphlab Workshop 2012
- Scaling Up Machine Learning, the Tutorial 2011
- Workshop on Algorithms for Modern Massive Data Sets 2006, 2008, 2010
Big Learning's three editions
The 2011 edition connected algorithms, systems and tools for learning at scale, with hardware among its focal areas. The 2012 edition grouped its scope around data, models and algorithms, applications, and tools and systems. These broad categories allowed the discussion to include data handling and practical implementation as well as statistical procedures and their analysis.
The 2013 edition foregrounded algorithms and data management. Its NeurIPS workshop description proposed an exchange between database researchers and machine-learning researchers. Online and batch learning, multicore and distributed algorithms, theoretical analysis and implementation studies all formed part of that focus. The programme therefore linked the work of fitting a model with the work of managing the data and computation around it.
Later ML Systems workshops
LearningSys was announced as a workshop name for NIPS 2015. The later ML Systems workshop page for NIPS 2017 described a meeting at the crossroads of machine learning, systems design and software engineering. Its scope included learning platforms, programming languages, data structures, distributed learning and GPU processing. It also asked how research in that area should be evaluated.
Those subjects overlapped with the scale questions discussed at Big Learning, while the later event had its own programme and name. The shared vocabulary helps explain a continuing research concern: a learning method is executed within a system, and that system has a design that can itself be studied. The relationship does not require every venue to have the same organisers, format or emphasis.
A separate conference on machine learning and systems
The MLSys conference describes an annual interdisciplinary conference at the intersection of machine learning and systems. Its scope includes model training and inference, distributed algorithms, data preparation, programming models, compilers, specialised hardware and the testing of learning applications. It is a separate conference, with its own organisation and programme, rather than another Big Learning edition.
The continuity lies in the research questions. Data, algorithms and execution systems remain connected, even when a venue's title or format changes. The dataflow guide, parameter-server guide and database guide examine different parts of that relationship. The workshop programmes give a concrete setting in which those questions were discussed during 2011–2013.