The Big Learning 2012 programme accepted 26 papers: four contributed talks and 22 posters. Their titles covered probabilistic computation, graph methods, optimisation, online learning, data representation and scientific applications. This range fitted the workshop’s broad interest in algorithms, systems, applications and tools for learning at scale.

The two groups below retain the contributed-talk and poster divisions. The 2012 edition page explains the Lake Tahoe meeting and its four topic areas. The invited-talk page gives the complementary subjects on scalable topic models, temporal analytics and randomized sampling.

A stack of plain paper and a closed unmarked laptop on a wooden table.

Inference and statistical models

The 2012 workshop description at NeurIPS set out a meeting across algorithms, systems and real-world domains. Several accepted titles named Monte Carlo methods, Gaussian processes, topic models and Bayesian networks. Those terms place probabilistic modelling and inference within the programme’s larger systems discussion. Wikipedia’s Dirichlet-process introduction supplies background for one contributed title by explaining a distribution over probability distributions used in Bayesian inference.

Topic-model titles appeared in both hierarchical and online forms. A related NIPS publication, Online Learning for Latent Dirichlet Allocation, described an online variational-Bayes algorithm based on stochastic optimisation. That publication offers a separate conceptual reference for readers approaching online inference. The scalable-inference page connects topic models with variational methods and Monte Carlo vocabulary.

Graphs, streams and wide representations

The list included streaming graph partitioning, community detection, graph construction, locality-sensitive hashing and subspace methods. These title-level clusters draw attention to representation as well as learning. A graph title concerns relationships among objects, while a hashing title concerns how objects or features are represented for computation. Subspace and point-to-subspace titles introduce another way to think about the arrangement of data.

Streaming subjects also appeared in online expectation maximisation and distributed online learning. Wikipedia’s stream-processing account describes sequences of events in time as central objects of computation. This background helps distinguish a stream-oriented setting from a fixed input collection. The dataflow page provides a route from that vocabulary to programming models and the movement of data between operations.

Applications alongside implementation

Accepted titles named molecular biophysics, genomic variant calling, computer vision, query understanding and natural-language processing. The scientific titles connect this programme with the wider question of learning from large research datasets. The biology page takes up proteins, peptides and genomes as a field of machine-learning research. Here those connections arise from the printed subjects rather than from the contents of the workshop papers.

Memory, GPU computation, distributed optimisation and learning frameworks formed another cluster of titles. Together with the statistical and application subjects, they show the breadth of the four-area call. A reader can move between the groups to trace graph, probabilistic or online vocabulary across the programme. The divisions below identify the presentation formats used at the meeting; the titles and author order identify each accepted contribution.

Contributed Talks

  1. Parallel Markov Chain Monte Carlo for Dirichlet Process MixturesDan Lovell, Ryan P. Adams and Vikash K. Mansinghka
  2. Streaming Balanced Graph Partitioning Algorithms for Random GraphsIsabelle Stanton
  3. Conditional gradient algorithms for large-scale learningZaid Harchaoui, Anatoli Juditsky and Arkadi Nemirovski
  4. Statistical Inference for Big Data Problems in Molecular BiophysicsArvind Ramanathan, Andrej Savol, Virginia Burger, Shannon Quinn, Pratul Agarwal and Chakra Chennubhotla

Posters

  1. BOOT-TS: A Scalable Bootstrap for Massive Time-Series DataNikolay Laptev, Carlo Zaniolo and Tsai-Ching Lu
  2. A Novel Parallel Hierarchical Community Detection Method for Large NetworksPing Lu, Shengmei Luo, Lei Hu, Yunlong Lin, Junyang Zou, Qiwei Zhong, Kuangyan Zhu, Jian Lu and Qiao Wang
  3. Robust and Efficient Locality Sensitive Hashing for Nearest   Neighbor Search in Large Data SetsByungkon Kang and Kyomin Jung
  4. Parallel one-versus-rest SVM training on the GPUSander Dieleman, A√§ron van Den Oord and Benjamin Schrauwen
  5. Large-Scale Online Expectation Maximization with Spark StreamingTimothy Hunter, Matei Zaharia, Tathagata Das, Pieter Abbeel and Alexandre Bayen
  6. Scaling Bayesian Network Parameter Learning with Expectation Maximization using MapReduceErik Reed and Ole Mengshoel
  7. Distributed Structured Prediction for Big DataAlexander Schwing, Tamir Hazan, Marc Pollefeys and Raquel Urtasun
  8. Local Logistic Classifiers for Large Scale LearningMohammad Reza Yousefi and Thomas M. Breuel
  9. Efficient Point-to-Subspace Query in L1: Theory and Applications in Computer VisionJu Sun, Yuqian Zhang and John Wright
  10. Gaussian Processes for Big Data through Stochastic Variational InferenceJames Hensman and Neil Lawrence
  11. Divide-and-Conquer Subspace SegmentationAmeet Talwalkar, Lester Mackey, Yadong Mu, Shih-Fu Chang and Michael Jordan
  12. MLbase: A Distributed Machine Learning WrapperAmeet Talwalkar, Tim Kraska, Rean Griffith, John Duchi, Joseph Gonzalez, Denny Britz, Xinghao Pan, Virginia Smith, Evan Sparks, Andre Wibisono, Michael Franklin and Michael Jordan
  13. GraphBuilder – Large-Scale Graph Construction using Apache HadoopNilesh Jain, Theodore L. Willke and Haijie Gu
  14. Generalizing Elliptical Slice Sampling for Parallel MCMCRobert Nishihara, Iain Murray and Ryan Adams
  15. Large-Scale Hierarchical Topic ModelsJay Pujara and Peter Skomoroch
  16. FLAG: Fast Large-Scale Graph Construction for NLPAmit Goyal and Hal Daume Iii
  17. Big Learning with Little RAMD. Sculley, Daniel Golovin and Michael Young
  18. Distributed Pipeline for Genomic Variant CallingRichard Xia, Sara Sheehan, Yuchen Zhang, Ameet Talwalkar, Matei Zaharia, Jonathan Terhorst, Michael Jordan, Yun Song, Armando Fox and David Patterson
  19. Large-Scale Query UnderstandingKhaled Refaat, Sugato Basu, Deirdre O’brien and Liadan O’callaghan
  20. Distributed Online Learning for Latent Dirichlet AllocationJinyeong Bak, Dongwoo Kim and Alice Oh
  21. MapReduce for Bayesian Networks Parameter Learning using the EM AlgorithmAniruddha Basak, Irina Brinster and Ole Mengshoel
  22. Large-scale Distributed Optimization for Improving Accuracy at the TopStephen Boyd, Corinna Cortes, Chong Jiang, Mehryar Mohri, Ana Radovanovic and Joelle Skaf