Accepted papers, 2012
The Big Learning 2012 programme accepted 26 papers: four contributed talks and 22 posters. Their titles covered probabilistic computation, graph methods, optimisation, online learning, data representation and scientific applications. This range fitted the workshop’s broad interest in algorithms, systems, applications and tools for learning at scale.
The two groups below retain the contributed-talk and poster divisions. The 2012 edition page explains the Lake Tahoe meeting and its four topic areas. The invited-talk page gives the complementary subjects on scalable topic models, temporal analytics and randomized sampling.

Inference and statistical models
The 2012 workshop description at NeurIPS set out a meeting across algorithms, systems and real-world domains. Several accepted titles named Monte Carlo methods, Gaussian processes, topic models and Bayesian networks. Those terms place probabilistic modelling and inference within the programme’s larger systems discussion. Wikipedia’s Dirichlet-process introduction supplies background for one contributed title by explaining a distribution over probability distributions used in Bayesian inference.
Topic-model titles appeared in both hierarchical and online forms. A related NIPS publication, Online Learning for Latent Dirichlet Allocation, described an online variational-Bayes algorithm based on stochastic optimisation. That publication offers a separate conceptual reference for readers approaching online inference. The scalable-inference page connects topic models with variational methods and Monte Carlo vocabulary.
Graphs, streams and wide representations
The list included streaming graph partitioning, community detection, graph construction, locality-sensitive hashing and subspace methods. These title-level clusters draw attention to representation as well as learning. A graph title concerns relationships among objects, while a hashing title concerns how objects or features are represented for computation. Subspace and point-to-subspace titles introduce another way to think about the arrangement of data.
Streaming subjects also appeared in online expectation maximisation and distributed online learning. Wikipedia’s stream-processing account describes sequences of events in time as central objects of computation. This background helps distinguish a stream-oriented setting from a fixed input collection. The dataflow page provides a route from that vocabulary to programming models and the movement of data between operations.
Applications alongside implementation
Accepted titles named molecular biophysics, genomic variant calling, computer vision, query understanding and natural-language processing. The scientific titles connect this programme with the wider question of learning from large research datasets. The biology page takes up proteins, peptides and genomes as a field of machine-learning research. Here those connections arise from the printed subjects rather than from the contents of the workshop papers.
Memory, GPU computation, distributed optimisation and learning frameworks formed another cluster of titles. Together with the statistical and application subjects, they show the breadth of the four-area call. A reader can move between the groups to trace graph, probabilistic or online vocabulary across the programme. The divisions below identify the presentation formats used at the meeting; the titles and author order identify each accepted contribution.
Contributed Talks
- Parallel Markov Chain Monte Carlo for Dirichlet Process Mixtures
- Streaming Balanced Graph Partitioning Algorithms for Random Graphs
- Conditional gradient algorithms for large-scale learning
- Statistical Inference for Big Data Problems in Molecular Biophysics
Posters
- BOOT-TS: A Scalable Bootstrap for Massive Time-Series Data
- A Novel Parallel Hierarchical Community Detection Method for Large Networks
- Robust and Efficient Locality Sensitive Hashing for Nearest Neighbor Search in Large Data Sets
- Parallel one-versus-rest SVM training on the GPU
- Large-Scale Online Expectation Maximization with Spark Streaming
- Scaling Bayesian Network Parameter Learning with Expectation Maximization using MapReduce
- Distributed Structured Prediction for Big Data
- Local Logistic Classifiers for Large Scale Learning
- Efficient Point-to-Subspace Query in L1: Theory and Applications in Computer Vision
- Gaussian Processes for Big Data through Stochastic Variational Inference
- Divide-and-Conquer Subspace Segmentation
- MLbase: A Distributed Machine Learning Wrapper
- GraphBuilder – Large-Scale Graph Construction using Apache Hadoop
- Generalizing Elliptical Slice Sampling for Parallel MCMC
- Large-Scale Hierarchical Topic Models
- FLAG: Fast Large-Scale Graph Construction for NLP
- Big Learning with Little RAM
- Distributed Pipeline for Genomic Variant Calling
- Large-Scale Query Understanding
- Distributed Online Learning for Latent Dirichlet Allocation
- MapReduce for Bayesian Networks Parameter Learning using the EM Algorithm
- Large-scale Distributed Optimization for Improving Accuracy at the Top