Accepted papers, 2013
The Big Learning 2013 programme accepted 21 papers, divided into three contributed talks and 18 posters. Their titles connect distributed learning with data management, representation, statistical inference and parallel hardware. The contributed subjects were parameter servers, parallel inference for the Dirichlet process, and distributed PCA with k-means clustering.
The complete title and author groups follow below. The 2013 edition page introduces the workshop’s machine-learning and database theme, while the schedule places its contributed talks and poster sessions within the day. The invited-talk page gives the complementary invited subjects.

Shared parameters and distributed computation
The 2013 description at NeurIPS centred the workshop on an exchange between machine-learning and database research. Parameter servers and resource-elastic learning appeared among its accepted titles. For a separate research account of the parameter-server line, the 2014 USENIX paper described workers holding data and workloads while server nodes maintained shared parameters. Its description supplies systems vocabulary without substituting for the earlier workshop contribution.
Other titles concerned distributed communication, iterative-convergent learning and large-scale computation with memory constraints. These subjects sit beside the workshop’s interest in implementation studies and system properties. The parameter-server page develops shared-parameter coordination, while the dataflow page considers the organisation and movement of computational work.
Dimensions, clusters and statistical uncertainty
PCA appeared in both contributed and poster titles, alongside k-means, kernel methods and nearest-centroid ideas. Wikipedia’s PCA explanation describes a linear dimensionality-reduction technique. This gives a reader background for the dimensionality vocabulary, while the clustering titles identify another group of objectives. The list therefore provides several entry points for considering how data representation and learning tasks interact at scale.
Parallel inference for the Dirichlet process formed a contributed subject. Online bootstrapping appeared among the posters. Wikipedia’s bootstrap account explains resampling to estimate an estimator’s distribution, supplying background for uncertainty assessment. These subjects can be read alongside the inference topic page, which introduces probabilistic methods and repeated statistical computation in a general setting.
Features, models and application settings
The poster titles included feature hashing, similarity search, tensor factorisation, kernel learning, random forests and multiclass classification. Several explicitly named GPUs, cloud computation or distributed execution. Their wording brings together a learning objective and a computational setting. It also makes clear that the programme encompassed more than one model family, even within a workshop whose focal theme involved algorithms and data management.
Scientific documents, digital marketing and online learning appeared as further title-level subjects. A reader can follow representation terms across these application settings, or trace parallel-computation vocabulary across different model types. The hashing page and hardware page offer useful companions. The contributed-talk and poster headings below describe the programme’s presentation groups, with each title followed by its printed author order.
Contributed Talks
- Parameter Server for Distributed Machine Learning
- Pitfalls in the use of Parallel Inference for the Dirichlet Process
- Distributed PCA and k-Means Clustering
Posters
- Ensembles of Budgeted Kernel Support Vector Machines for Parallel Large Scale Learning
- Efficient Online Bootstrapping for Large Scale Learning
- Distributed and Scalable PCA in the Cloud
- Towards Distributed Reinforcement Learning for Digital Marketing with Spark
- Lost in Publications? How to Find Your Way in 50 Million Scientific Documents
- cnidaria: A Generative Communication Approach to Scalable, Distributed Learning
- Beyond Pairwise: Provably Fast Algorithms for Approximate k-Way Similarity Search
- Petuum: A System for Iterative-Convergent Distributed ML
- Online Imbalanced Learning with Kernels
- FLEXIFACT: Scalable Flexible Factorization of Coupled Tensors on Hadoop
- A Distributed Approximation Algorithm for Mixed Packing-Covering Linear Programs
- Task-driven Greedy Learning of Feature Hashing Functions
- Approximate Nearest Centroid Embedding for Kernel $k$-Means
- Learning Random Forests on the GPU
- Towards Resource-Elastic Machine Learning
- Building Multiclass Nonlinear Classifiers with GPUs
- BIDMach: Large-scale Learning with Zero Memory Allocation
- Jubatus: An Open Source Platform for Distributed Online Machine Learning