Probabilistic inference at scale
Probabilistic inference asks what a model can say about unknown quantities after observations become available. Those quantities may include parameters, hidden topic assignments or distributions over possible explanations. Scale complicates the calculation because a large collection of observations can interact with a large set of unknown variables.
Several methods make the problem tractable through approximation or repeated computation. Variational inference optimises an approximate distribution; Markov chain Monte Carlo uses samples; the bootstrap uses resampled data to study an estimator. Each method carries a different kind of computational work and a different question about the quality of its result.

Topics connect local documents and shared structure
A topic model describes patterns in a document collection through unobserved topics. Latent Dirichlet allocation, or LDA, represents a document as a mixture of topics and represents each topic through a probability distribution over words. The observed text provides evidence about these hidden quantities. The task is therefore larger than assigning a single label to each document.
This structure connects local and shared information. A document has its own topic mixture, while topic-word patterns are relevant across the collection. Adding documents increases the amount of local information to process and supplies more evidence about shared patterns. The representation also depends on the vocabulary and the number of topics being considered, so document count alone does not describe the computational problem.
Variational Bayes turns inference into optimisation
Wikipedia’s variational-Bayesian-methods article describes techniques for approximating difficult integrals in Bayesian inference. Rather than calculate a complex posterior distribution directly, a variational method chooses an approximate distribution from a specified family and adjusts it through optimisation. That choice of family limits what relationships the approximation can express.
The approximation and the optimisation must be assessed separately. A procedure can solve its chosen optimisation problem thoroughly while the selected distribution family remains restrictive. Conversely, a useful family may still require substantial work to fit. At scale, these statistical choices meet systems questions about repeated data access, intermediate state and which quantities can be updated locally.
Online LDA processes portions of a collection
The NIPS 2010 paper Online Learning for Latent Dirichlet Allocation developed an online variational-Bayes algorithm based on stochastic optimisation with a natural-gradient step. Its description included document collections arriving in a stream. This connected an inference objective to updates that could proceed through portions of the data.
Online inference distinguishes the work needed for the current portion from changes to quantities shared across the collection. The update schedule determines how new information influences existing estimates. A stream also leaves the final amount of data unspecified at the start. The stochastic-optimisation page explains the broader relationship between inexpensive updates, noise and the progress made under a computing budget.
Sampling requires attention to dependence
Wikipedia’s Markov chain Monte Carlo article describes algorithms for drawing samples from a probability distribution by constructing a Markov chain. The aim is for the chain’s equilibrium distribution to match the target distribution. Samples then support calculations about a distribution that is difficult to analyse directly.
Successive states of a chain are related, so producing a long sequence is not the same as obtaining unrelated observations. Parallel computation can distribute separate chains or work within a sampling calculation, but the transitions still have to respect the target distribution. Splitting a dependent calculation arbitrarily can change the method. The statistical validity of the transition rule remains central even when more processors are available.
Dirichlet processes and model flexibility
A Dirichlet process is a distribution over probability distributions and can serve as a prior in an infinite mixture model. This offers a way to express clustering without fixing a finite number of components in advance. It changes the modelling question, but it does not remove the inference problem. A practical procedure still has to represent the relevant state and use observations to update beliefs about it.
Greater flexibility can bring additional bookkeeping. The inference method must identify what information remains local and what must be shared. Graph-parallel computation considers systems that expose sparse dependencies, including research examples involving sampling. The useful connection is between the dependency structure of an inference procedure and the operations its execution system permits.
Resampling and uncertainty about an estimator
Wikipedia’s bootstrap article describes estimating an estimator’s distribution by resampling data or a fitted model. Repeating a statistic on resampled collections gives information about its variability. This addresses a different question from inferring latent variables in a topic model, though both can require substantial repeated computation.
Separate bootstrap calculations offer opportunities for parallel execution when each repetition can run independently. Their expense still depends on the statistic being fitted and the amount of data each repetition needs. Repeatedly reading a large collection can dominate a seemingly simple estimator. The dataflow topic explains why reuse of data and intermediate results matters for such repeated tasks.
Across these methods, a useful result requires a clearly stated target and a clearly stated approximation. Computing more updates or more samples provides evidence only in relation to the method’s assumptions. Statistical checks and the cost of obtaining them therefore belong in the same plan as the main inference calculation.