Big learning meets biology: proteins, peptides and genomes
Biology was part of Big Learning's subject range from the beginning. The NeurIPS description of the 2011 workshop named bioinformatics among the areas where learning at scale arose. The 2012 accepted programme included the paper below. Its title connects the workshop to molecular research; the wider field now brings large collections of sequences, structures and simulation data into machine-learning studies.
The scale question is concrete: how can models learn from biological data whose volume, complexity and relationships exceed a simple sequential workflow? Protein structure prediction, proteomics and genomics each pose a different version of it. The relevant research concerns data representation, computational models, evaluation and resources, alongside the molecular questions being investigated.


Statistical Inference for Big Data Problems in Molecular Biophysics
Sequences, shapes and motion
Proteins contain long chains of amino-acid residues. Amino acids are organic molecules with amino and carboxylic-acid groups; peptide bonds link residues into chains. Peptides are shorter chains of this kind. For computational work, a sequence describes the order of residues, while a structure describes their arrangement in space. Those representations are related, but they give different information about the molecule being studied.
The protein-folding problem has several parts. A 2012 review in Science described questions about how sequence determines native structure, how folding occurs, and how a computer algorithm might predict a structure from sequence. It treated them as a research problem stretching back decades. Structure prediction addresses the computational question; it does not by itself supply a complete account of the physical folding process.
Molecular dynamics studies motion through computer simulation of interacting atoms and molecules. Its trajectories provide a different form of data from a static structure. Wikipedia's account notes that numerical integration errors can accumulate in long simulations. This matters for learning from simulation data: the computational procedure and its assumptions are part of the evidence, alongside the patterns a later model may extract.
Structure prediction and protein language models
The 2021 AlphaFold paper in Nature reported a neural-network approach to structure prediction that incorporated physical and biological knowledge and multiple-sequence alignments. Its evaluation concerned predicted structures. A companion study of the human proteome reported confidence measures and regions likely to be disordered, showing that the resulting dataset needed interpretation rather than one uniform confidence label.
A 2023 paper in Science reported protein structure inference from sequence using a large language model. The authors described structural information emerging in learned representations as models of protein sequences were scaled. This was a sequence-learning approach to a structural question. Its computational predictions belonged to research modelling; they were distinct from structures determined through physical measurements.
Public data underlie this work. Wikipedia’s AlphaFold account describes training with protein structures from the public Protein Data Bank. The Protein Data Bank holds molecular structures deposited from experimental studies, and RCSB PDB supplies tools for searching, visualisation and analysis of structural data. EMBL-EBI's account states that its open data contributed to AlphaFold's development. The AlphaFold Protein Structure Database makes predictions openly available, while also acknowledging limitations in the system. Experimental structures and predicted models are different kinds of record.
Design and molecular interactions
The 2024 chemistry Nobel announcement recognised computational protein design and protein structure prediction. Its accompanying explanation distinguished predicting a sequence's structure from designing proteins with new sequences. These are related computational directions with different questions: one starts from a sequence to estimate a structure; the other explores possible molecular designs. The recognition concerned that research, rather than a general claim about any particular application.
A 2024 review of functional protein design organised learning models around sequence data, structures and functional labels. It discussed outstanding challenges and the role of assays and benchmarks. Another 2024 review, focused on peptide structure prediction and design, highlighted structural flexibility and limited data as additional difficulties. Both accounts make the available evidence and the choice of representation central to the computational task.
A 2024 review in RSC Chemical Biology examined models for predicting peptide–protein interactions. It described peptide flexibility, limited structural information and the computational cost of conventional approaches as challenges. Predicting a complex's arrangement or interaction is a research question about molecules and models. The review's scope does not make a computed interaction an experimentally established observation.
Proteomics, genomics and broader research
Proteomics studies proteins across biological systems; genomics studies an organism's genome and the relationships among its elements. The National Human Genome Research Institute defines genomics around all of an organism's DNA. The Human Genome Project provided an earlier example of coordinated mapping and sequencing, while bioinformatics develops computational approaches for interpreting large and complex biological datasets.
The review Deep Learning in Proteomics examined tasks including peptide sequencing, spectrum prediction and protein structure prediction, and discussed limitations and future directions. A genomics primer examined deep learning for genome analysis. A separate review of health research discussed computer vision, language processing and other computational techniques. NIBIB's educational account emphasises dataset quality and fairness across groups, keeping evaluation and data preparation within the research problem.
Public infrastructure and computing cost
EMBL-EBI provides biological data resources and bioinformatics services. NCBI provides access to biomedical and genomic information, and NLM maintains information collections and electronic resources while supporting biomedical informatics research. These public institutions connect molecular data with tools and literature. Their work is part of the infrastructure that allows researchers to compare models and analyse data across studies.
The review of environmental impacts in protein science examined the energy and carbon costs of molecular simulations, interaction inference and structure prediction. Its findings make computing cost part of the scientific workflow. That returns to Big Learning's central question: model design and data scale have consequences for the systems and resources required to carry out the research.