Glossary

Proteomics

The large-scale study of the full complement of proteins expressed by a cell, tissue, or organism — connecting genomic information to biological function, and a field where AI is accelerating both structure prediction and abundance analysis.


What it means

Proteomics is the systematic study of the proteome — the complete set of proteins expressed in a cell, tissue, or organism at a given time and condition. Unlike the genome (which is largely static), the proteome is highly dynamic: protein abundance, modification state, and localization change in response to disease, environment, development, and treatment.

The main branches of proteomics:

Expression proteomics — measuring how much of each protein is present under different conditions, using mass spectrometry. Compares protein abundance between healthy and diseased tissue, treated and untreated cells, or developmental stages.

Structural proteomics — determining the 3D structure of proteins and protein complexes. This is where AlphaFold has had its largest impact: before AF2, structures required years of experimental work; AF2 provided predicted structures for virtually all known proteins in weeks.

Post-translational modification (PTM) proteomics — mapping phosphorylation, ubiquitination, glycosylation, and other chemical modifications that regulate protein function. PTMs are not encoded in the genome and cannot be predicted from sequence alone.

Interaction proteomics — identifying which proteins physically interact (the “interactome”), revealing signaling pathways, protein complexes, and functional networks.

Mass spectrometry: the core technology

Most quantitative proteomics is done by mass spectrometry (MS): proteins are digested into peptides, separated by liquid chromatography, and the mass and fragmentation pattern of each peptide is measured. Database searching matches measured spectra to theoretical peptides from a protein sequence database.

Modern MS-based proteomics can routinely identify and quantify 5,000–10,000+ proteins from a single experiment.

AI in proteomics

AlphaFold 3 extended structure prediction to protein complexes, protein-DNA interfaces, and protein-small molecule interactions — directly enabling structural proteomics at scale.

DIA-NN and MSFragger — AI-accelerated tools for processing mass spectrometry data with improved sensitivity for protein identification and quantification.

Protein language models (ESM3, ESM2) generate sequence embeddings that capture evolutionary and functional relationships, used for property prediction and variant effect assessment without requiring structure prediction.

AlphaMissense predicted functional consequences of all possible missense variants in human proteins — effectively an AI proteomics annotation tool.

Multi-omics integration

Proteomics is most powerful when integrated with genomics (which genes are active?), transcriptomics (which genes are being transcribed?), and metabolomics (what small molecules are present?). Multi-omics integration — combining these data layers to build a more complete picture of cellular state — is an active area where AI dimensionality reduction and network methods are increasingly used.