AI Maps the Protein Shape Universe, Revealing Hundreds of Thousands of New Structural Families
Scientists have taken a major step toward decoding the full diversity of life’s molecular machinery by organizing millions of predicted protein structures into related families...

Scientists have taken a major step toward decoding the full diversity of life’s molecular machinery by organizing millions of predicted protein structures into related families using artificial intelligence and a powerful new comparison technique. The effort reveals around 700,000 previously unclassified protein shape groups and identifies a small number that appear to be unique to humans, offering fresh insights into evolution, biology and drug discovery.
The work draws on the AlphaFold Protein Structure Database, a public resource containing AI-predicted 3D structures for more than 200 million proteins from organisms across the tree of life. By systematically clustering these structures for the first time, researchers have created what they describe as a panoramic map of the “protein shape universe.”
A New Way to See Protein Diversity
The study was led by Martin Steinegger, assistant professor at the School of Biological Sciences, Seoul National University, whose research focuses on ultra-fast methods for comparing massive biological datasets.
Traditionally, protein structure comparisons were carried out in small, targeted studies—often limited to a single protein family. Steinegger’s team instead clustered all available AlphaFold models at once, revealing large-scale structural patterns that only emerge at this global level.
“We are entering a phase where computational approaches allow us to explore protein space at an unprecedented scale,” Steinegger said, noting that older methods could have required many years of continuous computing to achieve a similar result.
How the Clustering Worked
To manage the enormous dataset, the researchers used Foldseek Cluster, an algorithm that converts complex protein shapes into a compact structural alphabet. This allows structures to be compared far more quickly than traditional alignment methods.
Using this approach, the team condensed hundreds of millions of predicted structures into roughly 2.3 million representative structural families. Many of the smallest clusters lacked functional annotations in existing databases, suggesting they may contain entirely new protein folds or biological activities.
Because the algorithm scales efficiently with data size, it can handle databases containing hundreds of millions of proteins without becoming computationally impractical.
Shedding Light on “Dark” Proteins
A major focus of the study was on so-called “dark clusters”—groups of proteins that do not resemble any previously known structures. From these, researchers selected tens of thousands with high-confidence predictions and searched them for features such as binding pockets or catalytic sites.
Some clusters were found to exist in only a single species, consistent with theories of de novo gene birth, where new protein-coding genes evolve from noncoding DNA.
When the team examined human proteins specifically, they found that truly human-only protein folds are rare. Instead, most human proteins belong to structural families that extend deep into evolutionary history, reinforcing the idea that evolution repeatedly reuses ancient molecular building blocks.
Unexpected Links in the Immune System
One of the most striking findings involved proteins of the human immune system. Several clusters connected human immune proteins with bacterial counterparts, implying shared structural solutions that predate complex animals.
For example, gasdermins—proteins that form pores in cell membranes during inflammatory cell death—were found to share a core structural domain with bacterial proteins. Similarly, the human bactericidal permeability-increasing protein (BPI) clustered with bacterial proteins of similar architecture, hinting at deep evolutionary roots.
The analysis also linked human DNA-sensing proteins, such as AIM2, to proteins found in gut bacteria. These connections would have been nearly impossible to detect through sequence comparisons alone, because the amino acid sequences have diverged extensively over time.
Why Structure Matters
Protein structure tends to be conserved far longer than protein sequence, making it a powerful tool for tracing evolutionary relationships that sequence-based methods miss. By organizing AlphaFold’s predictions into structural families, the new database provides researchers with an atlas for hypothesis generation, helping them infer functions for uncharacterized proteins.
The implications extend beyond basic biology. For drug discovery, dark clusters containing predicted binding pockets may represent untapped therapeutic targets—proteins that have never been studied experimentally and are not addressed by existing medicines.
A Foundation for Future Discovery
The researchers say their clustering framework offers a foundation for future experimental work, guiding scientists toward proteins and structures that have never been explored in the lab.
As AI-driven structure prediction continues to expand, this kind of large-scale organization may prove essential for turning an overwhelming flood of data into actionable biological knowledge—reshaping how scientists understand proteins, evolution, and disease.
