Thousands of protein pairs from disease-causing viruses, such as monkeypox virus, have been added to the protein-structure database AlphaFold.Credit: Kateryna Kon/SPL
A database of predicted structures for nearly every known protein on Earth is being upgraded to better represent some of the least known and deadliest species: viruses.
Researchers added more than 8,000 virus protein dimers — pairs of interacting molecules — to the AlphaFold Protein Structure Database today. The resource contains predictions of the 3D structure of proteins, which are generated using the AlphaFold2 artificial-intelligence tool and made publicly available. The added proteins come from 23 virus families, all with human-infecting members, and are part of the launch of a ‘pandemic preparedness portal’ within the widely used and freely available AlphaFold database.
Earlier this year, researchers added predicted structures for 1.7 million pairs of interacting proteins from 20 widely studied organisms, including humans, mice and tuberculosis-causing bacteria. This marked the first time that protein complexes — such as enzymes comprised of two identical protein strands — had been included in the AlphaFold database, which is maintained by the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI) in Hinxton, UK.
“Many viral proteins do not act individually, they act in concert with partners,” says Joe Grove, a molecular virologist at the University of Glasgow, UK, who was part of the effort to add viral complexes to the database.
Blind Spot
Although the database, which has more than three million users, holds predicted structures for most known proteins, viruses have been a blind spot, says Grove. High-quality entries are lacking for many individual proteins, such as those found in flaviviruses, which include Zika and dengue viruses.
This is because when some viruses replicate, their RNA molecules are first translated into what’s known as a ‘polyprotein’, which is then cut up into individual functional molecules. Therefore, it’s not always clear from a virus’s genetic sequence where a viral protein begins and ends, which can lead to incomplete structure predictions.
To improve on this, researchers at the Swiss Institute of Bioinformatics in Geneva identified accurate sequences for thousands of viral proteins that are cut from polyproteins. In total, Grove and researchers at various organizations from around the world analysed the sequences of 41,774 proteins from around 2,800 viruses, including ones that cause mpox, measles and hepatitis B.
With these sequences, the researchers used AlphaFold2 to predict the structures for 40,746 homodimers — which are pairs of identical protein strands interacting — and nearly 1.7 million heterodimers, in which the pairs are different molecules. But, of these, only 2,749 of the homodimers and 5,279 of the heterodimers were deemed accurate enough to include in the AlphaFold database (although all of the predictions were made publicly available).
The predictions lack the sugar molecules that adorn many viral proteins, helping them to evade immune detection. Many viral proteins also work in complexes larger than dimers. For example, the spike protein that enables SARS-CoV-2 and other coronaviruses to infect host cells is made up of three identical proteins, as is HIV’s envelope-entry protein. Dimer predictions for such complexes won’t be included in the AlphaFold database, in many cases, because their dimer predictions weren’t accurate enough, says Grove.
Sameer Velankar, a bioinformatician at EMBL-EBI who is part of the project, says that including such ‘trimer’ predictions is an obvious next step. But for bigger complexes, it’s not always clear how many copies of each component they are made of.
Grove is eager to dig into the data to further his research on viral entry proteins. “The proof of the pudding is going to be in putting that data out there and having teams such as my own start to drill down,” he says.
