Chinese researchers at Peking University and East China Normal University published a breakthrough framework in Nature on July 25, 2026. Entitled ContactSeek, the AI tool uses AlphaFold3 contact probability to systematically engineer gene editing tools, improving base editing precision and reducing off-target effects.
ContactSeek AI Framework Developed by Peking University and East China Normal University
Scientists working across academic institutions in China have unveiled a new computational framework designed to overhaul gene editing tool engineering. Developed by a Peking University team led by Professor Yi Chengqi in collaboration with Professor Li Dali team at East China Normal University, the system is called ContactSeek. The methodology and findings were published in Nature under the title Precise DNA Base Editing Using AlphaFold3-Based Contact Modelling, introducing ContactSeek an AI framework that leverages AlphaFold3 contact probability to systematically engineer gene editing tools with significantly improved precision. The paper represents a convergence of AI structural biology and gene editing, demonstrating how protein structure prediction models can directly guide the engineering of therapeutic tools.
Tackling the Off-Target Challenge in Base Editors
Base editors are critical instruments for modern molecular biology and gene therapy development. The ContactSeek framework addresses the fundamental challenge in base editing: maintaining high on-target editing activity while minimizing off-target effects. Base editors combine a Cas protein with a deaminase enzyme to directly convert one DNA base to another without requiring double-strand breaks, making them critical tools for both basic research and gene therapy development. However, off-target editing where the editor modifies unintended genomic sites has limited clinical application.
To overcome this hurdle, the research team turned to advanced protein structure prediction models. They utilized AlphaFold3, the DeepMind AI model for protein structure prediction extended to protein-nucleic acid complexes, to predict on-target and off-target DNA-RNA-protein ternary complexes and systematically compare their interaction differences. The key insight: AlphaFold3 contact probability, not structural prediction, proved most valuable.
Why Contact Probability Outperforms Structural Prediction
Updated versions were designed to handle interactions between proteins and nucleic acids, as well as complexes of multiple proteins. So the team fed AlphaFold versions of a target DNA sequence, along with a guide RNA, the Cas9 sequence, and an enzyme that chemically modifies bases and can stick to Cas9. Unfortunately, it choked, placing one of the proteins in what was clearly the wrong location. Undeterred, the team simplified things and fed AlphaFold only the DNA, RNA, and Cas9 protein, since the latter is the primary factor determining its sequence specificity. This worked much better, producing a structure that agreed with ones determined by experiments with actual nucleic acids and proteins.
By comparing the structures AlphaFold generated when fed different on- and off-target sites, the researchers found a general pattern. By comparing structures generated from on-target and off-target sites, the researchers identified a distinct pattern. Many (about two-thirds) of the off-target sites caused the Cas9 protein to adopt a slightly different structure. But nearly all (over 95 percent) of them altered which amino acids contacted the RNA. So there are clearly some cases where Cas9 maintains its normal structure but amino acids within it flex around in ways that accommodate the mispaired bases of off-target sites.
AlphaFold Contact Probability Analysis Setup
Conveniently, AlphaFold was already set up to identify what is termed the contact probability,
namely, the chance that any two items, such as amino acids or nucleotides, are within a very small distance (eight Angstroms). Contact probability sensitively captures changes from single-base-pair mismatches and single-amino-acid mutations that predicted structures miss. The researchers could take the output of the contact probability analysis for on- and off-target sites and compare them, identifying exactly which amino acids in Cas9 have altered contacts when there’s a mismatch between the guide RNA and the DNA. They termed this computerized analysis setup “ContactSeek.”

Engineering New Cas9 Variants for Precision Editing
Cas9 Variants Engineered for Adenine Base Editors
On its own, ContactSeek tended to produce a large list of amino acids that shift around when bound to an off-target site. So the researchers focused on regions of the Cas9 protein where these amino acids clustered, viewing this as a sign that these areas were adapting to the differences caused by mismatched bases. They then began to test versions of Cas9 with different amino acids at these sites.

By integrating contact probability with high-throughput off-target editing sequencing data, ContactSeek systematically identified key amino acid residues in both Cas proteins and deaminases that determine editing specificity.
The team leveraged these insights to engineer a series of novel Cas9 variants for adenine base editors that maintain efficient on-target editing while significantly reducing DNA off-target editing.