This collaborative publication marks the official launch of Phase 2 of Target 2035, an open science initiative aimed at developing pharmacological tools for the entire human proteome by the year 2035. The roadmap provides a framework for how academia, industry, contract research organizations (CROs), and small-to-medium enterprises (SMEs) can work together under open science principles to generate high-quality, standardized protein–ligand binding data at scale. This data will serve as the backbone for improving predictive models and for exploring the full landscape of human proteins, including many that have historically been neglected. The overarching goal is to make early-stage drug discovery faster, more affordable, and more accessible.
Central to the initiative is the creation of a publicly available, high-quality dataset of protein–small molecule binding interactions, produced using standardized, high-throughput screening methods. Researchers around the world are encouraged to contribute proteins for use in these screening campaigns, thereby powering the development of machine learning–ready binding datasets. The AI and machine learning community is being actively engaged through open benchmarking competitions designed to test, compare, and refine predictive models for small-molecule binding. All resulting data — both positive and negative — along with methods and materials, will be openly shared via public repositories, enabling global reuse and supporting continued model development.
Through this roadmap, the scientific community is committing to the open sharing of data, tools, and methodologies to foster a collaborative ecosystem capable of tackling some of the most urgent biomedical challenges. Organisations are encouraged to get involved by contributing purified proteins to support assay development, participating in open benchmarking challenges, joining the open science machine learning network MAINFRAME, and becoming members of the Target 2035 Consortium to help shape the future of universally accessible pharmacological tools.
Dr. Aled Edwards, CEO of the SGC, stated, “Our goal is to democratize early-stage drug discovery. We want every scientist, regardless of institution, to access reliable data and tools that can help them explore the potential of any protein.” Prof. Matthew Todd, UCL Lead for the SGC, added, “High quality datasets are key to effective machine learning. The roadmap paper describes an approach to generate trillions of robust, reliable experimental data points of protein-ligand interactions. By sharing such datasets openly, we can as a community improve our ability to predict small molecules that bind novel proteins. This would not only accelerate drug discovery, but improve our basic understanding of biology. That’s an exciting prospect.”
Links: