On September 10th, OpenAI introduced the research approach of the Cé sar de la Fuente laboratory at the University of Pennsylvania: The team uses their own developed deep learning models to search for molecules with potential antibacterial properties in existing and extinct organisms' genomic and protein databases. At the same time, they utilize ChatGPT and Codex to discuss hypotheses, write and modify code, process data, and integrate ideas from different disciplines. According to the officials, this method can reduce the initial search for candidate molecules from several years to just a few hours. The key terms here are “candidate” and “initial search,” not creating a new drug in just a few hours.
Antimicrobial resistance is causing existing drugs to become increasingly ineffective. Research cited by OpenAI shows that in 2021, approximately 5 million deaths were related to bacterial resistance. If this trend continues, the annual burden is likely to nearly double by 2050. Traditional drug development often starts with known compounds and makes structural modifications. As the more easily discoverable directions have been explored repeatedly, the potential for new discoveries has gradually diminished. De la The Fuente team took a different approach: they treated DNA and protein sequences as information, searching through life science databases for patterns that could lead to the formation of active peptides or other antibacterial molecules.
AI is best at narrowing down the search scope, not at announcing answers on behalf of the laboratory.
Digital genomics allows researchers to search for clues across plants, animals, microorganisms, and even extinct species. However, the larger the amount of data, the harder it is to find truly useful signals. Deep learning models can learn the associations between sequences and activity, filtering out less relevant candidates from a vast number of fragments. What researchers would have previously had to design, synthesize, and test one by one can now be focused on a smaller list of possibilities. The key savings are not in canceling experiments, but in reducing a large number of low-probability attempts.
ChatGPT and Codex undertake a different level of collaboration. Team members come from fields such as biology, chemistry, computer science, and engineering, with knowledge boundaries that do not overlap. Biologists can use Codex to create data processing programs, while programmers can also understand experimental issues more quickly; ChatGPT helps the team to develop hypotheses, organize literature, and discuss connections between different fields. De la Fuente regard the shared workspace as a collective discussion board where members contribute both successful ideas and directions that did not succeed, allowing subsequent explorations to retain context.
This approach is more realistic than simply having a general model predict drugs. Professional deep learning models are responsible for filtering specific sequences, general assistants handle code and knowledge collaboration, while human researchers decide on the problems to be studied, check the results, and design experiments. Each layer has its boundaries: large language models may generate incorrect citations or unreliable code, and professional models can also inherit biases from the training data. Teams must retain versions of the data, model parameters, filtering rules, and records of manual modifications in order to understand how a candidate makes it onto the experimental list.
It is only after a candidate compound appears that the work truly becomes difficult. Researchers need to confirm whether it can kill the target microorganisms, measure its effective concentration, and check whether it also damages human cells. Chemists may continue to modify the molecule to improve its stability, activity, and safety; subsequently, they must assess the toxic dose, the speed at which microorganisms develop resistance, the absorption, distribution, and metabolism of the molecule in the body, as well as whether it can be manufactured stably. Even if all pre-clinical steps are successful, it still has to go through regulatory review and phased clinical trials.
From "hit candidate" to available drug, the failure rate and evidence standards will not disappear due to AI.
The funnel for drug discovery is extremely narrow. Preliminary screening models can compress millions of sequences down to a few thousand or dozens of candidates, but a large number of molecules are eliminated at each stage. Some are effective in the test tube but degrade rapidly in the human body; others have strong bactericidal capabilities but are too toxic; and there are also those that are difficult to manufacture or cannot reach the site of infection. Increasing the speed of preliminary screening by several orders of magnitude would create more experimental opportunities, but it might also shift the bottlenecks to synthesis, animal testing, and clinical resources.
Therefore, when evaluating such AI research, it is not sufficient to look at how many “new molecules” the model generates. More meaningful indicators include the experimental hit rate, structural differences from known antibiotics, selectivity, the speed of resistance development, manufacturability, and whether results can be replicated by independent laboratories. If AI generates ten thousand candidates, but only a very small number pass the initial experiments, the computational speed does not automatically translate into research and development efficiency. Conversely, even if the number of candidates is not large, as long as the subsequent hit rate is improved, it may truly shorten the R&D cycle.
The sources of data also need to be scrutinized. It is known that active molecules are more likely to be included in databases, while there is less data on rare species and poorly studied environments. Models may favor sequences that they are familiar with, thereby missing truly novel mechanisms. The genomes of extinct organisms often contain gaps and uncertainties in reconstruction, and any candidate findings must be verified through experiments. Researchers also need to guard against leaks between the training set and the test set, to prevent models from merely memorizing similar molecules and being mistakenly credited for discovering new patterns.
For hospitals and the public, this case should not be interpreted as “AI new antibiotics are now available.” The official article describes the scientific research workflow and candidate screening process; no molecule has been announced to have received clinical approval. The treatment of drug-resistant infections should still be based on doctors’ judgment, drug sensitivity test results, and approved medications. Packaging early research results as treatment recommendations not only exaggerates progress but may also pose real risks.
What this work truly demonstrates is the organizational approach of interdisciplinary research. Specialized models quickly scan life databases, and universal AI reduces the barriers to programming and knowledge communication, while laboratories provide factual verification. The combination of these three elements may bring more rare sequences into the scope of research, but it is still reproducible experimental data that ultimately convinces the scientific community. AI can reduce the time required to find the first candidate compound from years to hours; as for which compound will become a drug, however, it still requires step-by-step validation.












