Select Page

August 26, 2026

In their article, Lang et al (2026) discuss the issues of assessing quality in preclinical research statistically in the age of the emerging use of artificial intelligence (AI). They aimed to show how large language models (LLMs) and more generally foundation models can enhance rather than undermine the quality of preclinical research while highlighting important role of human expertise, statistical rigor, and study design. They then mentioned two scenarios assuming that statistical support is currently optional: 1)When experts are present, AI should augment rather than replace their role and 2)When experts are absent, the introduction of autonomous AI by nonexperts poses additional risks beyond those already present in current practice.

They had described these stages as a hierarchy like a pyramid. For example, if the lower level is compromised in terms of quality than the upper levels will not be able to fix this. They suggested that the risk inherent in this dependency can be addressed by the pharmaceutical industry’s well-established concept of Quality by Design, which is a framework that promotes proactive planning and control to ensure that quality is built into each step of a process rather than tested into the output afterwards or addressed reactively at later stages (Yang et al. 2025; U.S. Food and Drug Administration, 2025).

Some of the well-known risks of autonomous AI tools include hallucinations, incomplete suggestions, and overconfidence while the benefits which include efficacy, as time-consuming tasks can be automated.  In terms of data generation, they state that autonomous AI can assist in creating drafts for statistical plans, particularly regarding sample size planning and randomization procedures and also that any document creation can be sped up immensely. However they do recommend that in order to ensure quality, supervision by a statistics expert is mandatory to assess whether all relevant aspects of the experimental design are covered and phrased correctly and also to ensure that the sample size and randomization process are appropriate.

Also, their suggestion for data analysis is that many programming and coding tasks can be effectively managed by autonomous AI support. However, once again, they recommend that accountability of results should be part of an expert role and not delegated to a nonexpert who only understand running the AI to create code but doesn’t understand how to create the proper checks and oversight.

As they mentioned, although it is technically feasible to delegate entire analytical tasks to AI systems such as ChatGPT, the risks associated with hallucinations, incorrect decisions, and a lack of reproducibility are too substantial to be ignored without thorough review (Dobler et al., 2025). Also, when autonomous AI is employed to suggest analytical methods, flaws become apparent. Another issue is that any training data for general foundation models often consists of more statistical examples chosen by nonexperts than by experts. Furthermore, this discrepancy arises from the lack of binding guidelines in the preclinical domain, which frequently results in statistical experts being excluded from the planning and analysis of experiments that are subsequently published.

They mentioned that the LLMs generate suggestions based on probability, but that they are more likely to propose statistical analyses that nonexperts would choose but that a statistical expert would contradict. One instance they mentioned is that the LLMs may recommend tests of assumptions on data derived from very small experiments, which are neither reliable nor conducive to reproducible choices.  Also, another critical point mentioned is that p values are often emphasized over effect sizes and confidence intervals, which are critical for comprehensive data interpretation (Wasserstein and Lazar, 2016).

Their final layer of the quality pyramid is data interpretation. Data scientists can create systems for data gathering and natural language applications to derive answers. For example, Bayer AG’s internal project PRINCE (Preclinical Information Center) to access and analyze preclinical safety data.  As they have stated, statistical expertise is one of the pillars for ensuring reliable and reproducible research results (Friedrich et al, 2022; Hoerl, 2025) and fostering collaboration among experts across all layers is paramount to enhancing the quality and reliability of preclinical research findings; the AI support can certainly improve research quality if experts are key stakeholders, ensuring responsibility and accountability.

Written by,

Usha Govindarajulu

Keywords: preclinical research, AI, LLMs, Quality by Design, ChatGPT

References:

 Dobler, D., H. Binder, A.-L. Boulesteix, et al. 2025. “ChatGPT as a Tool for Biostatisticians: A Tutorial on Applications, Opportunities, and Limitations.” Statistics in Medicine 44, no. 23–24: e70263. https://doi.org/10.1002/sim.70263.

Friedrich, S., G. Antes, S. Behr, et al. 2022. “Is There a Role for Statistics in Artificial Intelligence?” Advances in Data Analysis and Classification 16: 823–846. https://doi.org/10.1007/s11634-021-00455-6.

Hoerl, R. W. 2025. “The Future of Statistics in an AI Era.” Quality Engineering 1–13. https://doi.org/10.1080/08982112.2025.2556222.

Lang, T., Konietschke, F., Rahnenführer, J., Kubiak, R., Brendel, M., Binder, H. and Igl, B.-W. (2026), Ensuring Quality in Preclinical Research: The Importance of Being Human. Biometrical Journal., 68: e70145. https://doi.org/10.1002/bimj.70145

https://onlinelibrary.wiley.com/doi/full/10.1002/bimj.70145?saml_referrer

U.S. Food and Drug Administration. 2025. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products: Guidance for Industry and Other Interested Parties (Draft Guidance). FDA. https://www.regulations.gov/document/FDA-2024-D-4689-0003.

Vieira, V. E., C. Henrique, S. S. Kulkarni, et al. 2025. “From Data Silos to Insights: The PRINCE Multi-Agent Knowledge Engine for Preclinical Drug Development.” Frontiers in Artificial Intelligence 8: 1636809. https://doi.org/10.3389/frai.2025.1636809.

Yang, S., X. Hu, J. Zhu, et al. 2025. “Aspects and Implementation of Pharmaceutical Quality by Design by Design From Conceptual Frameworks to Industrial Applications.” Pharmaceutics 17, no. 5: 623. https://doi.org/10.3390/pharmaceutics17050623.

Wasserstein, R. L., and N. A. Lazar. 2016. “The ASA Statement on p-Values: Context, Process, and Purpose.” The American Statistician70, no. 2: 129–133. https://doi.org/10.1080/00031305.2016.1154108.