Original Article
Development of consensus-based rules for AI-assisted Clinical Trial Agreement review: a Delphi study
Abstract
Background: Clinical Trial Agreement (CTA) review is a critical bottleneck in trial initiation due to its complexity and reliance on specialized expertise. While artificial intelligence (AI) can enhance efficiency, the absence of standardized, expert-endorsed, and large language model (LLM)-utilizable rules limits its application. This study developed a consensus-based rule set for LLM-assisted CTA review within China’s regulatory context.
Methods: A preliminary set of 74 rules (categorized into legal, financial, operational, and quality management domains) was drafted based on International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use-Good Clinical Practice (ICH-GCP), Chinese Good Clinical Practice (GCP) guidelines, institutional Standard Operating Procedures (SOPs), and expert interviews. A Delphi survey of three rounds was carried out with the participation of 23 Chinese CTA experts recruited via purposive and snowball sampling from diverse regions and institutions. Eligibility criteria included ≥3 years of CTA review/management experience and no conflicts of interest. A rule was approved only if it met both pre-defined criteria: (I) achieved high or medium consensus, determined by mean score, standard deviation, and coefficient of variation, and (II) received an importance rating of “critical” or “important” from at least 85% of experts. Anonymity was maintained through online questionnaires to minimize bias.
Results: After three-round Delphi survey, high consensus was achieved on 76 rules (14 revised +2 new rules based on expert feedback). The panel showed strong engagement (response rate ≥95.65%), authority (composite reliability =0.89), and coordination (Kendall’s W =0.87–0.91, P<0.001). The final rule set covers 42 operational, 15 financial, 12 legal, and 7 quality management rules, with revisions focusing on linguistic precision (52.90%), structural optimization (28.40%), and critical additions.
Conclusions: This study establishes the first consensus-based, comprehensive, expert-approved rule set for LLM-assisted CTA review, providing a preliminary reliable framework to standardize and optimize the process. The Delphi methodology offers a potentially transferable blueprint for developing similar LLM-assisted agents in other jurisdictions, with adaptations to local regulatory and cultural contexts required. Limitations include a single-jurisdiction expert panel, which will be addressed in future research.

