🇮🇪Ireland
16°C Partly Cloudy · Dublin
Live Updates
--:--:-- IST
Contributor sign in
Latest
Astellas Expands Its 330 Million Euro Tralee Biopharma Facility with a Second Aseptic Filling Line to Double Drug-Product Capacity ◆ Xeolas Pharmaceuticals Opens 158,000 Sq Ft State-of-the-Art Baldoyle Facility to Scale Specialty Medicine Manufacturing ◆ Priya Life Science Partners with Fleming for the 9th Annual Corporate Compliance & Transparency in Life Sciences Summit in Zurich ◆ Ireland Has the Capital and the Lessons: Digital Project Management Is How They Become Delivery ◆ Dunbar Pharma Brings First Plant-Derived Dronabinol API to UK Market Through IPS Pharma ◆ Leveraging Priya Life Science as a Data Tracker: The Ultimate Use Case & Career Guide ◆ The €100K Reality Check: Why a Six-Figure Pharma Salary in Ireland Feels Different Than in Switzerland or Germany ◆ Ireland's €93.8 Billion Non-EU Pharma Export Engine: Trade Data, Destination Markets, and Economic Impact ◆ Astellas Expands Its 330 Million Euro Tralee Biopharma Facility with a Second Aseptic Filling Line to Double Drug-Product Capacity ◆ Xeolas Pharmaceuticals Opens 158,000 Sq Ft State-of-the-Art Baldoyle Facility to Scale Specialty Medicine Manufacturing ◆ Priya Life Science Partners with Fleming for the 9th Annual Corporate Compliance & Transparency in Life Sciences Summit in Zurich ◆ Ireland Has the Capital and the Lessons: Digital Project Management Is How They Become Delivery ◆ Dunbar Pharma Brings First Plant-Derived Dronabinol API to UK Market Through IPS Pharma ◆ Leveraging Priya Life Science as a Data Tracker: The Ultimate Use Case & Career Guide ◆ The €100K Reality Check: Why a Six-Figure Pharma Salary in Ireland Feels Different Than in Switzerland or Germany ◆ Ireland's €93.8 Billion Non-EU Pharma Export Engine: Trade Data, Destination Markets, and Economic Impact ◆
Clinical Trials in Spain / NCT07677202
Starting soon Observational

Large Language Models Versus Human Examiners for Grading Physiotherapy Clinical Cases

NCT07677202 · tracked via the Priya Life Science Spain tracker
Phase
Observational
Started
2026-08-01
Last updated
2026-06-30

Condition(s) studied

Educational AssessmentArtifical IntelligencePhysical Therapy Education

Investigational drug(s) / intervention(s)

LLM-based assessmentFaculty assessment (reference standard)

LLM-based assessment: Assessment of each anonymized examination by three large language models (for example, Claude, ChatGPT, and Gemini, in the versions available during data collection). Each model receives an identical standardized prompt embedding the study rubric and returns a score per criterion, a global score, and structured qualitative feedback. Each model is queried in duplicate in independent sessions under fixed generation parameters to estimate intra-model (test-retest) reliability, and outputs are compared across models to estimate inter-model agreement.

Faculty assessment (reference standard): Assessment of the same anonymized examinations by faculty with expertise in the course, applying the identical rubric, serving as the reference standard. In the preferred scenario, two faculty members score each examination independently (paired human correction); if faculty workload precludes this, a single expert faculty rating, or the official course grade already assigned, is used as the reference. Faculty and LLM raters are blinded to one another's scores.

Study summary

This study evaluates whether large language models (LLMs) can reliably assess written clinical-reasoning case examinations completed by undergraduate physiotherapy students, compared with faculty assessment. In the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree), students solve complex clinical cases that require clinical reasoning, technical knowledge, and therapeutic decision-making. These cases are traditionally graded by faculty, a time-consuming process that may show inter-rater variability.

A set of de-identified student case examinations will be assessed using the rubric currently applied in the course, which covers clarity and structure of clinical reasoning, integration of the biopsychosocial model (ICF and APTA frameworks), accuracy in identifying pain mechanisms, coherence between diagnosis, hypotheses, and treatment, originality and depth of analysis, and professional writing. Each examination will be scored independently by three LLMs (for example, Claude, ChatGPT, and Gemini), each receiving an identical standardized prompt that embeds the same rubric, and by faculty serving as the reference standard.

To avoid overloading faculty, full double human grading may not be feasible; the human reference will therefore consist of expert faculty grading by one independent rater or, when resources allow, two independent raters. In contrast, paired assessment is fully implemented across the AI models: each examination is scored by several LLMs, and each model is queried in duplicate, allowing the study to estimate agreement between models and the test-retest stability of each model.

The primary aim is to quantify agreement between LLM-generated scores and the faculty reference score. Secondary aims include agreement among the LLMs, test-retest reliability of each model, criterion-level agreement, the quality and usefulness of the qualitative feedback generated, the time and cost associated with each approach, and students' perceptions of the usefulness of human versus AI feedback.

The findings will clarify the strengths and limitations of LLMs as supportive tools for formative assessment in health-professions education and will inform criteria for their responsible and effective use. No LLM output will affect students' official grades, which remain the sole responsibility of faculty.

Eligibility

Sex
ALL
Min age
18 Years
Max age
—
Healthy volunteers
Accepted
Inclusion Criteria: * Students officially enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period. * Submission of a completed written clinical-reasoning case examination as part of the course. * Provision of informed consent for the anonymized examination to be used for educational-research purposes. Exclusion Criteria: * Refusal to provide, or withdrawal of, informed consent. * Blank, incomplete, or non-evaluable examinations (e.g., no developed written response). * Examinations that cannot be reliably de-identified prior to assessment.

Primary outcome measure(s)

Trial sites (1)

FacilityCityRegionStatus
Centro Superior de Estudios Universitarios La Salle Madrid Madrid

More Neuron, Spain trials in Spain

Official registry record

This page summarises publicly available registry data for informational purposes — not medical advice. Eligibility is determined by each study team; patients should discuss participation with their clinician.

View NCT07677202 on ClinicalTrials.gov ↗ ← All trials in Spain