About the Project
Building Better Feedback Across
UT Medical Schools
This project — "Scaling and Validating AI-Enabled Simulation Assessment Across University of Texas Medical Schools" — is an 18-month, $300,000 initiative funded by the UT REAL Health AI Pilot Program, which supports collaborative, implementation-focused AI initiatives across UT Health-Related Institutions.
How It Works
The UTSW production workflow moves from OSCE capture to AI grading and faculty review, turning recorded clinical encounters into actionable feedback for students.

In the UTSW production workflow, 1. Capture — multi-angle video from the simulation center flows into a learning management system; 2. Platform — SimRubrics authors AI-compatible rubrics, Elephant catalogs the multimedia data, and MAPLES orchestrates AI grading; and 3. Review — faculty review flagged items, adjudicate disagreements, and grades are reported to students with full provenance.
What is an OSCE?
An Objective Structured Clinical Examination (OSCE) is how medical schools assess whether students can demonstrate clinical skills in a standardized medical scenario.
Students are assessed on communication, history taking, physical examination, clinical reasoning, and medical knowledge. They rotate through rooms, each with a standardized patient presenting a different scenario. Many programs record encounters and use detailed rubrics to support structured clinical competency assessment.
The result is an enormous volume of encounters that need expert review. That's where AI comes in.

A UTSW OSCE workflow with medical students rotating through stations, encounters recorded from overhead cameras, and performance evaluated with clinical rubrics.
The Challenge
The educators who train tomorrow's doctors deserve better tools.
- 1
Grading Takes Too Long
Large OSCE administrations can require many hours of faculty and standardized-patient educator review. That workload can delay feedback that is most useful when delivered promptly.
- 2
Human Grading Varies
In a 300-encounter retrospective medRxiv preprint, standard human evaluators achieved quadratic weighted kappa = 0.732 against a physician-adjudicated reference. Consistent rubric interpretation and review remain important parts of reliable assessment.
- 3
Assessment Workflows Are Local
Assessment tools, rubrics, and infrastructure are commonly managed within each school. That makes it harder to compare practices, share validated tools, and adopt what works at peer institutions.
Our Approach
The MAPLES Platform
AI that helps educators give students faster, more consistent, and more detailed feedback on their clinical skills.
Preprint Agreement Result
A 300-encounter retrospective medRxiv preprint reported three-camera MAPLES agreement of quadratic weighted kappa = 0.830 against a physician-adjudicated reference, compared with 0.732 for standard human evaluators against the same reference. Faculty retain final authority and review cases routed for human judgment.
Three Years in Production
More than 7,000 encounters have been graded in UTSW production. The grant proposal estimates that the six participating medical schools collectively represent more than 3,200 medical students.
Infrastructure Feasibility Review
UT-HIP is one infrastructure option under technical feasibility review alongside UTSW-managed and site-specific paths. Final hosting, governance, and access controls have not been selected and remain subject to site-specific approvals.
Multimodal Assessment
Beyond text-based note grading: the platform supports audio, transcript, and video analysis for physical examination identification and clinical reasoning assessment.
Flexible Rubric Framework
AI-compatible rubric design that each school can customize to their own educational philosophy — shared tools and best practices, not one-size-fits-all templates.
Open Governance Playbook
Developing a reusable framework for multi-site IRB routing, data use agreements, and AI governance in medical education — designed for adoption beyond the UT System.
Program at a Glance
UTSW production, a preprint agreement result, and the six-school consortium.
Guiding Principles
The values that shape how we build and deploy AI in medical education.
- 1
Human-in-the-Loop
AI augments, not replaces, expert judgment. In the UTSW production workflow, prespecified low-scoring cases are routed for mandatory human review, and faculty retain final authority over grades.
- 2
Phased Deployment
Sites activate at their own pace along two axes: modality (text → transcript → video) and study design (retrospective → prospective). No site is forced into lockstep.
- 3
Transparency and Reproducibility
Every AI grading decision is logged with full provenance. Rubric versions, model parameters, and confidence scores are captured for audit and research.
Funding & Support
This project is one of several initiatives funded by the UT REAL Health AI Pilot Program ($300,000, awarded March 2026, 18 months). Infrastructure planning is being coordinated with UT Southwestern Information Resources and Enterprise Data Services, while the UT Health Intelligence Platform remains a prospective UT System path subject to hosting, governance, and site approvals. The project originated at the UTSW Simulation Center, built by the Jamieson Lab.