The open-source system, called xvr, fits a separate model to each patient so it can line up X-rays taken during an operation with the CT or MRI scan the patient had beforehand. Each match comes back in seconds, accurate to within a millimeter. The paper, by researchers at MIT and collaborating hospitals, appeared in Nature in September.
In keyhole procedures, clinicians steer catheters and endoscopes by live X-ray. A flat image makes it hard to tell exactly where an instrument is and which way it points, and misjudging that risks complications. Matching the image to the 3D scan, a step called registration, helps locate the instrument inside the body. Done by hand, it means entering coordinates or marking anatomical landmarks on a monitor.
Earlier AI tools for the job needed large hand-labeled datasets and stayed tied to the anatomy they were trained on. By simulating the physics of imaging, xvr instead renders thousands of training X-rays from the 3D scan of the one patient it will serve. Lead author Vivek Gopalakrishnan told MIT News that because these images come from the patient's actual body rather than from a generative model, the system has nothing to hallucinate.
Training such a model from scratch needs about 12 hours, too long for an emergency. So the team pretrained a foundation model on whole-body scans from more than 2,000 patients. Fitting it to a new patient finishes in five minutes or so with no loss of accuracy.
On real X-rays from five hospitals, covering dozens of bones and organ systems in adults and children, xvr was an order of magnitude more accurate than earlier AI methods.
xvr is still research. The team is working with surgical-robot makers and clinical groups to turn it into navigation tools, and wants it fast enough to run in real time.