ReVision3D
ReVision3D · research preview

An AI that works out why it misses tumours, and fixes itself.

How it works

One loop. The AI runs it, and the rules check it.

findings misses loss report one change score better: new reader not better: discard, propose again fixed rules · the AI can't edit them Readevery CT slice Checkvs expert outlines Trace misseswhere each was lost Proposethe AI designer Testshort training run Beatsnoise?
The loop that turned the hand-built reader into the one the AI designed. Select any step, or let it play.
The evidence

Everything above comes from the data below. Every number here is recomputed from the model's saved reads and matches the paper.

ReVision3D · held-out test patients

Watch the loop improve a CT lesion detector

Each step below is one version of the system, from a fixed pipeline to the version the loop designed. Step through them and watch lesions on real test patients go from missed to found. Every dot is a lesion of at least 1 mL in the reference mask, and every × is a false positive. All numbers come from the reads shipped in this repository, and J matches Table 1 exactly.

Qwen3-VL-4B readerLiTS liver · 31 patientsKiTS23 kidney · 70 patientstest set, read once

From fixed pipeline to the loop's design

The first two steps were set by hand. In step 3, loss attribution showed that most lesions were never detected on any slice, so the designer revised the perception level. Step 4 is the inference program the designer wrote over the candidates from all four of the system's readers.

Organ
False-positive budget per patient
Or use the ← and → keys.
detected missed found at this step lost at this step ×false positive dot size ∝ log volume

Select a patient to see where each version placed its findings.

Recall at each false-positive rate

Where the lesions were lost

never detected on any slicedetected, dropped to meet the budgetkept

One patient, four versions

All views are in radiological convention, with the patient's right on the viewer's left. The CT viewer shows the axial slices the reader actually read: the slice where each lesion is largest, and the slices of kept false positives. Amber boxes are the reference lesions from the mask. The reader marks a point per finding, not a box. Below the viewer, each version's findings are projected onto one axial plane. The patient shown first is the one that gained the most lesions from step 1 to step 4.

Patient
Version

How the designer chose step 3

These are the trials of the perception round on development patients. Each proposal changes one setting, gets a short warm-start run, and is scored by macro J. It is kept only if it clears the noise floor measured from two baseline seeds. Select a trial to read the designer's reasoning.

How the designer wrote step 4

At the inference level, the designer rewrites score(c), the function that ranks the merged candidates. This needs no training, so each revision is scored on development patients right away. A revision is accepted only when a paired bootstrap gives P(ΔJ > 0) ≥ 0.9.

Organ