TL;DR An annotated CT volume is an exact replay world: any view can be rendered and any localized prediction checked against the mask. That lets a self-improving system trace every missed lesion to the stage that lost it, and revise that stage.
One 4B reader, held-out patients, evaluated once

The paper in six points
- The problem. Medical imaging AI misses things in 3D scans. A miss can come from any stage: which slices were looked at, how well the model sees, how it was trained, or how its findings were combined.
- Why self-improvement is hard here. A system that improves itself needs to know which stage to fix, and its own feedback can't tell it.
- Key idea. An annotated CT scan is an exact replay world. Any view can be rendered on demand, and any predicted location can be checked against the expert's tumour outline.
- How it works. Every missed lesion is traced to the stage that lost it. A frozen language-model designer revises that stage, and a change is kept only if it beats measured seed noise. The checker and the objective never change.
- What it found. Better training recipes improved precision but hit a ceiling. Attribution then showed the real bottleneck was perception, meaning how the model sees the image.
- Result. Revising perception gives 79% liver and 83% kidney lesion recall at under 0.4 false positives per patient, ahead of 13 frozen frontier models, with the biggest gains on small lesions.
Read the full abstract
Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-generated feedback and logged trajectories provide little guidance on which component should change. We introduce ReVision3D, an RSI system that leverages 3D volumes with spatially grounded annotations to determine where visual evidence is lost and recursively improve the corresponding visual capability. A frozen language-model designer proposes revisions to acquisition, perception, training, or inference, while the verifier and system-level objective remain fixed.
Our key insight is that an annotated volume forms an exact replay world for view rendering and spatial verification: unvisited views can be rendered on demand, and localized predictions can be checked directly against reference masks. This grounded feedback directs targeted revision, while only changes that improve beyond measured seed noise are retained. Each accepted change triggers renewed attribution, allowing the dominant bottleneck to shift across rounds. On abdominal CT, attribution identifies perception as the dominant remaining limitation. Revising that level enables ReVision3D to achieve 79% liver recall and 83% kidney recall at under 0.4 false positives per patient, outperforming the evaluated frozen multimodal foundation models, with the largest gains on small lesions.
Volumetric replay
Annotated volumes become exact replay environments: previously unvisited views are rendered on demand, and localized predictions are grounded in one shared 3D coordinate system.
Attribution-guided RSI
Failures are attributed across acquisition, perception, training and inference, and improvement is directed to the matching intervention space. The verifier, objective and acceptance rule stay fixed.
Bounded visual self-improvement
Different interventions move different failure modes. Perception adaptation lifts liver and kidney recall after recipe-level progress saturates, and the loop rejects changes that stay within seed noise.
The gain is concentrated on small lesions
Recall at ≤ 0.5 false positives per patient on the held-out test patients (paper, Table 1). Frontier models read 8 uniform slices, ReVision3D every slice. Under the same every-slice protocol, the strongest frozen reader reaches 0.34 liver recall at 0.5 false positives per patient (Appendix Table A3). Hover a dot for the model.
Full table
Paper figure

Only a better way of seeing shrinks the never-seen share
Training recipes (reinforcement learning, then a learnability-aware recipe) raised the objective J from 0.379 to 0.529, mostly by improving precision; the maximum reachable recall stayed near 0.57–0.58, and none of fourteen further recipe proposals beat noise. Attribution then pointed at perception. On the test patients, each lesion is either never flagged on any slice, flagged and then discarded to meet the false-positive budget, or kept.
Lesions ≥ 1 mL at 0.5 false positives per patient, test patients (paper, Figure 3b). The same 4B reader in all rows; only the perception configuration differs.
Paper figure

The designer's search, one change at a time
Mean liver and kidney development objective J̄ per round (paper, Figure 4a). Each proposal changes one setting from the best so far. The two-epoch design is accepted after an independent second seed confirms it.
Paper figure
