Preprint · arXiv 2026

ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception

The loop, on a real held-out patient

Every panel is real data from the paper's test evaluation.

findings misses loss report one change score better: new reader not better: discard, propose again fixed rules · the AI can't edit them Readevery CT slice Checkvs expert outlines Trace misseswhere each was lost Proposethe LLM designer Testshort training run Beatsnoise?

TL;DR An annotated CT volume is an exact replay world: any view can be rendered and any localized prediction checked against the mask. That lets a self-improving system trace every missed lesion to the stage that lost it, and revise that stage.

Results at a glance

One 4B reader, held-out patients, evaluated once

Overview of ReVision3D: volumetric replay renders axial, coronal and sagittal views under a fixed verifier; loss attribution traces every miss to its stage; an LLM designer revises the responsible level (acquisition, perception, training recipe or inference), keeping a change only if the development objective J improves beyond seed noise.
Overview. (1) Volumetric replay renders alternative CT views and verifies localized predictions with a fixed, mask-grounded verifier. (2) Loss attribution traces each missed lesion to the stage where the evidence was lost. (3) A frozen LLM designer revises the responsible level. A change is kept only if the development objective J improves beyond measured seed noise, after which attribution is recomputed.
Abstract

The paper in six points

Read the full abstract

Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-generated feedback and logged trajectories provide little guidance on which component should change. We introduce ReVision3D, an RSI system that leverages 3D volumes with spatially grounded annotations to determine where visual evidence is lost and recursively improve the corresponding visual capability. A frozen language-model designer proposes revisions to acquisition, perception, training, or inference, while the verifier and system-level objective remain fixed.

Our key insight is that an annotated volume forms an exact replay world for view rendering and spatial verification: unvisited views can be rendered on demand, and localized predictions can be checked directly against reference masks. This grounded feedback directs targeted revision, while only changes that improve beyond measured seed noise are retained. Each accepted change triggers renewed attribution, allowing the dominant bottleneck to shift across rounds. On abdominal CT, attribution identifies perception as the dominant remaining limitation. Revising that level enables ReVision3D to achieve 79% liver recall and 83% kidney recall at under 0.4 false positives per patient, outperforming the evaluated frozen multimodal foundation models, with the largest gains on small lesions.

Volumetric replay

Annotated volumes become exact replay environments: previously unvisited views are rendered on demand, and localized predictions are grounded in one shared 3D coordinate system.

Attribution-guided RSI

Failures are attributed across acquisition, perception, training and inference, and improvement is directed to the matching intervention space. The verifier, objective and acceptance rule stay fixed.

Bounded visual self-improvement

Different interventions move different failure modes. Perception adaptation lifts liver and kidney recall after recipe-level progress saturates, and the loop rejects changes that stay within seed noise.

Against 13 frozen frontier models

The gain is concentrated on small lesions

ReVision3Dbest frozen model in each size bineach of the 13 frozen models

Recall at ≤ 0.5 false positives per patient on the held-out test patients (paper, Table 1). Frontier models read 8 uniform slices, ReVision3D every slice. Under the same every-slice protocol, the strongest frozen reader reaches 0.34 liver recall at 0.5 false positives per patient (Appendix Table A3). Hover a dot for the model.

Full table
Paper figure
Paper Figure 2: ReVision3D against frozen frontier models on liver, recall versus false positives and recall by lesion volume.
Loss attribution

Only a better way of seeing shrinks the never-seen share

Training recipes (reinforcement learning, then a learnability-aware recipe) raised the objective J from 0.379 to 0.529, mostly by improving precision; the maximum reachable recall stayed near 0.57–0.58, and none of fourteen further recipe proposals beat noise. Attribution then pointed at perception. On the test patients, each lesion is either never flagged on any slice, flagged and then discarded to meet the false-positive budget, or kept.

never seenseen, then discardedkept

Lesions ≥ 1 mL at 0.5 false positives per patient, test patients (paper, Figure 3b). The same 4B reader in all rows; only the perception configuration differs.

Paper figure
Paper Figure 3: accepted revisions on liver development patients and where lesions are lost on the test set.
Perception-level self-improvement

The designer's search, one change at a time

trialsecond seedbest so fardefault design, two seeds

Mean liver and kidney development objective J̄ per round (paper, Figure 4a). Each proposal changes one setting from the best so far. The two-epoch design is accepted after an independent second seed confirms it.

Paper figure
Paper Figure 4: recursive improvement of the perception configuration and the never-detected share.
Citation

BibTeX