Everything above comes from the data below. Every number here is recomputed from the model's saved reads and matches the paper.
Each step below is one version of the system, from a fixed pipeline to the version the loop designed. Step through them and watch lesions on real test patients go from missed to found. Every dot is a lesion of at least 1 mL in the reference mask, and every × is a false positive. All numbers come from the reads shipped in this repository, and J matches Table 1 exactly.
The first two steps were set by hand. In step 3, loss attribution showed that most lesions were never detected on any slice, so the designer revised the perception level. Step 4 is the inference program the designer wrote over the candidates from all four of the system's readers.
Select a patient to see where each version placed its findings.
All views are in radiological convention, with the patient's right on the viewer's left. The CT viewer shows the axial slices the reader actually read: the slice where each lesion is largest, and the slices of kept false positives. Amber boxes are the reference lesions from the mask. The reader marks a point per finding, not a box. Below the viewer, each version's findings are projected onto one axial plane. The patient shown first is the one that gained the most lesions from step 1 to step 4.
These are the trials of the perception round on development patients. Each proposal changes one setting, gets a short warm-start run, and is scored by macro J. It is kept only if it clears the noise floor measured from two baseline seeds. Select a trial to read the designer's reasoning.
At the inference level, the designer rewrites score(c), the function that ranks the merged candidates. This needs no training, so each revision is scored on development patients right away. A revision is accepted only when a paired bootstrap gives P(ΔJ > 0) ≥ 0.9.