7,000+ runs · paired by world · 95% bootstrap intervals over seeds
Method comparison, paired by world.
Every method is run on the same simulated worlds, so differences are measured within a world and only then averaged. Intervals are 95% bootstrap over seeds; pixels are never treated as samples.
At the canonical world
Simulation Two numbers per method, never one: how much contamination it removed and how much clean signal it kept.
Clean signal transfer
signal no template can mimic that survived · 1 = intact
Contamination removed
1 = fully removed · 0 = untouched
Against signal–systematic correlation
Signal transfer as ρ grows
grey: the other realisable methods
Total transfer falls as 1 − ρ² for every template method: that loss is forced. Clean transfer is flat in
ρ for every method: whatever a method loses beyond the forced part, it loses whether or not the signal resembles the
templates.
Against signal strength
Aggregate scientific loss by signal-to-noise
lower is better · ρ = 0.3 · the components are shown on the other panels
Across worlds
Is any method consistently safe?
Share of contaminated worlds where the method is safe
across rho, SNR, family, sample-size and observing-condition slices
No single score ranks methods robustly: under the 27 weightings of the aggregate loss's three components, the best method changes with the weights. That is why every page reports the components.