Linear mixing
THE CASE FILE
One claim.
Nine ways to test it.
Read the evidence in order. Each chapter removes one convenient assumption and asks whether learning finally earns a robust advantage.
An interactive observatory for testing what HyperMix actually demonstrates: in this benchmark, a well-calibrated spatial matched filter leads or ties the learned detector.
MF 0.0110 vs unmixer 0.0136
unmixer interval +0.0096 wider
THE CASE FILE
Read the evidence in order. Each chapter removes one convenient assumption and asks whether learning finally earns a robust advantage.
CHAPTER 01 · SIGNAL
Target SNR measures the target contribution against noise, not the energy of the entire scene. Move the control to inspect the low-signal regime.
20 log₁₀(target RMS / noise RMS)SPECTRAL MISMATCH
The implanted target does not change. Only the signature supplied to the detector is shifted along the spectral index.
Now remove the convenient assumption that the lab signature reaches the sensor unchanged.
↓CHAPTER 02 · PHYSICS
Measured spectra, spectral response, atmosphere, and bilinear mixing are opt-in controls. The oracle target already knows the transformation; the lab target does not.
Linear mixing
USGS + bioHSI
Gaussian response
Mismatch appears
Full forward model
Use only declared sensor metadata to transform the lab signature before detection.
↓CHAPTER 03 · TRANSFER
Across 45 paired cases, the detector receives no target labels or score feedback. It only transforms the measured lab spectrum using declared wavelength, SRF, atmosphere, and illumination metadata.
Can a physical target family improve over the unchanged lab signature and approach the sensor-space oracle?
This is a deterministic physical transform on synthetic implants, not a learned win and not biological validation. The broad family was the pre-specified primary test, and it performed substantially worse.
HOLD OUT THE TARGETA narrow physical prior nearly reaches the oracle.Now remove the exact target entirely and keep the information available to every comparator explicit.
↓CHAPTER 04 · TARGET KNOWLEDGE
In this leave-one-host-out test, E. coli can use only P. putida as its family reference, and vice versa. The held-out spectrum appears only in implantation and the oracle ceiling.
SELECTED SUMMARY
Lower target-relative signal across the same scenes and holdouts
Correct reading: the family MLP loses to the other-host MF by 0.014 AUC and 0.042 Pd. The fully blind MLP does not beat spatial RX. Two measured hosts support only narrow family robustness, not broad chemical generalization.
FINAL CAUSAL TESTThe other host survives; learning still does not win.Now remove even the family and let a model learn background statistics without labels.
↓CHAPTER 05 · BACKGROUND
A shallow autoencoder learns only from unlabeled spectra in the test scene. It never receives labels, a target mask, or the target signature during training.
If real clutter is non-Gaussian, can scene-level background learning beat the spatial matched filter?
Both 95% confidence intervals are below zero.
This pre-specified shallow autoencoder is significantly worse on both metrics. It closes this simple instantiation, not every possible background-density model.
Global RX
0.539AUC · Pd 0.001Raw background AE
0.869AUC · Pd 0.108Ariel rewards calibrated uncertainty, not AUC alone. Ask whether learning can win on probability quality.
↓CHAPTER 06 · CALIBRATION
MF scores receive Platt scaling. The learned detector receives temperature scaling with bias correction, alone and as a three-member ensemble. Calibration and evaluation use disjoint target implants.
Can the learned ensemble beat the spatial matched filter on calibrated uncertainty while detection remains tied?
Both 95% intervals are above zero. The learned probabilities are significantly worse.
The pre-specified criterion required favorable NLL and ECE intervals. Neither was favorable, even after a fair calibration split.
NLL CI 0.00448–0.01778If detection is nearly saturated, inspect how many spectral channels actually carry the matched-filter result.
↓CHAPTER 07 · BANDS
Bands are ranked without implanted labels by the absolute full-scene matched-filter coefficient |C⁻¹(t−μ)|. The spatial MF is then recomputed using only the top-k bands.
Smallest k within 0.005 of the full-model mean. This is descriptive, not equivalence proof.
Top-3 carry only 9.8%–16.1% of absolute coefficient weight across the three scenes.
The crop-classification result does not transfer directly: fewer than three bands were not enough for this target-detection benchmark.
Keep the leading baseline and ask how much implanted signal is needed at a fixed false-alarm budget.
↓CHAPTER 08 · DETECTION LIMIT
Target-free scenes calibrate the threshold before evaluation. Noise stays fixed while abundance falls. The limit is the first tested grid point that sustains Pd 0.80 at every higher abundance.
Now calibrate abundance on separate implants and require honest prediction intervals.
↓CHAPTER 09 · CALIBRATED QUANTITY
Training, affine scale calibration, conformal residual calibration, and evaluation use disjoint implants. Scene-seed cases, not pixels, are the bootstrap units.
No calibrated abundance advantage
No calibrated abundance advantage. MAE is not significantly different. Both intervals cover above 90%, but the unmixer has significantly larger absolute bias and wider intervals.
Bring a score map, inspect its threshold, then finish with the boundaries of every claim.
↓LOCAL RESULT VIEWER
Drop a detector score map and inspect candidate pixels at different thresholds. Processing stays entirely in your browser.
PNG, JPEG, or WebP · up to 12 MB
Visualization only. Pixel brightness is treated as detector score. This does not run HyperMix inference on an RGB image.
The last chapter states exactly what this benchmark can and cannot support.
↓READ BEFORE CLAIMING
There is no naturally occurring remote biological target. Backgrounds can be real or measured, but targets are implanted.
Pellets are not remote surfaces. Beer-Lambert converts absorbance into a reflectance-like target.
The MLP does not see the raw cube. It recombines MF, ACE, and smoothed versions built from the nominal target.
Three scenes are not a population. The hierarchical intervals describe this benchmark, not every sensor or ecosystem.
One background model is not the whole model class. T7a closes the pre-specified shallow autoencoder, not every density estimator.
A calibrated score is still benchmark-specific. The split uses independent implants in the same three real backgrounds, not a new sensor population.
Band sparsity is target-aware here. The ranking knows the target signature and does not establish a universal three-band sensor.
The nominal transfer result is synthetic. It nearly reaches the oracle on implanted targets, but still needs independent sensor and biological validation.
The held-out family has only two measured hosts. Near-oracle transfer between E. coli and P. putida is narrow robustness, not broad chemical generalization.
The detection limit is a simulator fraction. It uses an exact sensor-space target and implanted blobs, not biological concentration or field validation.
Abundance intervals are conditional on detection. They cover pixels already declared as target and do not include the uncertainty of missing the target itself.