HHYPERMIXOBSERVATORY
GitHub
OPEN BENCHMARK · AUDITED 06 AUG 2026

Detection without
the victory lap.

An interactive observatory for testing what HyperMix actually demonstrates: in this benchmark, a well-calibrated spatial matched filter leads or ties the learned detector.

103passing tests
3real backgrounds
2 / 4calibration / eval seeds
MITopen source
LATEST AUDIT · T12
CALIBRATED MAENo learned win

MF 0.0110 vs unmixer 0.0136

90% INTERVALSBoth cover

unmixer interval +0.0096 wider

Inspect the new evidence

THE CASE FILE

One claim.
Nine ways to test it.

Read the evidence in order. Each chapter removes one convenient assumption and asks whether learning finally earns a robust advantage.

01SignalDoes learning hold when the target fades?Spatial MF leads02PhysicsDoes sensor realism reverse the result?Mismatch dominates03TransferCan declared physics carry the lab target to the sensor?Narrow prior works04Target knowledgeWhat survives when the exact target is hidden?A narrow family survives05BackgroundCan raw scene statistics rescue learning?Not in this test06CalibrationCan learning win on honest probabilities?MF still wins07BandsIs the signal really carried by three bands?Not here08Detection limitWhat abundance reaches operational Pd?15–20% at FAR 1e-209QuantityDoes calibration rescue learned abundance?No aggregate advantage
01

CHAPTER 01 · SIGNAL

Turn down the signal.
See what holds.

Target SNR measures the target contribution against noise, not the energy of the entire scene. Move the control to inspect the low-signal regime.

TARGET SNR0 dB
DEFINITION20 log₁₀(target RMS / noise RMS)
AGGREGATED AUCINDIAN PINES · SALINAS · PAVIA U.
Spatial MF0.982
Learned detector0.972
Pixel MF0.908
ACE0.811
0.50 chance1.00 perfect

SPECTRAL MISMATCH

When the signature is wrong

The implanted target does not change. Only the signature supplied to the detector is shifted along the spectral index.

MF0.940AUC
SPATIAL MF0.990AUC
LEARNED0.987AUC
NEXT QUESTIONThe baseline survives low signal.

Now remove the convenient assumption that the lab signature reaches the sensor unchanged.

02

CHAPTER 02 · PHYSICS

From the lab
to the sensor.

Measured spectra, spectral response, atmosphere, and bilinear mixing are opt-in controls. The oracle target already knows the transformation; the lab target does not.

01

Linear mixing

Stylized control

ORACLE0.983
LAB TARGET0.983
Δ 0.000
02

USGS + bioHSI

Measured spectra

ORACLE0.995
LAB TARGET0.995
Δ 0.000
03

Gaussian response

10 nm SRF

ORACLE0.994
LAB TARGET0.995
Δ 0.001
04

Mismatch appears

SRF + atmosphere

ORACLE0.994
LAB TARGET0.913
Δ -0.081
05

Full forward model

Bilinear scenario

ORACLE0.983
LAB TARGET0.906
Δ -0.077
Key Phase B evidence

With SRF + atmosphere, knowing the target at the sensor is worth 0.081 AUC. Physical mismatch matters more than swapping the detector.

TRANSFER TESTPhysical mismatch hurts more than model choice.

Use only declared sensor metadata to transform the lab signature before detection.

03

CHAPTER 03 · TRANSFER

A narrow prior
closes the gap.

Across 45 paired cases, the detector receives no target labels or score feedback. It only transforms the measured lab spectrum using declared wavelength, SRF, atmosphere, and illumination metadata.

PRE-SPECIFIED TEST / T9

Can a physical target family improve over the unchanged lab signature and approach the sensor-space oracle?

01

Unchanged lab target

AUC0.96895% CI 0.962–0.974
Pd@FAR 1e-30.62695% CI 0.587–0.672
02

Nominal metadata transfer

AUC0.99095% CI 0.988–0.992
Pd@FAR 1e-30.76095% CI 0.727–0.787
03

Broad family subspace

AUC0.77795% CI 0.749–0.799
Pd@FAR 1e-30.36995% CI 0.340–0.394
04

Sensor-space oracle

AUC0.99295% CI 0.991–0.994
Pd@FAR 1e-30.77895% CI 0.744–0.805
PRIMARY · BROAD FAMILY MINUS LABAUC −0.191 [−0.217, −0.171] · Pd −0.257 [−0.293, −0.221]Failed decisively
SECONDARY · NOMINAL TRANSFER MINUS LABAUC +0.022 [0.016, 0.028] · Pd +0.134 [0.097, 0.169]Near oracle, but not the primary criterion

This is a deterministic physical transform on synthetic implants, not a learned win and not biological validation. The broad family was the pre-specified primary test, and it performed substantially worse.

HOLD OUT THE TARGETA narrow physical prior nearly reaches the oracle.

Now remove the exact target entirely and keep the information available to every comparator explicit.

04

CHAPTER 04 · TARGET KNOWLEDGE

Hide the target.
Audit what remains.

In this leave-one-host-out test, E. coli can use only P. putida as its family reference, and vice versa. The held-out spectrum appears only in implantation and the oracle ceiling.

SELECTED SUMMARY

0 dB

Lower target-relative signal across the same scenes and holdouts

Blind methods remain near chance
Exact-target oracle0.981
Other-host family MF0.979
Family MLP0.963
Fully blind MLP0.508

Correct reading: the family MLP loses to the other-host MF by 0.014 AUC and 0.042 Pd. The fully blind MLP does not beat spatial RX. Two measured hosts support only narrow family robustness, not broad chemical generalization.

FINAL CAUSAL TESTThe other host survives; learning still does not win.

Now remove even the family and let a model learn background statistics without labels.

05

CHAPTER 05 · BACKGROUND

The last
honest test.

A shallow autoencoder learns only from unlabeled spectra in the test scene. It never receives labels, a target mask, or the target signature during training.

RESEARCH QUESTION / T7A

If real clutter is non-Gaussian, can scene-level background learning beat the spatial matched filter?

01

Spatial matched filter

AUC0.98795% CI 0.968–0.997
Pd@FAR 1e-30.65095% CI 0.227–0.872
02

Spatial background autoencoder

AUC0.97695% CI 0.945–0.994
Pd@FAR 1e-30.32495% CI 0.087–0.544
PAIRED DIFFERENCE · AUTOENCODER MINUS MFAUC -0.011 · Pd -0.325

Both 95% confidence intervals are below zero.

No causal advantage

This pre-specified shallow autoencoder is significantly worse on both metrics. It closes this simple instantiation, not every possible background-density model.

Target-agnostic checks

Global RX

0.539AUC · Pd 0.001

Raw background AE

0.869AUC · Pd 0.108
NEW SCORECARDThe last simple causal detection test also fails.

Ariel rewards calibrated uncertainty, not AUC alone. Ask whether learning can win on probability quality.

06

CHAPTER 06 · CALIBRATION

A score is not
a probability.

MF scores receive Platt scaling. The learned detector receives temperature scaling with bias correction, alone and as a three-member ensemble. Calibration and evaluation use disjoint target implants.

T7B

Can the learned ensemble beat the spatial matched filter on calibrated uncertainty while detection remains tied?

01

Spatial MF + Platt

NLL0.05766lower is better
Brier0.01540lower is better
ECE0.00896lower is better
AUC reference0.9863 scenes · 24 cases
02

Learned ensemble + temperature

NLL0.06792lower is better
Brier0.01940lower is better
ECE0.01293lower is better
AUC reference0.9803 scenes · 24 cases
PAIRED DIFFERENCE · ENSEMBLE MINUS SPATIAL MFNLL +0.01026 · ECE +0.00397

Both 95% intervals are above zero. The learned probabilities are significantly worse.

Reliability at 0 dB Spatial MF + Platt Learned ensemble + temperature
Observed frequency
ideal
Predicted probability
No uncertainty advantage

The pre-specified criterion required favorable NLL and ECE intervals. Neither was favorable, even after a fair calibration split.

NLL CI 0.00448–0.01778
ECE CI 0.00208–0.00560
BAND AUDITLearning also loses on calibrated uncertainty.

If detection is nearly saturated, inspect how many spectral channels actually carry the matched-filter result.

07

CHAPTER 07 · BANDS

How sparse is
the signal?

Bands are ranked without implanted labels by the absolute full-scene matched-filter coefficient |C⁻¹(t−μ)|. The spatial MF is then recomputed using only the top-k bands.

SPATIAL MF AUC · 95% CI IN RESULTS
0.838
1
0.949
2
0.948
3
0.963
5
0.969
10
0.983
20
0.986
40
0.987
80
0.984
all
top-kall = 103–204 bands
TOP-3 MINUS ALL BANDS−0.036 AUC [−0.092, −0.000]
20 bands

Smallest k within 0.005 of the full-model mean. This is descriptive, not equivalence proof.

Top-3 carry only 9.8%–16.1% of absolute coefficient weight across the three scenes.

RESULT

The crop-classification result does not transfer directly: fewer than three bands were not enough for this target-detection benchmark.

OPERATIONAL QUESTIONThree bands are not enough in this benchmark.

Keep the leading baseline and ask how much implanted signal is needed at a fixed false-alarm budget.

08

CHAPTER 08 · DETECTION LIMIT

How much signal
is enough?

Target-free scenes calibrate the threshold before evaluation. Noise stays fixed while abundance falls. The limit is the first tested grid point that sustains Pd 0.80 at every higher abundance.

SENSOR FWHMREALIZED FARNOMINAL LODCONSERVATIVE LOD
8 nm0.0068015%20%
12 nm0.0063715%20%
20 nm0.0028520%>20%
QUANTIFY THE TARGETA detection limit says when, not how much.

Now calibrate abundance on separate implants and require honest prediction intervals.

09

CHAPTER 09 · CALIBRATED QUANTITY

An estimate needs
an honest interval.

Training, affine scale calibration, conformal residual calibration, and evaluation use disjoint implants. Scene-seed cases, not pixels, are the bootstrap units.

METHODMAE90% COVERAGEINTERVAL WIDTH
Calibrated MF0.01100.9710.0719
Calibrated unmixer0.01360.9870.0815
PAIRED DIFFERENCE · UNMIXER MINUS MFMAE +0.0026 [−0.0006, 0.0059] · width +0.0096 [0.0090, 0.0102]

No calibrated abundance advantage

RESULT

No calibrated abundance advantage. MAE is not significantly different. Both intervals cover above 90%, but the unmixer has significantly larger absolute bias and wider intervals.

INSPECT THE ARTIFACTCoverage alone is not interval efficiency.

Bring a score map, inspect its threshold, then finish with the boundaries of every claim.

LOCAL RESULT VIEWER

Bring your own
score map.

Drop a detector score map and inspect candidate pixels at different thresholds. Processing stays entirely in your browser.

Drop an image here

PNG, JPEG, or WebP · up to 12 MB

Awaiting a score map Local only
SOURCE MAP
H
THRESHOLDED VIEW
72
DISPLAY THRESHOLD72%
Your image never leaves this device.

Visualization only. Pixel brightness is treated as detector score. This does not run HyperMix inference on an RGB image.

FINAL READINGA convincing map is not biological validation.

The last chapter states exactly what this benchmark can and cannot support.

READ BEFORE CLAIMING

What this project
does not demonstrate.

  1. 01

    There is no naturally occurring remote biological target. Backgrounds can be real or measured, but targets are implanted.

  2. 02

    Pellets are not remote surfaces. Beer-Lambert converts absorbance into a reflectance-like target.

  3. 03

    The MLP does not see the raw cube. It recombines MF, ACE, and smoothed versions built from the nominal target.

  4. 04

    Three scenes are not a population. The hierarchical intervals describe this benchmark, not every sensor or ecosystem.

  5. 05

    One background model is not the whole model class. T7a closes the pre-specified shallow autoencoder, not every density estimator.

  6. 06

    A calibrated score is still benchmark-specific. The split uses independent implants in the same three real backgrounds, not a new sensor population.

  7. 07

    Band sparsity is target-aware here. The ranking knows the target signature and does not establish a universal three-band sensor.

  8. 08

    The nominal transfer result is synthetic. It nearly reaches the oracle on implanted targets, but still needs independent sensor and biological validation.

  9. 09

    The held-out family has only two measured hosts. Near-oracle transfer between E. coli and P. putida is narrow robustness, not broad chemical generalization.

  10. 010

    The detection limit is a simulator fraction. It uses an exact sensor-space target and implanted blobs, not biological concentration or field validation.

  11. 011

    Abundance intervals are conditional on detection. They cover pixels already declared as target and do not include the uncertainty of missing the target itself.