← Original confirmation audit
Public 3W research · Separately frozen foundation study

Which well-event model earns a second look?

Compare fifteen retained methods on the same causal decision grid. Inspect false alarms, missing-data abstention and individual groups before interpreting a headline score.

Validation onlyIndustrial deployment unqualified8 reserved groups unopened
The apparent winner is concentrated in one group.
The selected logistic discovers 13/24 known onsets within 300 seconds; 11 come from one filename group. Full pretrained TSPulse finds no new onset at its selected alarm budget. Both TCNs fail the separate strict Mac checkpoint replay. None is promoted.

These are production-event indications after labelled onset, not pre-onset forecasting, physical well diagnosis or a controller. Public filename groups do not authenticate physical well identity; public-pretraining contamination remains unresolved.

Complete validation replay now verifies a current-only numerical path.
Both unchanged TCNs exactly reproduce all 121,325 eligible scores with native batches and eight copies of only the current input. Ordinary one-reading inference fails on 46,577 / 22,294 scores; frozen alarms do not change. Extra compute and every failure remain. No model is promoted; original Mac failure and all eight unopened reserved groups remain. Inspect six full-stream profiles, eight groups and retained call costs
Linux replay: batch size changes the result.
In 96 saved-probe controls, 40 runs reproduce all 128 original scores exactly; 56 fail the unchanged tolerance. Larger tested batches pass with MKLDNN enabled. Batch size 1 fails, so one-reading-at-a-time inference remains unqualified.
The original Mac audit stays FAILED; that 128-probe study left complete-stream qualification open. No model is promoted. Inspect all controls, source and included-compute evidence
Loading the frozen result audit…