Which well-event model earns a second look?
Compare fifteen retained methods on the same causal decision grid. Inspect false alarms, missing-data abstention and individual groups before interpreting a headline score.
The selected logistic discovers 13/24 known onsets within 300 seconds; 11 come from one filename group. Full pretrained TSPulse finds no new onset at its selected alarm budget. Both TCNs fail the separate strict Mac checkpoint replay. None is promoted.
These are production-event indications after labelled onset, not pre-onset forecasting, physical well diagnosis or a controller. Public filename groups do not authenticate physical well identity; public-pretraining contamination remains unresolved.
Both unchanged TCNs exactly reproduce all 121,325 eligible scores with native batches and eight copies of only the current input. Ordinary one-reading inference fails on 46,577 / 22,294 scores; frozen alarms do not change. Extra compute and every failure remain. No model is promoted; original Mac failure and all eight unopened reserved groups remain. Inspect six full-stream profiles, eight groups and retained call costs
In 96 saved-probe controls, 40 runs reproduce all 128 original scores exactly; 56 fail the unchanged tolerance. Larger tested batches pass with MKLDNN enabled. Batch size 1 fails, so one-reading-at-a-time inference remains unqualified.
The original Mac audit stays FAILED; that 128-probe study left complete-stream qualification open. No model is promoted. Inspect all controls, source and included-compute evidence
What the input supports
Unknown labels, warmup and unusable inputs are not correct normal predictions. Full-grid and eligible denominators are retained separately.
Does the saved checkpoint reproduce?
The included Linux job completed. The separate Mac audit remains FAILED: thirteen arms pass; both TCNs fail. A reproducible validation score alone does not establish physical transfer.
Inspect the unchanged checkpointAll fifteen retained methods
Thresholds and the primary method were selected on validation. Both alarm budgets must be ≤1 false alarm per eligible normal day and ≤1% eligible normal time in alarm, with ≥0.5 eligible normal days. Disabled alarms remain visible and do not qualify a useful detector. The same frozen thresholds apply to both displayed cohorts.
| Method | Episode coverage | New onset ≤300s | False alarms/day | Normal time in alarm | Selected alarm | Mac replay |
|---|
Where the outcome comes from
Eight validation filename groups; rows with no eligible normal exposure have no false-alarm rate. They are not zero-risk evidence.
| Filename group | Coverage | New onset ≤300s | Carried-in coverage | False alarms/day | Normal time in alarm | Normal eligibility |
|---|
Group uncertainty
1,000 bootstrap replicates resample filename groups, not authenticated independent physical wells. Undefined replicates are explicit. These intervals describe selected validation data.
Recorded study boundaries
357 training recordings / 16 groups; 33 validation recordings / 8 groups; 113 reserved recordings / 8 groups.
Eleven compact fits, full pretrained TSPulse with zero pretrained fits, and three conventional references. Causal 512-second history, a 64-second reconstruction tail and decisions every 30 seconds.
No reserved measurements or predictions were opened. Scalers are training-only; alarm calibration is validation-only. This is an original causal adaptation, not an exact offline-paper replication.
Inspect every labelled episode
| Source recording | Event | Eligible decisions | Covered | Carried in | First coverage delay | New onset ≤300s |
|---|
What the 48 TCN diagnostic controls established
Two saved checkpoints × six batch sizes × two deterministic settings × two repetitions. Every condition repeats exactly on the observed Mac host. Batch size changes scores; the deterministic flag changes none. Every configuration still fails the unchanged original tolerance.
Batch 256 duplicates the 128 saved probes; it does not reconstruct the full original validation batches. No configuration was adopted. The exact kernel/architecture contribution remains unisolated. No new fits, retuning or raw/reserved access occurred.
Inspect and reproduce the evidence
Original output SHA256:
Selection SHA256:
Artifact content revision: