Evidence
oracle7-deepspace: fault management that decides onboard, maintains itself under a gate, and explains itself from its own records
The problem
Beyond the inner solar system, the ground cannot be in the loop for anything that unfolds faster than a round trip of light: . Flight fault protection handles anticipated faults with responses written before launch; when those cannot resolve an anomaly, the fallback is safe mode: stop science, keep the spacecraft safe, and wait for Earth. Safe mode is robust against the unknown, but it is a pause rather than a fix, and its cost grows with light time.
oracle7-deepspace sits between pre-written responses and safe mode, with a decision loop that fits in about a hundred kilobytes and a few milliseconds:
- Facts from telemetry, with persistence — a single-sample event never licenses a physical action.
- Hypotheses that could produce those facts.
- Admissible responses, each operator's preconditions checked against echoed state.
- Lookahead on an onboard digital twin built from telemetry (never from truth), physics perturbed ±10%: every response is simulated under every live hypothesis for 90 minutes and the best worst case is chosen.
- A receipt naming the facts, hypotheses, scored alternatives and path — which is what the craft later cites when asked why.
Two further layers: self-maintenance (rules proposed only from episodes the craft verified itself, admitted only if a validation-suite gate shows no correct response lost, no harmful response added, and less compute; versioned, ground-vetoable) and self-report (every sentence built from a telemetry frame or receipt and cited; anything no record answers is declined).
Single faults — fresh held-out seeds
Four spacecraft profiles (Voyager 1, New Horizons, Juno, Europa Clipper) × their fault classes × 50 seeds (5000–5049), never seen during development. 95% Wilson intervals. The simplified limit-to-safe-mode baseline enters safe mode on any confirmed fact and has no fault-specific responses. It is not a model of real flight fault protection, which pairs fault-specific responses (closer to the expert table) with safe mode as a fallback; it shows what safe mode alone achieves on these faults.
The expert table is an upper bound: hand-authored with full knowledge of every fault class. oracle7-deepspace comes within about one point of it with no fault table at all, from general physics lookahead — the property that matters for faults nobody anticipated. Its misses are all one case: a small solar deficit on Juno the twin cannot resolve.
Two simultaneous faults
12 fault pairs × 20 fresh seeds (6000–6019); the second fault 30 minutes after the first.
Self-maintenance gate
Ablations
Final code, fresh seeds 5000–5019 (680 scenarios each).
Twin model mismatch
Evidence persistence
One-tick persistence acts on radiation spikes and drifting sensors; two ticks or more removes every harmful action.
Physics the onboard model does not know
The twin shares the simulator's equations, so the ablation above only varies constants. Here the flight code is frozen (same SHA-256) and the spacecraft flies truth physics its twin does not model: inter-zone conduction and radiative (T⁴) heat loss, heater power ×0.8–1.25; coulombic loss, a drifting unmodelled load and current noise; sensor quantisation and heavy-tailed outliers; 3× pointing jitter; and all of these plus 2% telemetry dropout. Faults were also injected weak (just above the detection thresholds) and strong (far beyond the design range). Fresh seeds 7000–7049, run once.
Strict = the designed response on the right target. Outcome = the spacecraft came to no harm (no damage, no harmful action, battery above 30%, no zone out of band for more than 10 minutes). Outside the design range the designed response is not necessarily the right one, so the outcome is the fairer measure there.
The twin is not brittle to physics it does not model: with every term active, oracle7-deepspace and the expert table lose about the same, which points at the telemetry-to-facts layer they share. The dominant case: a healthy but under-powered heater, or a stuck-on heater whose extra heat the added losses absorb, can look exactly like a failed one. Outside the design range every strict miss on base physics was harmless (the twin declined to shed science for deficits it predicted the craft could carry). At the extreme, strong faults under every mismatch term, no system stayed unharmed in every run; oracle7-deepspace had 72 damage events against the table's 100.
Beyond what it was built on
Four more tests, written and run once on fresh seeds 8000–8049 with the flight code frozen (same SHA-256).
Spacecraft added after the freeze
On Voyager 2, OSIRIS-APEX and the Mars Reconnaissance Orbiter it was correct in all 1,300 scenarios, and the rules learned on the original four spacecraft transferred without loss. Parker Solar Probe, the only hot environment, exposes an assumption in the facts layer that the expert table shares: a zone warming above its thermostat while its heater is off is read as a stuck heater, which is false when the environment itself is hot. Parker is therefore not flown in the live simulation.
Faults it has no hypothesis for
Three of five were handled without harm. A frozen temperature sensor is never recognised (no fact tests for a reading that stops moving) and damages the spacecraft in 11% of runs; a load-current sensor reading 3.2 A high is taken for a short, a false safe mode every time. The expert table behaves identically, so both gaps are in the facts layer.
Determinism and endurance
Real spacecraft telemetry — reported honestly
The evidence layer was also run, pre-registered and unsupervised, on NASA's labelled SMAP and MSL (Curiosity) telemetry anomalies (Hundman et al., KDD 2018).
It is not competitive with the published LSTM detector. The contribution is the decision, maintenance and explanation layer, which is complementary to a learned detector: the natural next step is a learned detector feeding oracle7-deepspace facts.
Compute
Pure Python, standard library only. In a run capped at 5% of one CPU core and 64 MB (systemd cgroup), the flight logic made the same decisions as uncapped.
Reproduce
Every table on this page comes from one code version, identified by the SHA-256 of the deepspace package sources:
The simulation on this site runs the same code; its Compute tab shows the hash computed from the files it actually loaded.
| Command (Python 3, standard library only) | Produces | Result file |
|---|---|---|
| python3 eval/eval_fresh.py | Single faults, compound faults | fresh_eval.json |
| python3 eval/eval_final.py | Gate, ablations, self-report, compute | final_supporting.json |
| python3 eval/eval_structural.py | Structural mismatch, out-of-range faults | structural_eval.json |
| python3 eval/eval_generalization.py | New spacecraft, unknown faults, determinism, 60-day soak | generalization_eval.json |
| python3 eval/eval_smap_msl.py | Real SMAP/MSL telemetry | smap_msl_eval.json |
| python3 eval/eval_paper.py | Pre-registered first held-out run (pre-fix code) | paper_eval.json |
Versions. Two fixes were made after the first held-out run (a fact is marked handled only by one of its own
responses; a decision to wait expires after 30 minutes). paper_eval.json is the pre-fix run, kept because it holds the
pre-registered numbers the paper discloses (96.5% single, 92.5% compound). Its gate log therefore shows a validation-suite baseline
of 97/102 rather than the final code's 101/102; an early draft of the paper's gate table used those pre-fix numbers. The current
paper and this site use the final code throughout.
Design your own test
Every scenario above was written by the author. eval/run_scenarios.py flies a JSON file of scenarios written by
anyone (craft, physics variant, faults with onset, target and optional strength beyond the design range) against the frozen
flight code and both baselines, and stamps the report with the SHA-256 of the scenario file and of the code. An evaluator can write
a test without seeing the code; nothing is changed to run it. Example: example_scenarios.json
and its report.
Source code is available from the author on request.
Limits
- The spacecraft models are simplified power, thermal and attitude physics. The profiles decide which subsystems and faults exist; they are not engineering-fidelity models of the real vehicles.
- Fault classes are ones the simulator implements. The twin reasons from physics rather than a fault table, but novel physics is not yet tested.
- No flight-like hardware, RTOS or F Prime integration yet; the next step is a C port behind an F Prime component, and a hybrid with a learned anomaly detector.