A model wrapped in the loop that rewrites it
Requests climb the spine: parse, route, remember, speak. Around everything runs the self-improvement loop, and no change reaches the served model without passing its gate. Drag to orbit; click a layer to open it.
Record, propose, examine, serve
Record
A wrong turn becomes a structured failure record in the inbox.
Propose
A small, named repair or a new organ. w480 connected the math, relation and dialogue organs; w486 made the code organ grade its own confidence.
Examine
Unseen conversation turns and 2,369 capability checks. Lose one correct answer, or assert one wrong conclusion, and it is rejected.
Serve
The tip moves in HEAD_MANIFEST; every service resolves the same verified head.
What the gate has caught
Capability checks passed after the math, relation, dialogue and code organs were connected to the served model (w480, w482). No check lost; no new wrong assertion.
Confidently asserted wrong programs on 500 held-out coding tasks, once the code organ grades its own evidence (w486). Every program is still shown; thin ones are marked tentative and ask for one more example.
An unanswered turn. 71% of all turn time was one 122 MB file rewritten on every refusal; the gate rejected the speed fix alone because it has no latency column.
Pointer bound to an acknowledgement word: 15 of 20 test conversations before w472, 0 of 20 after. No other metric moved.
Deliberately broken rule sets submitted to oracle7-deepspace's gate. All four rejected; nine real improvements admitted in 45 seconds.
oracle7-deepspace on fresh spacecraft faults, against 29.4% for a simplified limit-to-safe-mode baseline. Zero harmful actions.
Six models, one substrate, one gate each
| Model | What it does | Measured on held-out data | Status |
|---|---|---|---|
| oracle7 | Conversation from 20.5M cited sentences, a 143,355-item textbook store and typed relations | Math 1,113 right / 191 wrong on held-out textbook exercises · code 137/500 MBPP, 1 wrong | head w492 |
| -deepspace | Spacecraft fault management, self-maintaining | 98.9% vs 29.4% baseline; 0 harmful | paper + live |
| -vision | Recognition that cites the examples it matched | MNIST 98.58% · Fashion 90.46% · CIFAR-10 62.99% | internal 2.0 |
| -trafficcop | Routes each request to what should answer it | 30/32 held-out routes | panel pending |
| -chameleon | Builds and re-skins websites | 1/4 first-pass; 0 false completions | internal 2.0 |
| -artist | Composes images from cited parts | 553 prompts × 4 seeds, running | in evaluation |
Where the core model falls short today
Measured 2026-09-25 on 600 fresh questions, before the organs were connected: most never reached retrieval, because the router did not recognise the wording. The relation, math and dialogue organs now answer some of what stopped here; this set has not yet been re-measured. Still open: most textbook word problems chain several steps, and the organ reads only single-step ones (5 of 244 held out, none wrong).