Adversarial & substrate waves (v8–v39)

v16 - Regime-Shift Envelope (Lucas Sweep)

In plain language

July 9, 2026. Companion to RESULTS - v16.0 Regime-Shift Envelope.

What we did

Every simulation so far used our best guesses for how people and institutions behave — wages, software budgets, how flaky sponsors get in a recession. A famous economics objection (the Lucas critique) says: the moment you change the system, the guesses change too, so test whether your conclusions survive wrong guesses. So we shook all nine of the most important guesses at once — each one anywhere from half to one-and-a-half times our estimate, in every combination — and re-ran the headline tests inside the shaken worlds.

What survived the shaking

The peacetime safety net: everything. In all 38 shaken worlds, a big sponsor quitting caused no missed groceries and no money-printing. That headline is solid no matter how wrong our guesses are (within reason).

The governance cap: everything, to machine precision. Across 400 shaken worlds — including every setting of its own dials the constitution allows — no contributor ever exceeded their voting ceiling, and buying a majority always took at least twice the required minimum of colluders. (We also learned the earlier single-pass version of the cap would have leaked in about 1 world in 13 — the fix adopted this morning genuinely matters.)

The jury system: holds right up to a sharp wall. Packing 201-person juries stays under a 1%-of-the-time nuisance as long as attackers control up to 25% of the pool — exactly the design's assumption — then breaks fast at 28%. Three points of margin. Good to know precisely where the cliff is.

What did not survive

The storm test. Our committed run — recession, two sponsors quitting, one chronically late — passed on a thin margin. Shaken, it holds in only about half the worlds. The follow-up experiments found the exact reason, and it's one dial: how much sponsors keep paying during the recession itself. (Those follow-ups are now properly saved in the results file — a July 11 fix; they'd originally been quoted from unsaved probes, and saving them properly changed the story at one point.) Keep paying 60% of their commitment: fine, in both random worlds we test. 50%: the two worlds disagree — one shows a single close call, the other shows half a year under the line. 45% or less: the poorest go under the essentials line for months — over a year in the worst world. So 50% isn't a safe floor; it's the edge of the cliff. (Notably, the system never printed money to paper over it — it failed honestly, by rationing, exactly as designed.)

One of our own predictions also flopped, usefully: we bet that a slow handover with a too-shallow deposit would cause missed groceries even in calm times. It never did — calm times absorb the gap. The deposit rule matters in storms, which is precisely why nobody should weaken it during the good years when it looks unnecessary.

The one-line takeaway

Shake every assumption at once and the design's promises hold — except one, which turns out to be a promise someone else has to keep: sponsors must contractually commit to keep paying most of their share even in a recession — 60% is safe, 50% is the edge of the cliff — or the storm math doesn't work. That's now a proposed contract term (gate-20 candidate) with the safety margin priced in, not an assumption.


Update (July 11): three more machines shaken, and the two hardest questions

Earlier the shake test could only rattle some of our machines; three of the toughest (the ones for data-batcher collusion, price-oracle capture, and multi-layer attacks) were written in a way we couldn't shake without rebuilding them first. That rebuild is now done, so we shook them too — same "half to one-and-a-half times every guess" treatment.

What held up: the batcher-collusion machine's conclusions all survived cleanly — cheating still loses money in every shaken world. The oracle machine's headline (capturing the price feed costs ~$100M and can only break things, never profit) held in every world. And the cross-layer fix — one combined cap instead of four separate ones — worked in every single shaken world. What turned out to be more fragile than we'd said: two things we already knew were close calls got confirmed as close calls (the oracle's worst-case delivery, and whether the essentials "lane" is really 5× the priciest thing to fake — sometimes yes, sometimes only 1.5×). And one new hairline: in exactly one shaken world out of 32, a captured oracle running high inflation dropped delivery a whisker below our 0.95 promise (0.9498). We're reporting it rather than rounding it away.

The recession-floor cliff, pinned down. A companion run (v31) had put the danger zone between 50% and 55%. We zoomed in with a finer ruler (1% steps) and found the real edge: the cliff is between 51% and 52%, and you need 53% before nobody misses any groceries. So the recommendation to write 60% into the sponsor contract isn't just safe — it sits a comfortable eight-or-nine points above the true edge.

The billion-dollar question: what if people won't pay for their data? Our whole revenue story rests on people and firms valuing data/software/machine-access at certain levels (~$600, ~$900, ~$240 a year). Our sharpest critic says those could be ten times too high. So we ran the "$60 world" — every one of those guesses cut to a tenth. The result is clarifying: - The safety net still holds. Nobody misses essentials, and the system still never prints money — because the floor is paid by sponsors, not by data revenue. (The cushion gets thin, though: in the storm the poorest end up right at the line, 1.008.) - Sponsoring EDEN still beats running old-fashioned welfare — that comparison doesn't care about data prices, because both sides shrink together. - The builder economy collapses. A typical builder's income falls from ~$420 to ~$45 a month; the share of builders earning a real living ($200+/mo) drops from 88% to under 3%. - The floor never pays for itself and, in the $60 world, gets even further from doing so — it leans harder on sponsors.

One-line takeaway of the update: shake the three newest machines and the safety results hold; the safety net survives even the nightmare "data is worthless" world (because sponsors, not data sales, pay for it) — but the builder economy and the dream of a self-funding floor do not. Those two ride entirely on the one experiment we haven't run yet: do people actually value their data? Everything else we can stress-test in software; that one needs the real world.

Figures

stray_fig_v15_import_PRE-FIX_archive.png

Technical results

Run: July 9, 2026. Spec: v16 SPEC - Regime-Shift Envelope (registered).md — bars L1–L4 fixed before code. Engine: lucas_sweep.py (sweeps the committed engines as committed: v6.8 run_fed imported, v15 functions imported, v11 via the closed form proven exact in theory_proofs.py P2). Committed: results_v16.json. Seeds 7 + 11 throughout. Every number traces to results_v16.json → bars / E_B / E_C / sweep_EA.

Verdict in one line: the peacetime federation headline and the governance/committee results are envelope-robust — but the storm headline is not: it FAILED the registered 75% bar (holds in only 47% of the joint ±50% envelope), and the isolation slices locate the break — storm delivery is conditional on sponsors continuing to pay a recession fraction whose breach threshold sits between 0.45 and 0.60 depending on seed (committed July 11; the two seeds disagree at exactly 0.50), so the sponsor-contract austerity floor must be priced at 0.60, or at 0.50 only with escrow sized to the residual gap.

Bar summary (L1/L3/L4 pass; L2 FAILED as a finding; the L1 bracket failed-to-fail, informatively)

Bar Registered Measured Result
L1 peacetime exit, lag ≤ 6 ≥ 90% of envelope holds delivery, zero printing 100% (38/38 cells) PASS
L1 bracket: lag > 6 (escrow fixed at 6) breaches appear (E*=lag load-bearing in peacetime) 0 breaches in 26 cells FAILED-TO-FAIL — finding F1
L2 storm, lag ≤ 6 ≥ 75% of envelope holds 47.4% (18/38) FAIL — finding F2, the headline of this run
L3 governance containment (governed κ, α ranges × population envelope) 100% contain; cartel ≥ 1/κ 100% (400 cells; worst overshoot 7×10⁻¹⁸ = machine epsilon; min cartel×κ = 2.02) PASS
L4 committee censorship k=201 ≤ 1%/epoch for attacker share ≤ 25%, all H PASS at all H; s* = 28% everywhere (0.48% at 25% → 5.5% at 28% → 16.9% at 30%) PASS + measured bracket

Findings

F1 — E* = 6 = step-up lag is a storm result, and the sweep proved its scope. The registered bracket predicted delivery breaches wherever step-up lag (3–9 swept) exceeds the committed 6-month escrow in peacetime. Zero occurred (0/26). Re-reading the committed artifacts confirms why: v6.8's J3a escrow sweep — the source of E*=lag — was a storm cell; in peacetime the surviving members' payments plus a healthy economy absorb the lag−escrow gap without the poorest noticing. Registered prediction wrong in public, per tradition; the finding sharpens the gate language: escrow depth ≥ step-up lag is a storm-conditions requirement, not a general one — which is also why it must never be relaxed on peacetime evidence.

F2 — The storm headline is dial-local; the dial is austerity. Under joint ±50% perturbation of nine dials, the registered storm (recession + two exits + chronic delayer) holds delivery in only 47.4% of the lag ≤ 6 envelope — and 18 of the 20 flip-cells breach the program's own 3-consecutive-month failure criterion (worst: 22 months below floor, p10 trough 0.77). Zero printing anywhere — the no-print rule held in all 128 cells; the failures are honest rationing. Isolation slices (all else central, storm scenario, both seeds — committed July 11 to results_v16.json → isolation; the numbers below were originally quoted from uncommitted in-session probes, Verification v4 D2) make the mechanism causal, not correlational. Sponsor recession-payment fraction, seed 7 / seed 11: 0.60 → holds both seeds (troughs 1.088 / 1.062); 0.50 → the seeds disagree — seed 7 hairline (1 month below, trough 0.994, no sustained breach), seed 11 a real breach (6 months below, max-consec 5, trough 0.982); 0.45 → 11 / 12 months below; 0.35 → 13 / 14 months (troughs 0.869 / 0.856). So the breach threshold g* sits between 0.45 and 0.60, straddling 0.50 with seed-level variance — the original "between 0.50 and 0.45" was the seed-7 story; the committed two-seed record shows 0.50 is inside the uncertainty band, not safely above it. A correlational inversion in the LHS medians (flips at higher builder ARPU) was checked by isolation and found to be a small-sample confound — isolated, higher ARPU is monotonically protective in both seeds (months below 14→13 across 450→1350 ARPU at harsh austerity on seed 7; 16→12 on seed 11; troughs improving throughout — the originally-quoted "14→11" matched neither seed and is corrected here). Design consequence (gate-20 candidate, sharpened by the committed slices): the sponsor contract's austerity clause is load-bearing — and because g* straddles 0.50, the safe floor is ≥ 0.60 of obligations, or ≥ 0.50 plus escrow explicitly sized against the residual austerity gap (not against the exit lag alone). The J4 lesson ("price the backstop") extends: price the austerity floor too, with margin.

F3 — The governance cap function is envelope-robust across its entire governed range. 400 cells over κ ∈ [0.02, 0.10], α ∈ [0.3, 0.7], σ ∈ [1, 3], whale ∈ [10–50%], N ∈ {2k, 10k, 50k}: water-fill containment held in every cell to machine precision, and the cheapest majority stayed ≥ 2/κ (double the registered bar) everywhere. The single-pass cap would have leaked in 7.75% of the governed envelope — the v0.2 water-fill mandate is not a corner-case fix; it binds in ~1 of 13 plausible configurations.

F4 — The k=201 committee has a sharp, previously-unmeasured wall at s* = 28%. Censorship stays under 1%/epoch through the design's 25% assumption at every pool size, then breaks steeply (5.5% at 28%, 16.9% at 30%). The comfort margin between the cohort-cap regime (≤15–25%) and the censorship wall is ~3 percentage points of attacker share — thin enough to justify the cross-layer joint cap (v14/gate-18) from this direction too.

Honest limits

Nine dials over two engine cell-families plus two closed-form families is the start of a Lucas answer, not the end: v12/v13-oracle/v14 engines are top-to-bottom scripts and couldn't be swept without refactor (owed — the refactor is mechanical); n=32 LHS points gives coverage, not density (the g* boundary was located by follow-up slices, not by the LHS itself); and the deepest Lucas objection — that the structure, not the dials, shifts under the regime — is answerable only by the pilot. The sweep demotes one committed conclusion (storm robustness → conditional on the austerity floor) and confirms three; no conclusion is promoted for surviving.

Run and written July 9, 2026, verification session (self-labeled Fable per the file header conventions; per the owner's record this sitting ran as Opus 4.8 — see INDEX provenance note). The L2 FAIL and the L1 failed-to-fail bracket are reported, bars unmoved, per house rules.

Postscript, July 11, 2026 (Verification v4 action 2, executed by the Fable 5 verification session): the F2 isolation slices — previously quoted from uncommitted in-session probes — are now committed (--isolate gov|arpupartial_v16_iso_*.pklresults_v16.json → isolation, both seeds). Seed 7 reproduced every quoted number; seed 11 moved the story at gov_pay 0.50 (sustained 5-consecutive-month breach), so g* is reported as (0.45, 0.60) straddling 0.50, the ARPU slice's "14→11" was corrected to the committed 14→13 (s7) / 16→12 (s11), and the gate-20 candidate's floor language was sharpened to ≥0.60-or-escrowed. Bars L1–L4 and every pre-existing JSON leaf verified unchanged.


Backlog #9 extension — July 11, 2026 (blocker #8 cleared): sweep the refactored engines, densify g*, add the FAIL-tier WTP row

#8 made v12/v13-oracle/v14 importable with byte-identical committed output, discharging this run's stated honest-limit. #9 executed the three owed pieces via new engine modes (--sweep-layers, --densify-gstar, --wtppartial_v16_layers/gstar/wtp.pkl), committed under THREE NEW top-level keys in results_v16.json (sweep_layers, g_star_densification, wtp_fail_tier_sensitivity). Every pre-existing leaf (bars/E_B/E_C/isolation/sweep_EA) is byte-identical — the file minus the three new blocks reproduces the prior results_v16.json byte-for-byte (91,581 bytes), and the assemble is deterministic (run twice → identical). Engines swept AS COMMITTED (imported run()/anchors, never reimplemented); the tested "bars" are each engine's own committed RESULTS headlines (no new bar introduced). Seeds fixed at [7,11] for the three engines (their cross-seed checks hardcode 7/11 — the #8 caveat); the behavioral/economic dials are swept, not the seed set. n=32 joint log-uniform ±50% LHS per engine.

9.1 — Layer sweep: which headlines are envelope-robust, which are dial-local

Engine · headline Committed ±50% envelope (n=32) Verdict
v12 R1 audit deters (q*≤5%) 0.62% q* ∈ [0.41%, 1.03%], 32/32 ROBUST
v12 R2 honest cost ≤0.5% 0.010% [0.005%, 0.015%], 32/32 ROBUST
v12 R3 vesting bracket (V0 profit / V168 deterred) pass V168 profit = $0 in 32/32 ROBUST
v12 R4 blast ≤0.5% annual mint 0.0063% [0.0012%, 0.019%], 32/32 ROBUST
v13 O1 capture ≥$25M $99.5M [$42.7M, $169M], 32/32 ROBUST
v13 O2 sabotage-only (ROI<0, ≥3×) 80× ratio [39×, 305×], 32/32 ROBUST
v13 O3-central delivery ≥0.95 0.960 [0.9498, 0.985], 31/32 DIAL-LOCAL (F6)
v13 O3-stress delivery ≥0.90 0.888 (baseline FAIL) [0.823, 0.967], 18/32 pass dial-local (confirms F3)
v13 O4 lane ≥5× reporter 4.3× (baseline FAIL) [1.55×, 12.0×], 7/32 pass dial-local (confirms F4)
v13 O5 single-source dominated 1/995 32/32 ROBUST
v14 X1 composition discount (joint<sum) 0.66 ratio [0.46, 0.87], 32/32 ROBUST
v14 X2 blast past 2% gate 3.74% comp [1.32%, 6.31%], 25/32 DIAL-LOCAL (F7)
v14 X3 delivery ≥1.0 (reserve drains) 1.0 / 4mo drain [1.8, 15.5]mo, 32/32 ROBUST
v14 X4 joint cap restores ≤2% 0.96% [0.48%, 1.41%], 32/32 ROBUST

No robust headline flipped. All four v12 conclusions (audit almost-free-and-deters, cheap honest cost, vesting load-bearing, bounded blast radius), v13's expensive-capture / sabotage-only / single-source-dominated conclusions, and v14's discount / delivery-protection / joint-cap-fix conclusions survive the joint ±50% shake. Two committed findings that already failed at baseline (v13 O3-stress F3, O4 F4) are confirmed dial-dependent — each passes in part of the envelope, sharpening rather than overturning them (F3's breach is driven by re-validation duration × inflation; O4's "lane is dearest" ordering holds — the lane is always ≥1.55× the reporter class — but the specific ≥5× magnitude clears in only 7/32).

New flip-findings (reported, not hidden):

F6 — v13 O3-central is not quite universal (1/32). Granted-capture central floor delivery dips to 0.9498 — a hair under the 0.95 bar — in one envelope corner: true inflation ~1.5× anchor combined with a short (0.58×) capture horizon. The margin is razor-thin (0.0002) and the mechanism is the same as F3 (inflation erodes real delivery through the quarantine); it says the ≥0.95 granted-capture guarantee should be stated as "at anchor inflation," not unconditionally.

F7 — v14's "blast composes past the 2% gate" is dial-local (25/32). In 7/32 points the composed identity blast radius stays within the 2% gate (range 1.32%–6.31%); the flips occur when the identity-layer breach fraction (v10 anchor 1.37%) and per-identity downstream leverage are both near/below their anchors (e.g., identity-fraud ×0.5, leverage-scale ×0.61). So the X2 danger finding is conditional on the identity keystone being at least as bad as v10 measured — which the honest limit already flagged as the load-bearing unknown. Crucially, the fix (X4 joint cap) restores fraud ≤2% in 32/32 regardless, so the gate-18 recommendation is unaffected by X2's dial-locality.

9.2 — g* austerity boundary densified (0.01 grid) and reconciled with v31

v31 (Austerity Floor & Escrow) already densified g* on the committed v6.8 engine to a 0.05 grid {0.45, 0.50, 0.55, 0.60}, bracketing the 3-consecutive-month breach edge to (0.50, 0.55] and recommending ≥0.60. Rather than duplicate it, #9 added resolution v31 lacks: a 0.01 grid over [0.45, 0.60], both seeds, identical committed path (v6.8 run_fed, v16 storm, escrow E=6, same build_params route as the sweep + isolation slices). The finer grid reproduces v31's 0.50 and 0.55 cells exactly (0.50: s7 mb=1/mc=1, s11 mb=6/mc=5; 0.55: both 0/0) and refines the edge: on the binding seed (11), the sustained breach persists through g=0.51 (mc=3) and clears at g=0.52 (mc=2), so g* ∈ (0.51, 0.52] — tighter than v31's (0.50, 0.55]. Zero months below the floor in both seeds requires g ≥ 0.53 (seed 7 clears at 0.51). max_consec and months_below are monotone in the floor (v31's AB1 holds on the finer grid); the only non-monotonicity is a ≤0.011 p10-trough wiggle at 0.55→0.56 on seed 11, below the breach criterion. Net: v31's ≥0.60 recommendation stands with even more visible margin — the measured sustained-breach edge is ~0.51–0.52, so a 0.60 floor sits ~8–9 hundredths above it (troughs 1.088/1.062). Committed to results_v16.json → g_star_densification (cells + per-seed boundaries + v31 reconciliation).

9.3 — A9 FAIL-tier WTP sensitivity (the program's single hinge)

The registered field experiment sets WTP FAIL bars an order of magnitude below the anchors (data ~$600/contributor-yr; code $300–1,800; machine-pay $240/agent-yr). #9 divides the three committed anchors jointly (600/900/240) by 5 ($120-world) and 10 ($60-world = the A9 FAIL tier) and runs the committed v6.8 run_fed, J1-core peace + J2-core storm, both seeds. At the $60-world:

Headline Verdict Evidence
Floor delivery (essentials floor keeps delivering) SURVIVES 0 months below floor, zero printing, peace + storm, both seeds — floor delivery is sponsor-backed, not WTP-backed. But the margin evaporates: storm p10 trough thins from 1.062 to 1.008 (razor-thin), peace payr falls 0.85→0.69.
Sponsor coverage (cheaper than traditional welfare) SURVIVES member cost-per-enrollee ÷ traditional-welfare ratio is WTP-invariant (0.785→0.785 peace; moves <0.01) — EDEN's sponsor cost and the welfare benchmark scale with the same obligation.
Floor self-funding (endogenous slice covers the floor) already weak; DEEPENS endogenous coverage (slice/obligations) is only ~0.13 peace / ~0.07 storm at anchor and roughly halves to ~0.07 / ~0.03 — the floor was never self-funding; the $60-world pushes any self-funding horizon "years to the right."
Builder viability (builders earn a living) BREAKS median builder income collapses ~10× ($420→$45); share of builders clearing the $200/mo jobs line falls 0.884 → 0.028. The builder economy thins to non-viability.

This is exactly the A9 prediction, made quantitative: at the $60-world the builder economy thins (breaks) and machine-pay shrinks (c_arpu 240→24) and the funded floor moves years to the right (self-funding deepens) — while the two sponsor-backed promises (essentials delivery and cheaper-than-welfare sponsorship) survive, because they never depended on data ARPU. The honest sharpening: EDEN's safety promises are WTP-robust; EDEN's builder-economy and self-funding promises are WTP-fragile and remain hostage to the pending field experiment. Committed to results_v16.json → wtp_fail_tier_sensitivity (12 cells + four headline verdicts).

Backlog #9 run and written July 11, 2026 (Opus 4.8, owner-directed). No pre-existing bar or JSON leaf moved; two new flip-findings (F6 O3-central, F7 X2) reported per house rules; the g* densification and WTP row are decision-support (gate-20 / red-team A9), not new registered bars. Backup at outputs/backups/v16.

Raw data

⬇ results_v16.json