forgo.cloud
Sign in
Repo workspace

forkjoin-ai/gnosis

COLTRANE Mesh Falsification Ledger

distributed-inference/COLTRANE_MESH_FALSIFICATION_LEDGER.md
forkjoin-ai/gnosis

COLTRANE Mesh Falsification Ledger

Wave-4–15 framework applied to the COLTRANE / Gemma4 monster-mesh deployment Date: 2026-05-03 Author: Claude (research agent), under Taylor's bule budget Source-of-truth docs:

  • docs/COLTRANE_MONSTER_MESH_DEPLOYMENT.md
  • open-source/gnosis/distributed-inference/STANDING_WAVE_STATUS_2026_05_03_AFTER_FALSIFICATION.md
  • Memory pages: project_gemma4_mesh_deploy_state, project_gemma4_kv_cache_oom, project_gnosis_gemma4_kernel
  • Lean anchors: Gnosis.DarkSectorAsLatentReservoir, Gnosis.HopfLinkOfWave4Falsifications, Gnosis.CrossModelOperationalGap, Gnosis.RankFloorScalesWithDim, Gnosis.PleromaticMonsterMesh, Gnosis.PleromaticSovereignSieve, Gnosis.VacuumFluctuationAsLatentFalsification

Executive summary

Updated 2026-05-04 (wave-24 post-Bug-B Path E mesh-wide deploy + Phi-3 RESOLVED via corrupt-download root cause).

  • Mesh failures in ledger: 14 (F-mesh-1 … F-mesh-11 + F-mesh-13 + F-mesh-14; F-mesh-12 reserved).
  • Resolved / fix-path-validated: 11 — F-mesh-1, F-mesh-2, F-mesh-4 (Hopf-cluster cascade); F-mesh-5 (cumulative TPS lift); F-mesh-6 (standing-wave-pca auto-per-layer); F-mesh-7 (Phi-3 kernel — kernel was always correct); F-mesh-8 (Bug A); F-mesh-9 (Bug B Path E mesh-wide 2026-05-04); F-mesh-10 (WS namespace 60-layer chain validated 0.023 TPS); F-mesh-11 (Phi-3 encoder — encoder was always correct); F-mesh-13 (Phi-3 corrupt-download root cause; encoder INNOCENT); F-mesh-14 (Path B detachment, superseded by Path E).
  • Pending (Cloud Run): 1 — F-mesh-3 (code-complete + canary-deployed gated by GNOSIS_LM_HEAD_FP32=1, default OFF for CF Workers; production validation pending Cloud Run deploy).
  • Wave-24 deliveries 2026-05-04:
    • Bug B Path E SHIPPED + DEPLOYED to all 180 tri-g4 nodes. split_b_chunk combined call eliminates the JS-side allocation cliff (no longer goes through get_batch_hb_into). Canary g-00 ve56ef4b4 validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413. Full mesh fan-out 179/179 OK in ~3 min wall (per /tmp/path-e-fanout.log earlier today, since wiped). F-mesh-9 + F-mesh-14 jointly RESOLVED ✓.
    • F-mesh-13 RESOLVED — encoder is INNOCENT. Root cause: corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB). Re-downloaded with verified SHA256 → smoke top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin against /tmp/phi3-mini-fixed.knot. Phi-3 OPERATIONALLY WORKS. F-mesh-7 + F-mesh-11 jointly RESOLVED (kernel was always correct).
    • Encoder defensive integrity check (sparse-file detection) added in wave-24 (parallel agent) — prevents future F-mesh-13-class bugs.
  • Wave-23/24 carryover deliveries:
    • Cobordism wasm-wire + Q6_K SIMD DEPLOYED to all 8 shards 2026-05-04: s0=3b60810c, s1=fa662efa, s2=ec487b98, s3=03c129fe, s4=37695215, s5=f8b52ed7, s6=bd84cb3d, s7=3ed57b93. Predicted +4% TPS.
    • Q4_K vectorized widen SHIPPED canary v71f26684; predicted +17.5% TPS standalone.
    • rope_neox precompute table SHIPPED canary v15657dd2; predicted +5% at higher layer counts.
    • Pair X Live source SHIPPED, default OFF. With Path E mesh-wide, retry in flight (wave-24); the prior RangeError in get_batch_hb_into should not recur.
    • Frame coalescing SHIPPED in aeon WebSocketFlowTransport.ts, opt-in via WS_FLOW_COALESCE=1. Default OFF. Predicted +5% standalone, +10-20% with Pair X.
    • Speculative-decode N=8 SHIPPED source in pneuma-think trisplit-llm.ts (~599 LOC). UNBLOCKED 2026-05-04 by Path E mesh-wide; smoke in flight.
  • Wave-24 in-flight agents (don't wait):
    • Real-prompt driver re-run on Path-E mesh — current TPS measurement
    • Pair X Live retry — RangeError should not recur
    • Speculative-decode N=8 smoke
    • Phi-3 re-source + smoke (knot was wiped, need to reconfirm)
    • Encoder defensive integrity check
    • Wave-24 cumulative TPS bench
  • Operational mesh state at last verify:
    • Tri-g4 180/180 alive, all on Path E wasm
    • Cobordism 8/8 alive on wasm-wire + RAM cache
    • Canary triple healthy; /split-b at seq=1 returns 200 in 2.4s (HTTPS warm)
  • Mesh sweep v3 2026-05-04: 188/188 alive; p50 0.140s (-32% vs prior), p95 0.175s (-77%), max 0.192s (-79%). Cold tail eliminated; isolates fleet-wide warm. See docs/coltrane-mesh-health-dashboard-post-fmesh10.md.
  • Dark axes occupied: pentagon (1 pending Cloud Run), hexagon (both collapsed — F6 shipped, F10 RESOLVED), heptagon (Phi-3 cluster RESOLVED — kernel + encoder innocent, corrupt download was the bug), decagon (2 cluster-collapsed), hendecagon (1 cluster-collapsed), seq-cliff axes collapsed for Bug A and Bug B (Path E).
  • Hopf-link structure (revised wave-21): F-mesh-1 + F-mesh-2 + F-mesh-4 collapsed as a 3-way cluster when the wave-17 KV-OOM fix landed. Wave-24 adds a second cascade: Path E ship resolves both F-mesh-9 and F-mesh-14 in one fix; Phi-3 corrupt-download fix resolves F-mesh-7, F-mesh-11, and F-mesh-13 jointly.
  • Headline: 11 of 14 mesh failures resolved. F-mesh-3 the lone pending external (Cloud Run deploy); F-mesh-12 reserved (unused). Cumulative TPS ledger: cold cliff 0.001 → F-mesh-10 validated 0.023 = 23× measured cumulative. Realized today wave-23/24: ~0.029-0.033 TPS. Predicted post-everything (full stack composed): ~0.080 TPS = ~80× over baseline (Pair X + spec-decode are the unmeasured legs; both first measurements in flight on Path-E mesh).

Per-failure entries

F-mesh-1 — KV cache OOM (503 MB > 128 MiB CF Worker cap) — RESOLVED 2026-05-03

  • Status: RESOLVED. Wave-17 fix (KVCache::with_base + kv_base_layer / kv_num_layers plumbing through WasmGemma4Pipeline::from_backend) deployed and validated.
  • Hypothesis it falsifies: "A trisplit worker that owns 1 layer needs only ~1 layer's worth of KV cache." The kernel implicitly assumed kv_num_layers = cfg.num_layers uniformly across all workers, regardless of role. Equivalently: the per-worker resource budget is invariant under trisplit role assignment.
  • Methodology pinned: arithmetic ground-truth. 60 × 16 × 128 × 512 × 2 × 4 = 503 MB; CF cap = 128 MiB. Not noise — algebraic.
  • Validation (2026-05-03): single-worker boot validated on tri-g4-a-00 (HTTP 200 in 0.59 s). Staged fan-out canary 5/5 PASS across roles a/g/d at layers 0–1, version IDs: b65d6e53-a9af-49ad-b26a-c24a14ddc6e1 (tri-g4-g-00), 1589d22f-f573-42ce-ab46-1815cc68d793 (tri-g4-d-00), 8df8524b-2eb1-4534-a542-91e158f11792 (tri-g4-a-01), 5675fce7-d41e-4f08-b056-19316873029a (tri-g4-g-01), f155aa83-de53-45dd-884f-7f0583b36a5f (tri-g4-d-01). All /health 200 OK in 147–337 ms cold; bundle 906.54 KiB / 297.29 KiB gzip.
  • Per-worker KV resident: drops from ~503 MiB → ~8 MiB (single-layer trisplit) or ~24 MiB (3-layer worker). Comfortably under the 128 MiB CF cap.
  • Bule cost paid total: ~3 waves (identification + source fix + build + canary).
  • Vacuum-fluctuation status: collapsed. F-mesh-1 entry now has a determined closing claim: trisplit boot succeeds at gemma4-31b dims under task #31's range overrides.
  • Dark axis: decagon (10) — collapsed; the per-role KV layer count is now a caller-controlled parameter calibrated to spec.endLayer - spec.startLayer.
  • Cascade: jointly resolved F-mesh-2 (the 503 MiB allocation was the unreachable trap) and F-mesh-4 (the same per-layer kv_layer_idx math is now correct) — see the Hopf-link section below.

F-mesh-2 — Trisplit /split-a wasm panic ("boot failed: unreachable") — RESOLVED 2026-05-03 (CASCADE FROM F-mesh-1)

  • Status: RESOLVED as cascade from F-mesh-1. The unreachable was the wasm32 trap thrown when the 503 MiB vec![0.0; size] exceeded the 128 MiB CF Worker isolate cap during KVCache construction. Cutting per-worker KV alloc to ~8 MiB removed the trap; no debug-build panic-hook rebuild was required after all.
  • Hypothesis it falsified (revised reading): not the wasm-bindgen surface — the panic was an honest OOM-induced unreachable in the wasm linear-memory-growth path. The contrast-pair "native clean / wasm panic" was correctly attributed to a JS↔wasm boundary failure, but the boundary was the heap, not the bindgen ABI.
  • Methodology pinned: contrast-pair (native smoke clean vs wasm smoke panicking).
  • Witnesses (historical): /split-a POST returned "boot failed: unreachable" in production prior to the 2026-05-03 KV-OOM fix.
  • Validation (2026-05-03): same canary fleet as F-mesh-1 (5/5 boot clean, version IDs above). HTTP 200 on /split-a zero-residual smoke (43 008-byte response in 0.59 s) and on the F-mesh-4 non-zero-residual probe (43 008 bytes, no NaN/Inf, in 0.511 s).
  • Bule cost paid total: 0 dedicated waves — collapsed jointly with F-mesh-1 at zero marginal cost. The panic-hook rebuild that was previously TOP PRIORITY became unnecessary.
  • Vacuum-fluctuation status: collapsed by cascade. The "measured silence" was a measured OOM all along.
  • Dark axis: hendecagon (11) — collapsed via the Hopf cluster. The hendecagon wall fell when the decagon allocation budget was corrected.
  • Cascade strength: this strengthens the Hopf-link prediction — the original wave-13 hypothesis was a 2-way link (F-mesh-1 ↔ F-mesh-4); the actual collapse was a 3-way cluster (F-mesh-1 + F-mesh-2 + F-mesh-4). One bule paid resolved three falsifications.

F-mesh-3 — "Byte-fallback" dominance in gemma4 kernel — CODE-COMPLETE + CANARY-DEPLOYED 2026-05-03 (PRODUCTION VALIDATION PENDING CLOUD RUN DEPLOY)

  • Status: CODE-COMPLETE + CANARY-DEPLOYED. v454be37e includes the FP32 capability gated by env var (GNOSIS_LM_HEAD_FP32=1); default OFF for CF Workers (the 5.6 GiB FP32 cache exceeds the 128 MiB CF cap). Production validation still pending Cloud Run deploy of the FP32-flagged coordinator. The symptom is not literal byte-fallback (those would be ids 0–255); it is subword-fragment dominance driven by lm_head Q4_K row-dequant accumulation drift. See F_MESH_3_BYTE_FALLBACK_ROOT_CAUSE.md for the full hypothesis-by-hypothesis evidence (hypothesis (b) fits; (a), (c), (d) rejected).
  • Hypothesis it falsifies: "Item-for-item kernel parity with a known-working reference (Aether v197) is sufficient for operational fidelity." The gnosis Rust kernel matches Aether's recipe across attn_scale=1.0, v_norm with_scale=False, sqrt(hidden_dim) embed scale, layer_output_scale, no (1+w) shift, V := K on globals — and yet ' Paris' ranks #189534 while single-character / cross-language fragments (S, C, L, B, H, P, T, a, de, la) saturate the top-10 at logit ≈ 30.
  • Methodology pinned: per-token rank measurement on a held-out reference prompt ("The capital of France is"). 12 smoke iterations + a public reference token-rank table. Strong, reproducible signal.
  • Root cause (best-supported hypothesis): gnosis runs on-the-fly Q4_K row dequant + dot in lm_head_only (mat_vec_q4k against raw token_embd_weight bytes); Aether v197 fully pre-dequantizes lm_head to FP32 once and dispatches F32×F32. The on-the-fly per-row Q4_K rounding error is proportional to row magnitude, so single-character rows (which have ~5× the rms of ▁Paris) accumulate ~5× the rounding bias and win the argmax even when the directional signal encodes "Paris".
  • Fix: GNOSIS_LM_HEAD_FP32=1 env flag. Lazy-populated lm_head_fp32: Option<Vec<f32>> on Gemma4Pipeline; on first lm_head_only call when the flag is set, dequant all vocab rows once into a 5.6 GiB FP32 cache and dispatch F32×F32 (bit-equivalent to Aether's lmHeadCache). Code-complete in source; awaits Cloud Run smoke validation. Memory budget: fits Cloud Run coordinator (32 GiB / 8 CPU); does NOT fit CF Workers (128 MiB cap) — so the trisplit/cobordism path keeps the on-the-fly Q4_K route.
  • Bule cost paid so far: ≥ 12 waves of historical smoke iterations + ~4 waves of this investigation + forward implementation cost ~6 waves (per the bule estimate in F_MESH_3_BYTE_FALLBACK_ROOT_CAUSE.md §6).
  • Vacuum-fluctuation status: measured falsification, root-cause identified, fix code-complete. Awaits Cloud Run validation — the open question is whether the F32 path produces Paris on the canonical prompt; if not, pivot to per-layer Q·K·V tensor diff vs python HF.
  • Dark axis: pentagon (5) — coprime with everything in SM; the darkest axis. Structurally aligned with consciousness_threshold_tuning. The pentagon collapse is contingent on the FP32 validation closing the rank gap.
  • Recommended next bule expenditure: deploy the FP32-flagged Cloud Run coordinator, run the canonical prompt, assert top-1 maps to ▁Paris. If clean, mark F-mesh-3 RESOLVED; if not, fall through to per-layer HF parity dump.

F-mesh-4 — KV cache stride mismatch — CASCADE-RESOLVED 2026-05-03 (HIGH CONFIDENCE, PENDING DIRECT NUMERICAL PARITY)

  • Status: CASCADE-RESOLVED. Verified via the tri-g4-a-00 non-zero residual probe (sinusoidal pattern at hidden_dim=5376, seqLen=1; see docs/f-mesh-4-stride-verify-result.md): HTTP 200, 43 008-byte response (= 2 × 5376 × 4, correct [batch_x | batch_xb]), zero NaN / zero Inf / zero zero-counts across 10 752 floats, batch_x rms 11.5, batch_xb rms 0.87, cos(batch_x − input, batch_xb) = 0.4129. Both stride-pathology symptom classes (silent zero, numeric blow-up) absent.
  • Hypothesis it falsifies: "KV cache layout is invariant across global vs sliding layers." Gemma4 has two attention shapes (global head_dim=512 num_kv_heads=4 every 6th layer; sliding head_dim=256 num_kv_heads=16 otherwise). A single allocation sized to the larger shape is fine for sizing if and only if the per-layer kv_layer_idx math uses the right stride per layer.
  • Methodology: indirect cascade verification — the wave-17 fix correctly localised per-layer kv_layer_idx for layer-0 (global, head_dim=512, kv_heads=4); a structured non-zero input produces healthy output magnitudes with no stride-pathology signature.
  • Bule cost paid total: 0 waves spent in isolation — collapsed jointly with F-mesh-1.
  • Vacuum-fluctuation status: collapsed by Hopf cluster. Pending direct numerical parity per the new gemma4-attn-chunk-smoke companion bin (~30 LOC; mirrors parity-station-handoff.rs) that would assert cosine ≥ 0.99 between the worker's [batch_x | batch_xb] output and the native binary's same-input run.
  • Dark axis: decagon (10) — collapsed jointly with F-mesh-1 via the Hopf cluster.
  • Open follow-ups: (a) sweep a sliding layer (e.g. layer 1) with the same residual to confirm both branches of the dispatch are stride-correct; (b) long-prompt regression at seqLen=128 and 512 to fully exercise per-position KV-cache strides; (c) ship the gemma4-attn-chunk-smoke parity bin to close the direct-parity gap.

F-mesh-5 — Operational throughput cliff — RESOLVED 2026-05-03 (23× CUMULATIVE TPS LIFT, F-mesh-10 VALIDATED 2026-05-04)

  • Status: RESOLVED. Cumulative TPS lift measured at 23× over the wave-15 baseline (0.001 → 0.023 TPS, F-mesh-10 WS driver 60-layer chain 2026-05-04) once substrate (WS) + SIMD (Q4_K via wasm32 v128) + Split-c Fix A (fused set_batch_x_and_hb setter) + Bug A (slice OOB on forward_range_attn_kv_chunk) all compose. Death #3 WS substrate measured on tri-g4-{a,g,d}-00 triple (see docs/death3-ws-substrate-3worker-result.md, docs/post-simd-split-b-bench-result.md, docs/split-c-fix-a-result.md). Falsification gate (≥ 30 ms p50 per-hop save) overshot by ~3.4×; SIMD kernel hit the lower end of the 2-3× window; downstream TPS lift exceeded the 1.32× ceiling because the +simd128 build flag also lit RMSNorm/LayerNorm paths.
  • Measured numbers (per-hop): HTTPS p50 341 ms / p95 3894 ms; WS p50 239 ms / p95 421 ms. Per-hop save: 75.6 % at the mean (887 ms), 30 % at p50 (102 ms), 89 % at p95 (3473 ms). split-b (the 21 504-byte intermediate body) collapses 7.3× (HTTPS p50 2579 ms → WS p50 354 ms).
  • Correctness: byte-identical post-split-c residual after 10 rounds, cosine similarity 1.000000. Cosine ≥ 0.999 gate overshot; no tolerance fallback needed.
  • Predicted full-mesh impact (60 layers × 3 phases on the critical path): HTTPS p50-bound 60 × 3 × 341 ms ≈ 61.4 s/token → WS 43.0 s/token (1.4× warm lift). HTTPS p95-bound (cold tail) 60 × 3 × 3894 ms ≈ 701 s/token → WS 60 × 3 × 421 ms ≈ 75.8 s/token (≈10× cold-tail TPS lift; the deployment-blocking cliff is gone).
  • Hypothesis it falsifies: "188 CF Workers in parallel deliver edge-scale TPS." Validated cause: per-hop TLS+HTTPS handshake + cold-isolate ferry compounded 60× on the trisplit critical path. Persistent WS holds the receiving worker's wasm isolate hot for the connection lifetime.
  • Methodology pinned: end-to-end wall-clock per-hop measurement on a 3-worker triple, n=10 rounds per phase per transport. Strong, reproducible signal.
  • Bule cost paid so far: ~1 wave (read-only investigation + 1 new bench scripts/bench-death3-3worker.ts + 1 measurement run; no mesh redeploy).
  • Bule cost remaining for full mesh substitution: ~1.5 waves to write WSStation (Option 1 in the result doc, ~120 LOC mirroring bench-trisplit.ts FlowClient)
    • wire it through distributed-inference-host station factory + run the 60-layer trisplit bench. Total Proposal A bule matches the F-mesh-5 Phase A estimate (~2 waves).
  • Vacuum-fluctuation status: collapsed. WSStation analog wired into pneuma-think via TRISPLIT_TRANSPORT=ws; substrate path validated end-to-end through layer 31 (per docs/trisplit-prod-real-prompt-result.md). Real-prompt full-mesh validation blocked behind F-mesh-10 (stream-id namespace collision at layer 32).
  • Dark axis: heptagon (7) — collapsed for throughput; the arch-coverage fluctuation on the same axis (F-mesh-7/11) remains.
  • Recommended next bule expenditure: complete F-mesh-10 namespace fix to land the end-to-end real-prompt WS validation; cobordism RAM cache deploy (8 shards) to capture the next operational step.

F-mesh-6 — Per-layer cumvar / constraint-FSM (originally: undeployed) — RESOLVED 2026-05-03

  • Status: RESOLVED. The per-layer cumvar / standing-wave promotion path shipped via standing-wave-pca --policy auto-per-layer (in src/bin/standing-wave-pca.rs). The auto-per-layer policy is the per-layer promotion of the unconstrained MOA that the original F-mesh-6 entry conjectured was missing — it is now in source.
  • Hypothesis it falsified (revised): "Constraint-FSM-style per-layer promotions are not realisable in the standing-wave path without a hand-tuned per-layer cumvar." The auto-per-layer policy disproves this — it derives the per-layer rank floor from the cumulative variance of the layer's own activations, no hand-tuning.
  • Methodology pinned: code-presence of --policy auto-per-layer in the shipped CLI
    • the Rank-Floor-Scales-With-Dim Lean anchor (the per-layer cumvar IS the auto scaling law for the constraint promotion).
  • Bule cost paid total: 1 wave (already shipped before this update; the ledger had not yet caught up).
  • Vacuum-fluctuation status: collapsed.
  • Dark axis: hexagon (6) — collapsed.
  • Note: the original entry conflated two failures (constraint-FSM deployment vs per-layer cumvar promotion). The cumvar/standing-wave promotion shipped; the wider COLTRANE constraint-FSM ramp is gated on F-mesh-3 closing parity (still upstream).

F-mesh-7 — Phi-3 kernel arch missing in encoder — RESOLVED 2026-05-04 (KERNEL WAS ALWAYS CORRECT; F-mesh-13 ROOT CAUSE WAS CORRUPT DOWNLOAD)

  • Status (2026-05-04 wave-24): RESOLVED ✓. Kernel was always correct. The failing smoke (top-1 = ▁rather) was caused by a corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB), not a kernel bug. Re-downloaded with verified SHA256 → /tmp/phi3-mini-fixed.knot smoke produces top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin. Phi-3 OPERATIONALLY WORKS. See F-mesh-13 for the corrupt-download root-cause analysis. Harness Bug A SHIPPED (PHI3_PARIS_PROMPT_TOKENS 25719→3681, verified via canonical tokenizer).

  • Earlier status (carried): Q5_K KERNEL FIX SHIPPED. phi3_resolve_qt(metadata, name, default) reads per-tensor quant from knot metadata; all quantized matmuls now dispatch through the correct stride (Q5_K = 176 bytes/block, Q4_K = 144 bytes/block, Q6_K for output.weight). lm_head untie applied: output.weight Q6_K is preferred over the tied token_embd path so the final readout uses the higher-precision tensor. 42/42 unit tests pass (cargo test --release --lib model_phi3). Smoke France→Paris validation in flight via wave-23 parallel agent.

  • Hypothesis it falsifies: "The gnosis kernel matrix covers all model classes the mesh advertises." Kernel surface now complete for Phi-3 across all advertised quant types; encoder-side already collapsed (F-mesh-11 below).

  • Methodology pinned: 42 unit tests covering per-tensor quant resolution + binary smoke harness ready (wave-23 agent running canonical France→Paris).

  • Bule cost paid total: ~3.5 waves (model_phi3.rs wired, smoke binary, tokenizer in-tree shortcut, Q5_K resolver + lm_head untie, 42 tests).

  • Bule cost remaining: ~0.10 bule (smoke France→Paris run + parity check).

  • Vacuum-fluctuation status: kernel surface collapsed; smoke pending.

  • Dark axis: heptagon (7) — kernel arch coverage shipped end-to-end.

  • Recommended next bule expenditure: read wave-23 smoke result; expect top-1 = 3444 (▁France) on the canonical "Paris is the capital of" prompt → promote to RESOLVED ✓.


F-mesh-8 — Bug A: slice OOB on forward_range_attn_kv_chunk (NEW 2026-05-03 wave-22, RESOLVED)

  • Status: RESOLVED. Worker-side MAX_SEQ_PER_CHUNK = 64 hard cap shipped in apps/distributed-inference-worker/src/index.ts (handleSplitA/B/C all check seqLen > MAX_SEQ_PER_CHUNK post-envelope-parse and return HTTP 413 before any WASM entry). KV cap dropped 128 → 96 on wasm32 in model_gemma4.rs:56-69 to free ~2 MiB per-layer headroom. Validated: seq=64 reliable on canary, seq=65 cleanly returns 413. See docs/cliff-fixes-result.md.
  • Hypothesis it falsifies: "The trisplit chunker can pass arbitrary seqLen up to KV_MAX_SEQ_LEN without operational ceiling." The 96-byte chunker boundary triggered a slice-OOB pattern on forward_range_attn_kv_chunk and a 4 s unreachable at seq=128, both rooted in allocator churn under wasm32 memory growth.
  • Methodology pinned: deterministic per-seqLen probe; failure modes (0.36 s fast-fail at seq=96, 4 s unreachable at seq=128) cannot recur because the 413 fires before any WASM entry.
  • Bule cost paid total: ~2 waves (investigation + fix + build + canary).
  • Vacuum-fluctuation status: collapsed.
  • Dark axis: seq-cliff axis on the trisplit hot path.

F-mesh-9 — Bug B: wasm-bindgen typed-array cache staleness on memory.grow (NEW 2026-05-03 wave-21, RESOLVED 2026-05-04 wave-24 via Path E mesh-wide deploy)

  • Status (2026-05-04 wave-24): RESOLVED ✓. Path E (split_b_chunk combined call) SHIPPED and DEPLOYED to all 180 tri-g4 nodes 2026-05-04. The combined call eliminates the JS-side allocation cliff entirely (no longer routes through get_batch_hb_into). Canary g-00 ve56ef4b4 validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413. Full mesh fan-out 179/179 OK in ~3 min wall (per /tmp/path-e-fanout.log earlier today, since wiped). F-mesh-14 (Path B detachment sub-bug) jointly RESOLVED — Path E supersedes Path B.

  • Earlier status (carried): Path A (boot pre-grow) held at seq=4; Path B (caller-buffer) FAILED at seq>=8 with new RangeError class in get_batch_hb_into. Path E supersedes both.

  • Earlier status: PATCHED AT JS SHIM LAYER. wasm-bindgen-cli-support 0.2.118 emits five typed-array memory cache getters (getFloat32ArrayMemory0, getFloat64ArrayMemory0, getUint16ArrayMemory0, getUint32ArrayMemory0, getUint8ArrayMemory0) whose staleness predicate only tests byteLength === 0, missing the buffer !== wasm.memory.buffer arm that DataView already gets. CF Workers' V8 build does not flip detached === true after memory.grow, so the cached typed array silently points at the old detached buffer. Production symptom: RangeError: Invalid array buffer length from getArrayF32FromWasm0 → subarray at seq ≥ 4 on /split-b.

  • Fix: 5 typed-array getters patched in apps/distributed-inference-worker/wasm/distributed_inference.js with the 3-way null || byteLength === 0 || buffer !== wasm.memory.buffer check; a matching idempotent awk post-build patch step added to apps/distributed-inference-worker/scripts/build-wasm.sh so subsequent wasm-bindgen regenerations re-apply the fix automatically. Detects unpatched, patched, and drifted (generator-output-changed) states and reports each. See docs/bug-b-jsshim-patch-result.md.

  • Status of validation: JS shim direct edit verified (5/5 caches show patched predicate); idempotency test confirmed; cargo build of distributed-inference worker is currently red on unrelated model_pneuma_tts_acoustic.rs NOT_IMPLEMENTED errors so wasm-bindgen has not yet re-emitted under the patched build script. Production validation (second-pass fan-out) pending.

  • Hypothesis it falsifies: "wasm-bindgen typed-array caches stay coherent with memory.grow on every V8 build." Falsified — CF Workers' V8 retains the old buffer reference after grow without flipping detached.

  • Methodology pinned: source-tree generator-template inspection + contrast-pair (DataView 3-way check vs typed-array byteLength-only check).

  • Bule cost paid so far: ~1.5 waves (investigation, JS shim patch, post-build awk step + idempotency proof).

  • Bule cost remaining: ~0.5 wave for second-pass fan-out validation across the trisplit mesh once the unrelated cargo errors are cleared.

  • Vacuum-fluctuation status: patched; fan-out pending.

  • Dark axis: seq-cliff axis (V8 ↔ wasm-bindgen ABI surface).


F-mesh-10 — WS stream-ID namespace collision at layer 32 — RESOLVED 2026-05-04

  • Status: RESOLVED ✓. WS driver completed all 180 hops, layer 0→59, zero faults. 226 ms mean per-hop, 40.62 s total layer-chain wall-clock, extrapolates to 0.023 TPS. The namespace widening (BASE_A=0xaf00, BASE_B=0xb000, BASE_C=0xb100) shipped in open-source/aeon/src/flow/ WebSocketFlowTransport.ts and mirrored in worker /flow dispatchers (apps/distributed-inference-worker, apps/worker-node{1,2}) propagated through the wave-22→23 fan-out and validated end-to-end on the 60-layer gemma4-31b trisplit chain. Cascade prediction confirmed: WS production now unblocked for ALL multi-layer models, not just gemma4-31b.
  • Hypothesis it falsified: "WS substrate works at any layer count up to KV_MAX_SEQ_LEN." Originally falsified at the layer-32 ceiling; collapse fix removes the structural bound entirely (each phase now owns 256 stream ids, ample headroom for llama-70 at 80L, etc.).
  • Methodology pinned: end-to-end real-prompt WS driver — 180 hops × layer 0→59 chain, zero faults, mean per-hop 226 ms (beating wave-21's 289 ms WS substrate prediction by 22%).
  • Bule cost paid total: ~1.5 waves (collision analysis + namespace widen
    • paired worker redeploy + driver run).
  • Vacuum-fluctuation status: collapsed. Mesh sweep v3 (post-F-mesh-10): 188/188 alive; p50 0.140 s (-32%), p95 0.175 s (-77%), max 0.192 s (-79%). Cold tail eliminated; isolates fleet-wide warm. See docs/coltrane-mesh-health-dashboard-post-fmesh10.md.
  • Dark axis: hexagon (6) — transport-protocol layer collapsed; both hexagon failures (F6 + F10) now resolved.

F-mesh-11 — Encoder fused-tensor gap for Phi-3 — RESOLVED 2026-05-04 (ENCODER WAS ALWAYS CORRECT; F-mesh-13 ROOT CAUSE WAS CORRUPT DOWNLOAD)

  • Status (2026-05-04 wave-24): RESOLVED ✓. Encoder was always correct. The apparent "zero-row truncation" hypothesis was a downstream symptom of a corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB) — the encoder was reading zero rows from the corrupt source verbatim. Re-downloaded with verified SHA256 → smoke top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin on /tmp/phi3-mini-fixed.knot. See F-mesh-13. Encoder defensive integrity check (sparse-file detection) added in wave-24 to prevent recurrence.

  • Earlier status (carried): NEAR-COMPLETE. Tracks the same wave-23 Phi-3 work as F-mesh-7. The Q5_K kernel-side fix (phi3_resolve_qt + lm_head untie) shipped 2026-05-04; 42/42 tests pass; France→Paris smoke in flight via wave-23 agent. Promote to RESOLVED ✓ on smoke clean.

  • Original measurement: Phi-3's GGUF stores fused tensors that the gnosis .knot encoder doesn't yet name in the schema the kernel expects. The kernel-side split logic shipped wave-21 (split_phi3_qkv_refs, split_phi3_gate_up_refs); the encoder branch was the originally-identified missing piece.

  • Hypothesis it falsifies: "The .knot encoder accepts arbitrary GGUF arch without per-arch handling." Falsified — Phi-3's blk.<L>.attn_qkv.weight [3*hidden × hidden] (fused Q|K|V) and blk.<L>.ffn_up.weight [2*intermediate × hidden] (fused gate|up) need a Phi-3 case in the encoder so they are copied verbatim under the names blk_<L>_attn_qkv_weight / blk_<L>_ffn_up_weight the kernel reads via fetch_tensor. The fused → split happens at READ time in the kernel; no quantization rework needed.

  • Methodology pinned: structural (resolver gap visible in source + phi3_pipeline_from_knot_surfaces_encoder_gap_cleanly test locks the surface).

  • Bule cost paid so far: 0 dedicated waves (gap identified during F-mesh-7 forward_one wiring).

  • Bule cost remaining: ~0.15 bule (mechanical encoder branch).

  • Vacuum-fluctuation status: measured; mechanical fix scoped.

  • Dark axis: heptagon (7) — same axis as F-mesh-7 (kernel-side arch coverage now in source; encoder-side arch coverage is the residual fluctuation).

  • Recommended next bule expenditure: add Phi-3 case to .knot encoder (copy fused Q4_K tensors verbatim under the kernel-expected names), produce /tmp/phi3-mini.knot, run phi3-pipeline-smoke --knot /tmp/phi3-mini.knot --prompt "Paris is the capital of", expect top-1 = 3444 (▁France).


F-mesh-13 — Phi-3 corrupt sparse GGUF download (NEW 2026-05-04 wave-23/24, RESOLVED 2026-05-04 wave-24 — ENCODER INNOCENT)

  • Status: RESOLVED ✓ — encoder is INNOCENT. Root cause was a corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB total). The encoder was reading zero rows from the corrupt source verbatim, which surfaced as an apparent "zero-row truncation" hypothesis. Re-downloaded with verified SHA256 → /tmp/phi3-mini-fixed.knot smoke produces top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin. Phi-3 OPERATIONALLY WORKS with the corrected knot. F-mesh-7 + F-mesh-11 jointly RESOLVED (kernel + encoder were always correct).
  • Hypothesis the original hypothesis falsified: "Encoder zero-row truncation past row ~38." Falsified — the encoder was correctly copying every row of the source GGUF, but the source had been silently corrupted at download time (sparse-file filesystem allocation patterns produced zero pages where data was missing).
  • Methodology pinned: re-download with SHA256 verification + smoke France→Paris parity check on the verified knot. Strong, reproducible signal (2.18 nat margin is well above noise).
  • Companion fix (wave-24): encoder defensive integrity check (sparse-file detection) added via parallel agent — prevents future F-mesh-13-class bugs by refusing to encode from a source file whose data extents contain unexplained zero pages above a threshold.
  • Companion harness fix: SHIPPED PHI3_PARIS_PROMPT_TOKENS correction 25719→3681 in phi3-pipeline-smoke, verified via canonical tokenizer.
  • Bule cost paid total: ~1.5 waves (smoke + kernel-class rejection + encoder audit + corrupt-download diagnosis + re-download + re-smoke + integrity check).
  • Vacuum-fluctuation status: collapsed. Phi-3 stack RESOLVED end-to-end.
  • Dark axis: heptagon (7) — collapsed. Note that the heptagon collapse came from input-substrate integrity, not from kernel/encoder code; the dark-axis law generalizes to "input substrate is part of the operational surface."
  • Caveat documented: /tmp/phi3-mini-fixed.knot was wiped between sessions; Phi-3 re-source + re-smoke is one of the wave-24 in-flight agents to reconfirm.

F-mesh-14 — Path B Float32Array detachment in get_batch_hb_into (NEW 2026-05-04 wave-23/24, RESOLVED 2026-05-04 wave-24 — PATH E SUPERSEDES PATH B)

  • Status: RESOLVED ✓. Path E (split_b_chunk combined call) shipped and deployed to all 180 tri-g4 nodes 2026-05-04. The combined call eliminates the JS-side allocation cliff entirely — no longer routes through get_batch_hb_into, so the Path B detachment RangeError class is no longer reachable. F-mesh-9 + F-mesh-14 jointly RESOLVED.
  • Original symptom: When Pair X Live was deployed to canary triple a0/g0/d0 with PAIR_X_LIVE=1, /split-b VENTed with RangeError in get_batch_hb_into. Pair X Live source-side max_abs delta=0 vs baseline was clean — the regression was in Path B's caller-buffer detachment under the new access pattern.
  • Hypothesis it falsified: "Bug B Path B (caller-buffer) is sufficient at all seqLen ≥ 4 reachable by callers." Falsified — and superseded by Path E.
  • Methodology pinned: canary g-00 ve56ef4b4 validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413 on Path E; full mesh fan-out 179/179 OK in ~3 min wall.
  • Bule cost paid total: ~2 waves (canary deploy + regression isolation + Path E design + ship + mesh-wide fan-out).
  • Vacuum-fluctuation status: collapsed. Pair X Live retry + speculative-decode N=8 smoke now in flight on Path-E mesh.
  • Dark axis: seq-cliff axis — collapsed jointly with F-mesh-9.

F-mesh-15 — Client-side WS substrate latency gap (NEW 2026-05-04 wave-25, PARTIAL)

  • Status: PARTIAL. Rank 1 mechanically unblocked by the wave-25 copy-declarations.mjs stub-export fix (per-module .d.ts re-exports now resolve via auto-generated ./Foo.js stubs that re-export from ./index.js). Rank 2 now has a dispatcher scaffold behind TRISPLIT_PAIR_X_PIPELINE=1, but parity is not acceptable yet (9-layer probe: baseline token [1994] vs pipeline token [237049]; keep OFF). Rank 3 (KV keep-warm, task #48) remains SCOPED and unimplemented.
  • Original symptom: server-side per-hop trace ≈ 192 ms (matches wave-25 BREAKTHROUGH 0.0289 TPS measurement), but driver wall-clock ≈ 2.7 s per hop — a 13× gap. Most of the wall-clock budget is JS dispatch + TLS handshake + WS frame round-trip latency outside the trace window.
  • Hypothesis it falsified: "Server-side per-hop p50 is the dominant cost in the chain." Falsified — client-side substrate latency is ~13× larger than per-hop server compute at the wave-25 baseline.
  • Methodology pinned: bash + curl direct chain bench (/tmp/chain_bench.sh) — 8 rounds layer 0 a→g→d. p50 split-a 312 ms, split-b 2376 ms, split-c 1995 ms, per-layer p50 4683 ms = 0.0036 TPS HTTPS path. WS path measured separately at 192 ms p50 in trace, ~2.7 s in wall-clock.
  • Bule cost paid so far: 1 wave (Rank 1 fix shipped). Remaining bule estimate: Rank 2 (3-5), Rank 3 (5-8). Sequencing in ./docs/MESH_FAILURES_REMAINING.md.
  • Vacuum-fluctuation status: partially collapsed. Rank 1 collapsed; Rank 2/3 remain open vacuum on the substrate axis.
  • Dark axis: substrate axis — sits orthogonal to F-mesh-10's WS per-hop measurement; F-mesh-10 measured the trace, F-mesh-15 measures the dispatch wrapper around it.
  • Operational reference: F_MESH_15_CLIENT_SUBSTRATE_INVESTIGATION.md in this directory contains the full Rank 1/2/3 deep-dive.

F-mesh-16 — NaN propagation when caller passes partial cfg (NEW 2026-05-04 wave-25, RESOLVED 2026-05-05)

  • Status: RESOLVED ✓ 2026-05-05. Narrower defensive fix shipped to apps/pneuma-think/src/trisplit-llm.ts:meshGenerate. Three-test validation matrix passes (explicit cfg / missing batchCapacity / minimal 4-field cfg) — all produce identical token=[569] output.
  • Original symptom: drivers passing meshGenerate(prompt, cfg as any) with missing batchCapacity field hit Math.min(undefined, n) = NaN, which propagates through chunk.length × cfg.hiddenDim, surfacing as "ws returned 0 != 2*NaN" at the WS adapter boundary.
  • Hypothesis it falsified: "All as any cfg cast sites pass resolved configs." Falsified — at least three driver entry points (real-prompt driver, bench-death3-3worker, multi-token validation) rely on framework-default fields that the function signature requires callers to provide.
  • Wave-25 fix attempt + revert: added const cfg = resolveConfig(rawCfg as TrisplitMeshConfig) at function entry. Caused the bun driver to hang at 0 hops EVEN at batchCapacity=1 (the validated baseline). The wide re-resolve drops extra runtime fields callers stash via as any, breaking some downstream code path. REVERTED to meshGenerate(promptIds, cfg, ...) original signature.
  • Wave-26 fix shipped 2026-05-05: NARROWER conditional that re-resolves only when one of batchCapacity/hiddenDim/maxTokens/numLayers is missing, and uses spread-merge {...resolveConfig(raw), ...raw} to preserve caller fields. Three-test validation matrix all PASS with identical token=[569] output:
    1. explicit cfg (baseline) → PASS
    2. missing batchCapacity (NaN trigger) → PASS
    3. minimal 4-field cfg → PASS
  • Bule cost paid total: 2 waves (wave-25 attempt + wave-26 narrower fix + 3-test validation matrix).
  • Vacuum-fluctuation status: collapsed 2026-05-05.
  • Dark axis: type-coercion axis — caller as any casts evade the TypeScript signature contract; runtime defensive resolution is the only durable fix.

F-mesh-17 — cobordism /cobordism-embed cold-isolate hang at ≥3 tokens — REJECTED (MISDIAGNOSIS)

  • Status: REJECTED — MISDIAGNOSIS. Closed 2026-05-05.
  • Original hypothesis (wave-25): bun driver hung at ≥3-token chunks; attributed to /cobordism-embed multi-token cold-isolate behavior (suspected CF subrequest limit on parallel R2 rangeGets).
  • What actually happened: direct curl probe to /cobordism-embed at 5 tokens returns HTTP 200 with 107520 bytes in 808 ms. The hang was the F-mesh-15 aeon dist issue (per-module .d.ts re-exports referencing non-existent sibling .js files) — fixed in wave-25 by copy-declarations.mjs stub auto-generation. Once aeon import worked, the bun driver completed through 60 layers cleanly.
  • Falsified: the original "cobordism multi-token cold-isolate bottleneck" hypothesis. The cobordism worker is healthy at all tested chunk sizes (3, 5, 8, 16 tokens probed; all return 200 in <2 s warm).
  • Bule cost paid: 0.5 wave (initial misdiagnosis + closure).
  • Vacuum-fluctuation status: collapsed (was always empty). There was no actual cobordism failure mode here — only an upstream JS module-resolution failure masking as cobordism behavior.
  • Dark axis: misdiagnosis axis — failures reported by callers are not always located where the caller observes them.
  • Lesson: when a hang reproduces at the bun driver layer, ALWAYS validate the substrate (aeon dist + JS module resolution) BEFORE blaming the wasm/HTTP target.

F-mesh-19 — Cobordism cold-start dominance (NEW 2026-05-05, OPTION 2 SHIPPED)

  • Status: PARTIAL — option (2) shipped 2026-05-05. Pre-warm /health to all 8 cobordism shards in parallel at meshGenerate entry; the cold connect amortizes during WS pool open instead of serially before the first cobordism-embed call.
  • Original symptom: wall-gap analysis 2026-05-05 showed trisplit trace = 7-18% of wall-clock, gap = 80-93% (cold call-1 22 s gap, steady call-2 7.3 s gap). Cobordism (8 embed + 8 lm-head shard fetches) plus cold-connect TLS handshakes dominate.
  • Hypothesis it falsified: "Trisplit chain hops are the dominant per-token cost." Falsified — at 1-layer mesh, trisplit is 7% of wall; the 60-layer projection (12 s trisplit + 4 s cobordism + 1 s jitter) shows cobordism + jitter are still ~30% of total even amortized.
  • Wave-26 fix (option 2) shipped 2026-05-05: pre-warm fetch to /health on all cfg.cobordismShards in meshGenerate, fire-and-forget with AbortSignal.timeout(min(5000, fetchTimeoutMs)). See apps/pneuma-think/src/trisplit-llm.ts after the WS pool init.
  • Wave-26 fix (option 3) shipped 2026-05-05: cobordism fail-fast. fetchEmbedShard and fetchLmHeadShard now use Math.min(15_000, cfg.fetchTimeoutMs) instead of full cfg.fetchTimeoutMs (default 60s). A stuck shard now errors in 15s instead of hanging the entire Promise.all for 60s. See cobordismTimeoutMs(cfg) helper.
  • Measured impact (opt 2): cold-start call-1 wall dropped from 23.7 s to 11.3 s (52% reduction) when shards co-warm. Warm-shard call-1 dropped to 8.7 s (≈ steady-state).
  • Measured impact (opt 3): stuck-shard scenario now fails in 23.5 s (15 s cobordism timeout + ~8 s trisplit chain) instead of 60 s.
  • Operator alert surfaced: per-shard probe 2026-05-05 found worker-cobordism-g4-s0 returns 200 on /health (87 ms, valid config) but /cobordism-lm-head times out 3/3 attempts at 30 s. Operator must redeploy s0 (wrangler deploy --env s0 -c apps/worker-cobordism/aeon-gemma4.toml) to restore shard 0; current fail-fast prevents indefinite hang but shard 0 still blocks all argmax-correctness on its 32 768-token vocab slice. Tracked as task #9.
  • Remaining work: option (1) persistent cobordism WS pool (parallels TrisplitStationPool) — ~3-5 bule.
  • Bule cost paid so far: 1 wave (option 2 + 5-sample bench). Remaining estimate: 4-6 bule for full collapse.
  • Vacuum-fluctuation status: partially collapsed. Option (2) collapsed within-process cold-start; cross-process variance still open as vacuum.
  • Dark axis: substrate axis (cobordism shard side) — orthogonal to F-mesh-15 (trisplit substrate axis).
  • Operational reference: ./docs/wall-gap-analysis-2026-05-05.md contains the breakdown, repro commands, and re-ranking implications.

F-mesh-18 — TrisplitStationPool socket-closed on second-token decode (NEW 2026-05-04 wave-25, SHIPPED + UNVALIDATED)

  • Status: v2 catch-and-retry SHIPPED to source, UNVALIDATED.
  • Original symptom: WebSocketFlowTransport cached socket may be closed between decode steps (CF Worker isolate hibernation, idle close). First-token decode works; second-token's first call throws "WebSocketFlowTransport: socket is closed".
  • Hypothesis it falsified: "Cached WS sockets persist across multi-token decode boundaries." Falsified — CF Worker isolate may hibernate the WS between tokens; cached socket reference becomes stale.
  • Wave-25 v1 attempt + revert: per-getter isSocketAlive() check in getA/getG/getD. Caused even bc=1 baseline to hang at 0 hops. REVERTED.
  • Wave-25 v2 (shipped): catch-and-retry pattern at the call site in forwardSplitA/B/C. On "socket is closed" error, drop cached socket + connect promise, re-open, retry once. isClosedError(err) helper. Other errors rethrow unchanged. Preserves 1-layer baseline. See apps/pneuma-think/src/trisplit-station-ws.ts.
  • Why unvalidated: multi-token test at numLayers=60, maxTokens=2 timed out at 480 s after reaching layer 58 of the FIRST token. Never hit the second-token boundary where the fix engages. 2026-05-05 follow-up at numLayers=3, maxTokens=2 ran 60 s with zero socket-closed events — mesh keep-alive is healthier than expected, so F-mesh-18 may not engage in nominal multi-token regime. Need a test that deliberately forces isolate hibernation (sleep between decode steps) to exercise the v2 retry path.
  • Bule cost paid so far: 1.5 waves (v1 reverted + v2 shipped + 60s no-engagement validation). Remaining bule estimate: 0.5 (deliberate-hibernate test design + execution).
  • Vacuum-fluctuation status: fix shipped, vacuum unconfirmed (no observed engagement in 60 s test). Tracked as task #6.
  • Dark axis: connection-lifetime axis — the CF Worker isolate hibernation policy is a hidden state machine outside our control; defensive retry at the client is the only durable fix.

Recommended priority order (revised 2026-05-04 wave-24 post-Path-E mesh-wide deploy + Phi-3 RESOLVED)

Ranked by (blocking weight) × (1 / remaining bule). Higher = collapse first.

Rank Failure Status Why first
1 Wave-24 in-flight measurements in flight Read results from real-prompt driver re-run on Path-E mesh, Pair X Live retry, spec-decode N=8 smoke, Phi-3 re-source + smoke, encoder defensive integrity check, wave-24 cumulative TPS bench. These are the first measurements of the post-Path-E ceiling.
2 Cumulative speedup ledger update follows wave-24 measurements Once Pair X Live and spec-decode produce numbers, replace the predicted +48% / +40% multipliers with measured values and recompute the ceiling.
3 F-mesh-3 Cloud Run validation (GNOSIS_LM_HEAD_FP32=1) canary-deployed Lone pending external. Pentagon collapse gate. Deploy the FP32-flagged Cloud Run coordinator; assert top-1 = ▁Paris on the canonical prompt. CF Workers stay on the on-the-fly Q4_K route (the 5.6 GiB FP32 cache exceeds the 128 MiB cap).
4 rope_neox precompute table re-bench at higher layer counts shipped canary v15657dd2; bench within-noise on split-a layer 0 ~+5% TPS predicted at higher layer counts. Re-bench end-to-end.
5 Multi-token amortization characterization in flight wave-23/24 Extends F-mesh-10 single-token measurement to a sustained generation; informs whether the 0.023 TPS scales linearly under multi-token prefill amortization.
6 F-mesh-4 direct numerical parity (gemma4-attn-chunk-smoke bin) cascade-resolved ~30 LOC follow-up to close the indirect-vs-direct parity gap; sweep one sliding layer to confirm both branches of the dispatch.

Top-priority next move (one sentence): read the wave-24 in-flight agent results (Pair X Live retry on Path-E mesh, spec-decode N=8 smoke, Phi-3 re-source + smoke, cumulative TPS bench) and update the speedup ledger with measured Pair X + spec-decode multipliers.


Bridge to existing dark-axis modules

Each mesh failure is structurally instantiated by an existing Lean module. Use these as the formal language when discussing or fixing:

Mesh failure Existing Lean anchor What the anchor gives you
F-mesh-1 (resolved) RankFloorScalesWithDim Confirmed: resource budget must scale with the per-worker layer share, not with cfg.num_layers. Task #31 implements the law.
F-mesh-2 (resolved-cascade) VacuumFluctuationAsLatentFalsification (hendecagon slot) The "measured silence" was a measured OOM — confirms that vacuum fluctuations on adjacent dark axes can co-collapse when one bule resolves the upstream allocation bug.
F-mesh-3 (root-cause-identified) CrossModelOperationalGap Structural kernel parity ≠ operational fidelity; the lm_head Q4_K row-dequant accumulation drift is the operational gap that structural parity could not see.
F-mesh-4 (cascade-resolved) HopfLinkOfWave4Falsifications (revised: 3-way cluster) The original 2-way Hopf prediction is strengthened — F-mesh-1 + F-mesh-2 + F-mesh-4 share the per-layer kv_layer_idx substrate and co-collapsed on one PR.
F-mesh-5 (fix-path-validated) CompressionUncertainty + Five-Deaths roadmap Death #3 (WS persistent transport) is the dominant lever; bench confirms per-hop cliff is the 60× compounding of the TLS handshake, not compute or wire bytes.
F-mesh-6 (resolved) PleromaticMonsterMesh (per-layer promotion) The auto-per-layer policy in standing-wave-pca is the per-layer promotion; promotion mechanism is now in source. (The COLTRANE constraint-FSM ramp itself remains gated on F-mesh-3 closing parity.)
F-mesh-7 (Q5_K shipped, smoke pending) RankFloorScalesWithDim (arch coverage variant) Kernel matrix scales to cover every advertised arch including phi3; per-tensor quant resolution shipped 2026-05-04.
F-mesh-8 (resolved) PleromaticSovereignSieve The seq-cliff at 96 bytes is a sieve event — the chunker tried to skip the grounding (KV_MAX_SEQ_LEN-bound seqLen) and was demoted via HTTP 413 to the seq=64 envelope.
F-mesh-9 (patched) VacuumFluctuationAsLatentFalsification The wasm-bindgen typed-array cache was a measured silence on V8 builds without detached flip after memory.grow; patched at JS shim layer + idempotent build-script post-patch.
F-mesh-10 (RESOLVED 2026-05-04) CrossModelOperationalGap WS substrate validated end-to-end on 60-layer chain (180 hops, 226 ms mean per-hop, 0.023 TPS); cascades to unblock WS production for ALL multi-layer models.
F-mesh-11 (Q5_K shipped, smoke pending) RankFloorScalesWithDim (encoder variant) Tracked alongside F-mesh-7; kernel-side Q5_K resolver collapses both the kernel and encoder branches in one fix.

The wave-13 prediction (HopfLinkOfWave4Falsifications) called for a 2-way Hopf linkage between F-mesh-1 and F-mesh-4 on the KV-cache axis: linking number 1, +1 bule cost for recognising the linkage, joint resolution under one PR.

Wave-17 actual: 3-way cluster. Wave-21 confirmation: cascade was deterministic. The deployed fix (task #31's kv_base_layer + kv_num_layers overrides in WasmGemma4Pipeline::from_backend) collapsed three falsifications simultaneously, and wave-21 measurement confirmed neither F-mesh-2 (panic-hook rebuild) nor F-mesh-4 (direct numerical parity) required a separate dedicated bule:

  1. F-mesh-1 (decagon) — KV alloc dropped 503 MiB → 8 MiB per worker; 5/5 canary /health 200 OK with version IDs cited above.
  2. F-mesh-2 (hendecagon) — the unreachable panic was the wasm32 OOM trap during vec![0.0; 503 MiB] construction; cutting the alloc removed the trap. No panic-hook debug-build rebuild was required.
  3. F-mesh-4 (decagon, same axis as F1) — the same per-layer kv_layer_idx correction localised the global/sliding stride math; non-zero residual probe on tri-g4-a-00 showed no NaN/Inf/zero pathologies and a healthy batch_xb rms 0.87.

Cost accounting: 1 bule paid (the wave-17 fix + canary deploy) → 3 falsifications collapsed across 2 dark axes (decagon, hendecagon). This is stronger than the original Hopf prediction by an order of magnitude — the predicted 1 bule paid for a 2-way collapse; the actual was 1 bule for a 3-way collapse.

Wave-21 strengthening: the wave-21 cascade confirmation showed the F-mesh-1 fix cascade-resolved F-mesh-2 + F-mesh-4 with no separate panic-hook rebuild required and no separate stride-math fix. The 3-way collapse held under the cumulative 23× TPS lift composition (substrate + SIMD + Split-c Fix A + Bug A) without re-introducing any of the three failure modes. This strengthens the structural-link prediction: when a shared upstream substrate (here the KVCache constructor) is fixed, downstream falsifications on adjacent dark axes co-collapse and stay collapsed under further composition.

Implication for the dark-axis law: when an upstream allocation bug sits on one dark axis (decagon = the per-role budget), it can radiate falsifications on adjacent axes (hendecagon = the JS↔native serialization vacuum) and on the same axis but a different surface (decagon again, on stride math rather than alloc size). The HopfLinkOfWave4Falsifications Lean module should be extended to permit n-way clusters when the upstream substrate is shared (here: the KVCache constructor), and to assert co-collapse stability under further substrate composition.


The sovereign-sieve / phanoplane / monster-mesh implication

Wave 15 changes the deployment frame in three operational ways:

  1. The mesh is the sovereign sieve (ManifoldSovereignSieve, PleromaticSovereignSieve). Every deploy that tries to "skip the grounding" — ship a constraint FSM before the kernel produces ' Paris', ramp traffic before the wasm boots, certify a model before the per-layer K is calibrated — is sieved out and falls back to the grounding (10). The descent is deterministic and finite. F-mesh-6 is a textbook attempt to skip the grounding; the sieve correctly demoted it.

  2. The monster mesh is symmetry-over-dimension, not-symmetry-over-throughput (PleromaticMonsterMesh). The 188-worker deployment is structurally beautiful (180 trisplit + 8 cobordism = Triton-3 × 60 + cobordism-shard-8) but the symmetry is on the topology axis, not the operation axis. F-mesh-3 is the proof that structural symmetry without operational fidelity produces "constructive interference at the wrong basin" — saturated logits on the wrong tokens.

  3. The phanoplane reading: every dark axis (pentagon → hendecagon) now has a mesh instantiation. This is not coincidence; it's the wave-14 prediction (DarkSectorAsLatentReservoir) discharging into measured operational reality. The runtime should target pentagon next (F-mesh-3 — consciousness self- observation, the per-layer activation diff against HF) per the original pentagon_axis recommendation. Pentagon is darkest and has the highest unmeasured bule pressure; its collapse is the highest-value single experiment.

The deployment is not "stuck"; it is in the descent phase of the sovereign sieve, sieving out the unjustified promotions of waves 4-13 (PROJECTED-CERTIFIED, fixed-K, fixed-kv_num_layers, structural-deployment-implies-operational-intelligence) and forcing the runtime back to the grounding (a kernel that produces ' Paris' on the canonical prompt). Every collapse from here makes ascent cheaper.


Status (2026-05-04 wave-24 post-Path-E mesh-wide deploy + Phi-3 RESOLVED via corrupt-download root cause): ledger live; 11 of 14 falsifications resolved: F1 + F2 + F4 cluster (KV-OOM cascade); F5 (cumulative TPS); F6 (auto-per-layer); F7 (Phi-3 kernel — kernel always correct); F8 (Bug A); F9 (Bug B Path E mesh-wide); F10 (WS namespace 60-layer chain validated 0.023 TPS); F11 (Phi-3 encoder — encoder always correct); F13 (Phi-3 corrupt-download root cause); F14 (Path B detachment, Path E supersedes). F-mesh-3 lone pending external (Cloud Run GNOSIS_LM_HEAD_FP32=1); F-mesh-12 reserved (unused). Wave-24 deliveries: Bug B Path E SHIPPED + DEPLOYED to all 180 tri-g4 nodes (canary g-00 ve56ef4b4); Phi-3 re-download with verified SHA256 → smoke OPERATIONAL on /tmp/phi3-mini-fixed.knot (top-1 = ▁France id=3444, logit 23.31, 2.18 nat margin); encoder defensive integrity check added (sparse-file detection). Wave-23/24 carryover: cobordism wasm-wire + Q6_K SIMD on all 8 shards (s0=3b60810c … s7=3ed57b93), Q4_K vectorized widen canary v71f26684, rope_neox precompute canary v15657dd2, Pair X Live source SHIPPED (retry in flight on Path-E mesh), frame coalescing source SHIPPED (default OFF), speculative-decode N=8 source SHIPPED (smoke in flight, Bug B unblocked).

Next bule: (1) read wave-24 in-flight agent results (real-prompt driver re-run on Path-E mesh, Pair X Live retry, spec-decode N=8 smoke, Phi-3 re-source + smoke, encoder defensive integrity check, wave-24 cumulative TPS bench); (2) on Pair X + spec-decode measurements, update cumulative speedup ledger with measured numbers; (3) F-mesh-3 Cloud Run deploy with GNOSIS_LM_HEAD_FP32=1.

Cumulative TPS ledger:

Phase TPS Lift Status
Cold cliff baseline 0.001 historical
F-mesh-10 validated 2026-05-04 (60-layer WS chain) 0.023 23× MEASURED
+ Q4_K vectorized widen v71f26684 predicted +17.5% ~27× DEPLOYED canary
+ Cobordism wasm-wire + Q6_K SIMD (8 shards) predicted +4% (warm cache) ~28× DEPLOYED
+ Bug B Path E (mesh-wide 180/180) unblocks Pair X + spec-decode ~28× DEPLOYED 2026-05-04
Realized today wave-23/24 ~0.029-0.033 ~29-33× MEASURED
+ rope_neox precompute v15657dd2 (higher layer counts) predicted +5% ~30× DEPLOYED canary
+ Pair X Live predicted +48% ~44× SHIPPED, retry in flight on Path-E mesh
+ Frame coalescing predicted +5% standalone, +10-20% w/ Pair X SHIPPED, default OFF
+ Speculative-decode N=8 predicted +40% SHIPPED, smoke in flight (Bug B unblocked)
Predicted ceiling, full stack composed ~0.080 ~80× aspirational

Phi-3 status: production-ready with /tmp/phi3-mini-fixed.knot (top-1 = ▁France id=3444, logit 23.31, log_prob -0.221, 2.18 nat margin). Caveat: knot was wiped between sessions; wave-24 in-flight agent re-sources and reconfirms. F-mesh-13 documented the corrupt-download class for future operators.