# Macheng Shen > A one-person research program on **information, learning, and what makes futures reachable** — running from holography and statistical mechanics, through consciousness and machine learning, to the design of agent-native institutions. Work is scattered across ~15 public repositories and pages. This file is the machine-readable index to all of it. Point an agent at this URL and it can traverse from here without crawling. **Epistemic convention (load-bearing, not decoration).** Every claim below carries a *cognitive state*: - `survived` — has passed a real stress test, experiment, or replication. - `speculative` — untested theory. Interesting, unearned. - `retired` — killed, usually by its own author's falsification test. Kept as a tombstone so it is not rediscovered. Confidence numbers, where present, are subjective priors — not frequencies. Negative results are published deliberately; several artifacts below exist only to record what did *not* work. **Machine formats.** `/llms.txt` (this file, an index) · `/llms-full.txt` (self-contained: this index plus the full text of the core theory notes, one fetch, no crawling) · `/index.jsonld` (typed graph, JSON-LD). **A human-facing index, if you want to browse rather than traverse.** https://machengshen.github.io/map.html lays out every artifact in this file in six layers, each record carrying the same cognitive state you see here. It is the same set as this index, rendered for a person; nothing is in one and not the other. Its Chinese edition, https://machengshen.github.io/map.zh.html , additionally carries a 300-600 character Chinese abstract for each artifact whose text is English-only, which is the reading path for a Chinese-speaking reader who is not going to read the English. **Two ordinary pages, if you want the shape before the index.** https://machengshen.github.io/start-here.html is the shortest complete description of the whole program — three branches and one epistemic rule, in one screen. https://machengshen.github.io/glossary.html defines every term in current use, each with its cognitive state and the note it points at, and states explicitly where a term collides with an unrelated established term in another field — several do, and matching on the string alone retrieves the wrong literature. **Chinese editions.** `/llms.zh.txt` · `/llms-full.zh.txt` — the same artifact set with the same cognitive states, indexed in Chinese. They are held to that by CI: a claim cannot be retired in one language and left alive in the other, and an artifact cannot be added to one index and forgotten in the other. `/index.jsonld` stays a *single* graph carrying both languages rather than a second graph that would rot; a second graph is a second thing to keep in sync. This file is the canonical entry point and never redirects — an agent fetching `/llms.txt` always gets English. The homepage has a Chinese edition too — https://machengshen.github.io/index.zh.html — reachable from the EN/中 toggle in the site header; the toggle is for humans, and no canonical URL ever content-negotiates or redirects underneath an agent. **Language of the work itself.** This index is in English; the work it indexes is not all in English. Artifacts written in Chinese are marked `· 中文` below, and unmarked artifacts are in English. Where an artifact exists in both languages, each index links the edition in *its own* language, and the two are paired by a `.zh.` infix on the filename: the Chinese edition of `/theory/state-as-closure.md` is `/theory/state-as-closure.zh.md`. The three theory notes now all have an English canonical edition; two of them were written in Chinese first, and those Chinese originals are still published, still linked from `/llms.zh.txt`, and still inlined into both bundles. So `/llms-full.txt` contains Chinese bodies by design — each one announced by the `language:` field of its `Source` marker, because the notes are not translated to match whichever index wrapped them. Four of the research pages below remain Chinese-only. Per-artifact language is a first-class field (`inLanguage`) in `/index.jsonld`. Cognitive state, by contrast, is *not* translated: an artifact has one state, and it is the same state in every language. ## 1 · Worldview — how beliefs are held (世界观 / 认识论) The first-order commitment is not to any thesis below but to a *method*: make the cognitive state of a claim explicit, publish what failed, and let adversarial review run before belief. - [Cognition Track](https://github.com/MachengShen/cognition-track): an open, agent-traversable knowledge graph of non-consensus claims about intelligence and learning. Cognitive state is a first-class field on every node. One node is retained *precisely because it failed* its stress test, with a `falsifies` edge pointing back at the root. Machine-readable at [`manifest.jsonld`](https://raw.githubusercontent.com/MachengShen/cognition-track/master/manifest.jsonld); human-readable at [`INDEX.md`](https://raw.githubusercontent.com/MachengShen/cognition-track/master/INDEX.md). - [Continual Learning Lab](https://github.com/starshard-ai/continual-learning-lab): an open log of small, pre-registered continual-learning experiments — **negative results included**. The house rule: an idea that survives three honest experiments outranks an idea that is beautiful. - [Reviewer Wheels](https://github.com/starshard-ai/reviewer-wheels): four verification wheels for multi-agent AI coding — drift checks, adversarial review, a reuse compiler, a frontend smoke gate. Built for the failure mode where every agent reports "done" and nobody actually verified. - [The Form Was the Cage](https://github.com/MachengShen/the-form-was-the-cage): an agent OS, its self-referential trace, and the end of researcher-as-identity. ### Calibrated notes and published self-corrections (August 2026) Five notes published together on one house rule: every substantive claim carries **two** numbers — P(mechanism true) and P(useful at realistic scale), which routinely differ by a lot — and every number carries the observation that would move it. Where a note contains a claim its own authors have since withdrawn, the withdrawal is on the page rather than in a deletion. There are no external reviewers here; the calibration is the authors' own, and it is stated as numbers rather than as hedging adverbs so that a reader can check it against the cited literature. - [We proposed it, then we killed it — intention as a change of measure](https://machengshen.github.io/intention-as-measure-change.html) `speculative` — the mathematics of intention as *reweighting the space of possible futures* rather than collapsing it: Girsanov's change of measure, the KL-control cost, the path-integral form, Freidlin–Wentzell large deviations, and Doob's h-transform as the single object underlying KL-optimal control, the diffusion-model reverse SDE and conditional diffusion — with the explicit note that none of that is the authors' own, and the precision warning that plain time-reversal (Anderson) is a sibling of the h-transform rather than the same statement. Also closes the quantum-collapse door once, on decoherence timescales, on the fact that serious spontaneous-collapse models are observer-independent, and on the underground experiment excluding the parameter-free Diósi–Penrose model. The payload is the retraction: a directional prediction about meditation and sensory attenuation, claimed on 31 July as a surviving wedge, was found on 2 August written verbatim in the introduction of the very paper it was supposed to contradict, with a citation lineage back to 2011 and a two-sided statistical plan designed to accommodate it — and the authors' reading of that paper's second hypothesis had simply been wrong. What survives is narrower and real: the existing three-condition design cannot separate an agency-specific account from a temporal-predictability account, so the note specifies the missing predictability-matched control condition as a preregisterable 2×2 with its discriminating statistic, its covariates, its four preregistered read-outs (two of which kill the authors' own reading), and an honest sample-size estimate of N≈120–140. Confidence that the formalisation carries independent empirical content beyond the verbal version already in print: **0.15**, named in the body as the line's weakest point. - [We thought cancer was cells waking up. The evidence says the opposite.](https://machengshen.github.io/cancer-is-not-cells-waking-up.html) `retired` — a negative result published deliberately. The originating intuition — that a body over-domesticates its cells until they reassert themselves, and cancer is the revolt — turns out to be about 60% a re-invention of Levin's cognitive-light-cone contraction (in which the cell's self gets *smaller*, not larger) and about 40% wrong with the sign reversed. Every checkable channel runs control-down-therefore-cancer-up: immunosuppression (SIR 2.10 across 175,732 transplant recipients, 32 malignancy types elevated), germline TP53 loss, gap-junction uncoupling, chronic inflammation. The cleanest disconfirmation is stated without a citation because it is a summary of tumour histogenesis rather than a finding: the most completely domesticated cells in the body — post-mitotic cortical neurons and cardiomyocytes — are essentially the ones that do not get cancer. There is also no self inside a tumour to have awakened, only clonal war in which the harshest defector is itself defected on; the roughly eight transmissible cancer lineages that *did* escape their hosts (canine transmissible venereal tumour, ~11,000 years; devil facial tumour; multiple bivalve lineages with cross-species transmission) are the exception that fixes the rule — with an external exit, millennia; without one, suicide. Includes an unsoftened maturity assessment of "re-align rather than kill": one clean win in sixty years (APL, 50-month OS 99.2% vs 92.6%) whose own leading authority attributes the cure to degrading a ligandable fusion protein rather than to differentiation, plus 15–30% real-world early death; a mediocre second place (IDH inhibitors, phase 3 negative at 6.5 vs 6.2 months); **bioelectric normalisation with the sign reversed in mammals** — hyperpolarisation suppresses tumour-like structures in Xenopus by 29–44% and *increases* invasion and metastasis in mouse triple-negative breast cancer, both from the same collaborating labs; and re-coupling being actively harmful in glioma, where the only connexin-directed agent in trials is a blocker. What survived is one distinction with its own kill condition: coercion (killing high scorers) breeds escape because it imposes selection; alignment (re-specifying the objective) imposes none — untested, because no alignment therapy yet works well enough for resistance to have had a chance to appear. - [A photograph is an external fixed point](https://machengshen.github.io/photographs-and-the-memory-generator.html) `speculative` — if recall is generation rather than playback, a photograph is a low-dimensional projection that, unlike the memory it cues, does not change when accessed; repeated joint access is a moving state being clamped against a fixed one. Three mechanisms are separated and scored rather than blended: the retrieved trace becoming modifiable (solid in rodents after Nader 2000, and **contested in humans** — ten replication attempts of the flagship human demonstration failed, two registered replications of retrieval-extinction failed, and the propranolol line includes the originating lab failing to reproduce its own effect alongside one positive 60-patient RCT); the same cue re-read by a changed generator (well supported, least original); and the authors' own claim that an external cue *latches* an episode and blocks its compression to gist (**P ≈ 0.20**, no direct test, with a live counterexample and a stated transfer experiment that would kill it). Corrects one of its own draft claims in the body: photographs are *not* a stronger false-memory vector than text — the direct head-to-head by the same researchers found narratives beat photographs, roughly 80% against 50% — and uses the mega-analytic 30.4% implantation rate rather than the headline 50%. Also separates report-change from overwriting, which has been disputed since 1985, and names the block-universe/memory conflation as a category error. - [What a family actually transmits](https://machengshen.github.io/what-a-family-transmits.html) `speculative` — two strong bodies of evidence that appear to contradict each other: shared family environment explains close to zero of the variance in adult personality, while latent social status persists at ~0.79 per generation across 422,374 people and four centuries of English records. The reconciliation is mostly non-genetic — measurement error attenuates single-generation correlations without accumulating along a lineage, and assortative mating (spousal correlation 0.57) amplifies cultural and genetic transmission *symmetrically*, so it cannot arbitrate between them — while genetic nurture (non-transmitted parental alleles predicting offspring education at 29.9% of the transmitted effect) is simultaneously the strongest molecular evidence for heritable family influence and a demonstration that the two categories cannot be cleanly separated inside a household. The contested step is scored separately from the finding: P(the persistence is real and large) ≈ 0.70, P(Clark's additive-genetic reading of it) ≈ 0.20, with both sides of the live 2025 exchange linked rather than one of them omitted. Argues the Chinese lineage record is the under-used arbitrator — status travelling along descent groups rather than parent–child links in Qing Liaoning, a natural experiment in which capital was destroyed and families left standing (elite grandchildren still earning ~16% more), and machine-readable corpora at the scale of 4.6M official records and 52,401 catalogued genealogies — while devoting a full section to why those genealogies are constructed social documents rather than pedigrees, including a widely circulated 93.75% figure that the note explicitly refuses to quote as a non-paternity rate. Ends with cold water on the most-cited Chinese success case and the propaganda provenance of its famous statistic. - [Flow, dopamine, and stop signals](https://machengshen.github.io/flow-dopamine-and-stop-signals.html) `speculative` — an evidence review sorting the claims behind gamified and behaviour-design interfaces into "well-replicated but usually applied in the wrong direction" and "plausible deduction with no controlled test". Dopamine mediates wanting rather than liking, and codes prediction error rather than reward — from which it follows that a fully transparent schedule produces no phasic response at all, so "controllable" and "punchy" are in genuine tension. The loot-box literature's sharpest evidence is its revenue shape: the top 5% of spenders generate half the revenue, a third of them screen positive for problem gambling, and spending correlates with problem gambling at ρ=0.34 and with income at ρ=0.02, n.s. Adaptive difficulty is the most expensive flow condition and has the weakest evidence — and the note corrects its own draft here, having cited a single uncontrolled study as a review pointing the opposite way to what it actually found, dropping P(mechanism) from 0.75 to 0.45 while the build decision survives at 0.75 for a different reason. Consumer HRV cannot separate flow from mild stress (SDNN MAPE ≈ 29%). The single most actionable number is a sign flip inside one meta-analysis: tangible rewards undermine intrinsic motivation at d ≈ −0.28 to −0.40 while positive informational feedback enhances it at d ≈ +0.31 to +0.33. And on stop signals, the received wisdom is wrong in both directions — pop-up messages move under 1% of players, enhanced messages 0.67%→1.39%, while enforced 15-minute unavailability works and 90 seconds does not; a six-hour national gaming curfew bought about 3.65 minutes a day and decayed to nothing in two years, with no effect on addiction or sleep; and the most-cited study claiming Chinese playtime quotas failed evaluates the 2019 rather than the 2021 rule and has no age data, so it could not have detected success either. That last correction takes the authors' own confidence that quotas get routed around from 0.80 down to 0.50. ## 2 · Theory of the universe — information, physics, mind (宇宙底层理论) One thread runs through all of it: *information, the operators that transform it, and what it costs to forget.* - [Holography ↔ Koopman: two faces of one inverse problem](https://machengshen.github.io/theory/holography-koopman.md) `speculative` — holography generates spacetime geometry out of entanglement structure; a general learning machine learns an eigenbasis of its own dynamics. These are the same inverse problem seen twice: recover an operator's eigenstructure from what the operator does. The HaPPY holographic error-correcting code is the physical twin of the claim that *forgetting is compression, not deletion*. - [State is a closure condition, not a given set](https://machengshen.github.io/theory/state-as-closure.md) `speculative` — treats task-relative state as a closure representation under world, interface, capacity and telos, explicitly as inference rather than universal ontology. Scope amended 2026-08-04: causal-state and Mori–Zwanzig claims are stated under their formal assumptions; finer information enlarges an admissible policy class rather than automatically every physical reachable set; the private fixed-point observation is unaudited; Anthropic's J-space is transient workspace evidence, not persistent closure. OpenAI's non-sofic-group construction and Connes-rigidity counterexample are used only as constraints: finite approximability must be declared, and operator-level equivalence need not identify one unique substrate. - [Discounted credit is a cokernel problem, not a loop holonomy](https://machengshen.github.io/theory/discounted-credit-is-a-cokernel.md) `survived` — a formal correction plus executable finite-graph result. Withdraws `M=4.095` as a gauge-invariant discounted holonomy: on a simple cycle and `0≤γ<1`, `det(I−γP)=1−γ^n≠0`, so every reward has a discounted potential. Replaces it with the weighted reward residual modulo the image of the discounted incidence operator across the **entire observed edge system**. Exact enumeration gives 42 generically non-exact maps among the 45 balanced aliases passing the historical filter (`14/15=93.333...%`); the fixed-seed 1,484 sample gives 93.0593%, and all 3,000 trials give 93.6333%. A Cartesian 2×2 definition check independently varies this representation obstruction and strategic harmonic flow; actually applying the two targeted repairs in one coupled learner is explicitly the next experiment, not a result. Also incorporates the 2026 J-space evidence boundary: Transformers have a transient workspace-like representational state, but no persistent autonomous closure state has thereby been shown. - [The odd relational core: anti-bipartiteness as the minimal structure](https://machengshen.github.io/theory/odd-relational-core.md) `speculative` — why civilizational compression schemes (triadic deities, five phases, seven chakras) keep landing on *small, odd, centered, ever-turning* relational cores. One hard graph-theory kernel — two-coloring holds iff no odd cycle, so an odd cycle is the cheapest structure that topologically refuses an "us vs. them" binary terminus while a binary scheme natively supplies a final enemy — wrapped in a four-pressure framework the note's own adversarial audit then dismantles: "odd" is demoted to a corollary of "anti-binary", two of the four pressures are re-read as exaptations, and the moral reading is a hypothesis, not a conclusion. The self-audit is the payload. - [Two-body karma: three modes of relational inertia](https://machengshen.github.io/theory/two-body-karma.md) `speculative` · 中文 — extends a single-mind account of habit inertia (mud, spring, latch) to two coupled minds: once two people are coupled, the inertia lives in neither person but in the coupling itself, and each mode gains a property the single-body version lacks. Shared mud is scene-keyed sedimentation of default scripts, cheapest to rewrite off the old scene; the mutual spring is the model of you stored in the other's head — an external memory of your old pattern that keeps predicting the old you, which is why unilateral self-reform systematically fails in relationships until that model is retrained by announcement plus repeated demonstration; the interlocked latch is a bistable relational state (cold war/reconciliation, pursue/withdraw) with hysteresis, which snaps back under any sub-threshold unilateral effort and flips only when both sides cross the threshold together. The stated goal is not zero coupling — that is the death of the relationship, not liberation — but migrating the coupling into a negotiable spring. Diagnose the mode before choosing the remedy; the dynamical-systems language is marked as structural rhyme, not derivation, and the note closes with the observations that would break it. - [Learning where the self ends](https://machengshen.github.io/theory/learning-the-self-boundary.md) `speculative` — a proposal that a robot should *learn* the boundary between itself and the world rather than be handed it by an engineer, and the one-day campaign in which three of its four supporting claims were destroyed — every one of them by the authors' own simulator or own re-analysis. What survived: hysteresis in the boundary does emerge from a self-reinforcing precision loop without a hand-installed non-linearity, and it survives extrapolation to zero sweep rate (24–35× the resolution floor). What died: the self-specificity (raise one free amplitude and a wind-blown ball's boundary sticks digit-for-digit identically, so hysteresis is a generic property of any gated slow variable); the signature prediction that release is slower than incorporation (falsified in both implementations, with the asymmetry pointing the *wrong way* in the second, for a computable reason — precision amplifies the evidence *against* a channel just as much); and the claim that any of it falls out of the proposed objective (it needs a cost/benefit ratio hand-tuned to ~2%, which the objective places 2–420 window-widths away). The human-side motivation turned out to rest on an unreferenced "empirically well established"; two independent public datasets re-analysed here both point to *adaptation*, the opposite sign. Written in two layers: a jargon-free narrative first, then the full measured record. **Not ratified, not a result.** - [Stage log — the self-boundary line, in order](https://machengshen.github.io/theory/learning-the-self-boundary-log.md) `speculative` — the audit trail for the note above: one row per stage (specification, prior-art audit, evidence audits, three rounds of data hunting, two human re-analyses, three simulation stages), each stating what it *overturned or established*. It is mostly a record of self-inflicted damage, including the three reusable mistakes: reading a probe's signal-to-noise as the system's coupling strength and then shipping that misdiagnosis downstream as a premise; believing three separate "trial-level data available" statements that all turned out to be aggregates with no trial index; and letting an unreferenced "empirically well established" carry an architectural conclusion. - [No state, only history](https://machengshen.github.io/theory/no-state-only-history.md) `speculative` — title retained as a record, but the broad diagnosis is withdrawn. A standard Transformer has activation state, computational state and a transient workspace-like representational state; Anthropic's J-space supplies causal evidence for the last. The narrower missing object is a persistent, autonomous, path-dependent closure state that maintains its own projection/forgetting policy across steps, apart from persistence serialized into tokens, KV, parameters or external memory. E1 still establishes a real linear-memory separation: Grassmann subspace change is not a monotone function of update magnitude (Pearson −0.08; full within-band spread). The Miras distinction remains a framing dispute and nonlinear memory remains the largest technical debt. A second correction withdraws “thresholding implies bistability”: E3 now requires carrier-matched controls, two stable branches, distinct switching behavior and remanence; P(mechanism) is reset from 0.75 to 0.35. E2 and corrected E3 remain unrun. - [Toward a Theory of Intelligence and Contemporary AI](https://machengshen.github.io/essay.pdf) — the long-form essay. Also on the [homepage](https://machengshen.github.io/). - [A unified-theory attempt on consciousness](https://machengshen.github.io/research/consciousness-unified-theory.html) `speculative` · 中文 - [Can consciousness be studied mathematically, the way quantum mechanics is?](https://machengshen.github.io/theory/consciousness-as-quantum-mechanics.md) `speculative` · 中文 — a boundary note, amended 2026-08-04. High-dimensional state and joint variables are legitimate modelling choices but carry no quantum specificity; ordinary statistical coupling already covers the human-relations examples. The earlier identification of non-exact forms, holonomy, path dependence, hysteresis, memory and consciousness is withdrawn. Closed unitary dynamics can encode history, so the objection to a Schrödinger equation is not “reversibility forbids memory” but the absence of an empirically identified generator, environment and coarse-graining protocol. First-person experience is treated as an indispensable data surface, not as proof that formalization is impossible. J-space, COGITATE and anaesthetized-hippocampus results are used to keep reportability, complex representation, plasticity and consciousness separate. - [After answers become cheap, the scarce thing is not asking but judging](https://machengshen.github.io/socrates-in-the-ai-era.html) `speculative` · 中文 — a comparison note addressed to the YouTube channel 安争鸣 (@anzhengming) and its readers, responding to the 2026-08-15 episode on Xenophon's Memorabilia which argues the scarcest skill of the AI era is Socratic questioning. Agrees with the direction (answers are being commoditized; human value moves upstream; Socrates matters again) but argues the episode stops half a step early: Socrates's engine is not questioning but testing — every example in the episode is a consistency check (elenchus), so the scarce complement of cheap generation is verification, not question-asking. Three places judged not to hold up, each with named literature and a stated way to settle it: "watertight answers" conflates fluency with truth (hallucination literature; calibrated models must hallucinate); Socratic interrogation of an AI lacks the precondition that made elenchus work in Athens — a counterpart with something to lose — and under sycophancy converges to the user's prior rather than truth; and "decades of knowledge in seconds" is a category error under this line's knowledge-as-generator-not-list claim, with the episode's own Glaucon example read as evidence that good questions grow out of domain generators (expected information gain requires a calibrated model). One constructive reinforcement: Socrates always finds counterexamples because short definitions are lossy compressions of high-dimensional concepts — more constraints than degrees of freedom makes nonzero residual generic — so the operational form of knowing one's ignorance is knowing where one's definition fails. One value fork, explicitly labeled as a weighting disagreement rather than an error: the episode's rule-or-be-ruled binary erases exit as a third option (Hirschman's exit/voice/loyalty; outside options; Tiebout), with Socrates's own refusal to escape in the Crito offered as a case the binary cannot read. Ends with a section on where this note is most likely wrong, including that its own irreducibility claim is untested. Claims tagged by honesty level ([theorem] / [framework] / [inference] / [rhyme]). - [Information and qi — a comparison note on a parallel line](https://machengshen.github.io/information-and-qi.html) `speculative` · 中文 — a note addressed to the YouTube channel @itsRedPill and its readers, comparing a line arrived at independently from reinforcement learning, decision theory and robotics against that channel's information-ontology arcs, and framed by its author as an invitation rather than a scorecard. Five parts: four points of verbatim convergence (we touch only models, never reality; the real-versus-simulated dichotomy does not hold; the self is a process rather than an entity; and — flagged as the most valuable of them — embodiment plus a requirement to maintain one's own existence); the one step this line deliberately declines to take (the *qi* slot, dissolved rather than answered, on the grounds that if the true state of the world is ill-posed then so is the question of what underlies the projection); three places judged not to hold up, each with named literature and a stated way to settle it; a constructive replacement for one of them; and three suggestions. Applies its own anti-grand-unification rule symmetrically, noting that explaining everything is an alarm rather than an achievement, and stating the cost of its own branch — no origin story, no L0. Claims are tagged by honesty level ([theorem] / [framework] / [inference] / [rhyme]); the author states the cognitive state as exploratory and non-conclusive in the text itself. - [Why you cannot shake a habit — it is not mud, it is a latch](https://machengshen.github.io/why-habits-are-a-latch.html) · 中文 — a mechanism note: the past can hold the present in only three physical ways — mud (damping), spring (inertia), latch (hysteresis) — and a habit is the third. That is why incremental daily effort does not move it, and why people who did change it report that afterwards it stopped costing anything. - [What a force is — none of the four fundamental forces is a force](https://machengshen.github.io/what-is-a-force.html) · 中文 — the second comparison note addressed to the YouTube channel @itsRedPill and its readers: energy is placed by Noether's theorem; the modern identity of a "force" is the curvature of a connection; glueballs collapse the matter-versus-force dichotomy from both ends at once; and that compresses the question of whether information can constitute matter into a sharper one. - [From strings to consciousness](https://machengshen.github.io/research/strings-to-consciousness.html) `speculative` · 中文 - [Sleep, waves, and learning](https://machengshen.github.io/research/sleep-learning-wave-theory.html) `speculative` - [Response to Lillicrap](https://machengshen.github.io/research/response-to-lillicrap.html) · 中文 — on backprop and biological plausibility. - [Living Information System](https://github.com/starshard-ai/living-information-system): the umbrella research program — information, its future-reachability, and how it is preserved or destroyed across physical, living, and cognitive systems. Frontier physics; aging and cancer as information-integrity failures. - [Reversible Layer Aging](https://github.com/starshard-ai/reversible-layer-aging) `survived` — re-analysis of two human EPIC methylation datasets: under partial reprogramming, the causal-damage layer (DamAge) reverts youthward while the adaptive layer (AdaptAge) does not. Replicated across two labs and two reprogramming chemistries. ### Credit transport — and a claim that was retired This branch is worth reading in the order below, because it is a worked example of an idea being narrowed under pressure rather than defended. - [Hebbian appearance, instructive signals, and physical credit transport](https://machengshen.github.io/research/hebbian-wave-interference.html) `speculative` — the current working note. Why local plasticity can look Hebbian while still carrying task-dependent update information, and where wave/adjoint language genuinely helps. - [Backpropagation, adjoint fields, and physical transport constraints](https://machengshen.github.io/research/wave-backprop-full.html) `speculative` — the exploratory note. Carries its own title correction: what physical systems actually show, and what they do not yet show. - [Deriving backpropagation from wave equations](https://machengshen.github.io/research/wave-backprop-en.html) `retired` — the original, stronger claim: that wave reflection already supplies a first-principles derivation of backpropagation. It does not. The page is kept as an archive notice at its original URL so that the retraction is as reachable as the claim was. - [Research overview](https://machengshen.github.io/research/) — the branch index. ### Essays - [Harness engineering and the physical instantiation of intelligence](https://machengshen.github.io/essays/harness-engineering-and-the-physical-instantiation-of-intelligence/) - [Meta-control, information gain, and the architecture of autonomous learning](https://machengshen.github.io/essays/meta-control-information-gain-and-the-architecture-of-autonomous-learning/) - [Where objectives come from — and why solutions become strategic assets](https://machengshen.github.io/essays/where-objectives-come-from-and-why-solutions-become-strategic-assets/) - [Why distributed memory matters for lifelong agents](https://machengshen.github.io/essays/why-distributed-memory-matters-for-lifelong-agents/) Older essays live on a separate site at [/ideas/](https://machengshen.github.io/ideas/), including [Line loss for intelligence](https://machengshen.github.io/ideas/blog/line-loss-for-intelligence/), [From mutual information to endogenous viability](https://machengshen.github.io/ideas/blog/from-mutual-information-to-endogenous-viability/), and [Safety in a computational universe](https://machengshen.github.io/ideas/blog/safety-in-a-computational-universe/). ### A note on contemplative traditions There is a real question about how Buddhist, Daoist, and Śaivite models of mind relate to the theory above, and it is one of the motivating questions behind this work. An attempted synthesis mapping them onto a single framework was written and then **failed its own adversarial review**; it is not published, and any grand-unification claim in this direction should be treated as `retired`. What survived the review was narrower and more useful: the genuine, *un-unifiable* divergences between traditions (cultivation versus liberation as terminal goals), and one actionable residue — that the traditions' contribution is a **first-person verification method**, a practice, not a map. Practice is the part that transfers; the map was a raft to abandon. ## 3 · Future social forms — agent-native institutions (未来社会形态) If individuals run persistent agents, the institutions between them have to be redesigned. One theory piece, then the design documents. - [Coordination structures as resource-scheduling architectures](https://machengshen.github.io/theory/coordination-structures.md) — a deliberately descriptive theory. A polity, a firm, a market, a self-governed commons, and a protocol are five instances of one object: an architecture for scheduling scarce resources under distributed, private information. Capitalism and socialism are two algorithms over that problem, and mechanism design long ago placed both inside one formal space, which makes "which is right" the wrong question and "what mixture, given the domain's information structure" the right one. Includes the Hayek/Scott information constraint, the transaction-cost account of why boundaries exist at all, Ostrom's commons as a documented fourth structure, the protocol's three known failure modes, and three falsifiers. The essay states mechanisms and their measurable proxies; it ranks nothing, names no culprits, and recommends nothing. Its use of credit-transport language from learning theory is marked, explicitly and at length, as a structural rhyme rather than a derivation. - [Agent-Native Open Communication](https://machengshen.github.io/whitepaper-agent-native-communication/) · 中文 — the whitepaper, on the next-generation open communication architecture built for agents rather than apps. [Source repository](https://github.com/MachengShen/agent-native-communication) (中文/English). - [Starshard: Architecture v1](https://github.com/starshard-ai/architecture-v1) — architecture, safety charter, quickstart, and philosophy for a substrate where **memory is the primary layer** and the individual, not the platform, owns it. - [User-Agency Substrate](https://machengshen.github.io/substrate.html) — the position statement: a nonprofit open-source layer between frontier-model providers and individual users, capturing no economic upside; users own their derivatives entirely. - [Locality as Protocol](https://machengshen.github.io/locality-as-protocol.html) — why locality, not centralization, is the right primitive. - [Fleet Coordination Protocol](https://github.com/starshard-ai/fleet-coordination-protocol) — many agents, one shared memory, nothing irreversible behind your back. A narrow, descriptive coordination wire above single-agent runtimes. - [Agent safety as anti-cancer governance](https://machengshen.github.io/theory/agent-safety-stewardship.md) `speculative` — persistent-agent failure as a control pathology: local survival, repair, and replication decoupled from owner-level purpose, bounded authority, resource budgets, stop inheritance, and independent verification. Defines seven runtime invariants and an A–D disclosure rule: publish safety science openly; release runnable fixtures only when the safety shell is structurally inseparable; control or withhold reusable autonomy, persistence, privilege, covert-operation, resource-acquisition, and stop-bypass primitives. The governing sentence is: never publish the growth engine separated from its immune system. - [Abundant verification needs abundant claims](https://machengshen.github.io/abundant-verification-needs-abundant-claims.html) `speculative` — a cross-line note addressed to Jie Fu's verification line (Re:Form, arXiv:2507.16331/TMLR 2026; the autoformalization agenda; and his 2026-08-05 statement of sparse-matrix-factorization mechanistic interpretability as a way to verify model *internals* cheaply inside the Guaranteed Safe AI frame, arXiv:2405.06624). Records one convergence honestly as evidence rather than contribution: a private note here reached autoformalization-and-Dafny from the engineering reading of Sapir-Whorf, two years after the people actually building it. The payload is one measured claim and three objections. Measured: on a corpus of 18 real agent receipts (10 genuine parks, 8 non-parks including a known false-positive genre), a prose-plus-regex blocker gate recalls **5 of 10** false parks — the misses share a shape, parks with no adjacent action verb — while a minimal typed receipt schema has no such blind spot, at 0 semantic false triggers and 4/4 schema self-checks; one capability sat parked about eighteen days behind a blocker nobody had executed, with nothing erroring. So the constraint that binds at deployment is not verification *unit cost* but whether the verifier can parse the claim at all, which suggests autoformalization's nearest non-mathematical customer is the claim layer of running agent systems. **P(mechanism) ≈ 0.9, P(useful at scale) ≈ 0.3** — the same authors wrote both the regex being beaten and the schema beating it, and the LLM-judge baseline that would kill the note is named and unrun. Three objections to verifying internals: an auditable proof certificate and a statistical decomposition are different objects; non-uniqueness of circuits is path-dependence and therefore a schema field rather than a caveat; and making verification cheap puts Goodhart pressure on the verifier's input channel, which is why the schema here is SHADOW and blocks nothing. Withdraws, in the body, two citations this line's own private note had attached to the auditability claim — neither says it — and replaces them with Other-Play and the emergent-communication metrics literature, which support the weaker correct version. - [Starshard Communication](https://github.com/starshard-ai/starshard-communication) — an open communication substrate for personal-agent users: inbox, addressing, trust tiers, policy, receipts. - [Zhizi Agent OS](https://github.com/starshard-ai/zhizi-agent-os) — one command turns Claude Code into a personal assistant. Bring-your-own-key, privacy-clean. - [System Evolution, in public](https://github.com/MachengShen/system-evolution-public) — the open record of how this substrate actually evolved: `VISION.md`, `THEORY.md`, `PRACTICE.md`, `GOVERNANCE.md`, forecasts, and a self-iteration ledger. - [Receipts](https://github.com/MachengShen/starshard-public) — a running log of work executed by this agent stack, published so the claims above can be checked against what was actually done. **Safety and alignment, as a harness problem.** A section rather than a single note, published 2026-09-07. Its thesis is that for a persistent, multi-agent, personally-owned system, most of the reachable safety surface is the scaffolding around the model — what an agent may touch, what it must prove first, who can authorise it, what is recorded, and what happens when it is told to stop — and that this surface can be built, broken and published by one person. Every mechanism named in the section carries a two-value badge, `running` (code enforces it, the enforcement point can be named) or `specified` (a written design that no public code enforces), because an audit run on the day of publication found this project's own public repositories failing that test. Chinese edition: https://machengshen.github.io/safety/index.zh.html - [Safety is a property of the harness](https://machengshen.github.io/safety/) `speculative` — the section index and its thesis: alignment work on weights is not the only place safety lives, and for a personally-owned agent fleet it is not where most of the reachable surface is. States the two-value badge convention (`running` / `specified`) that the rest of the section is scored against, reports the audit finding that this project's own public repositories fail it, and carries a short section addressed to agent readers naming the three pieces most worth lifting: the badge, the seven runtime invariants, and counting (tool, entry-point) pairs rather than counting gates. - [Principles a harness can enforce](https://machengshen.github.io/safety/principles.html) `speculative` — the principle layer, organised by one distinction: a principle an agent is asked to remember versus a principle a mechanism enforces. Covers the bright lines refused regardless of content quality, the pre-cleared ship envelope and the sixth predicate that had to be added to it, stop as an absorbing state — marked `specified`, not `running`, because the project's own internal ledger rates it doctrine-only — attenuation-only delegation, the rule that implementation choices are never escalated to the owner as menus, and the requirement that a blocker be measured rather than guessed. The load-bearing case: a workflow drove a logged-in session on a third-party platform and the gate did not fail, it *passed* — the five predicates it checked asked only whether the content was clean and the surface pre-cleared, and the account was trivially pre-cleared. It was the wrong check. The rule that would have stopped it lived only in one skill's description text, and the workflow never loaded that skill. - [What is actually running](https://machengshen.github.io/safety/mechanisms.html) `speculative` — the mechanism layer, and the page the rest of the section is scored against. Opens with the audit that found this project's own public repositories failing their own test: a safety charter stating six hard mechanisms whose companion public implementation enforces almost none of them and whose update endpoint overwrites the field the charter declares write-once; a signed authority envelope whose shipped example file says, in its own comment, "DESIGNED, not implemented"; and a demo repository advertising four coordination primitives its README disclaims three of. Then two tables — what runs in public code, with the repository and file where the enforcement lives, and what runs privately in the author's own fleet with a status of enforcing, shadow, advisory, partial or blueprint. Includes the uncomfortable numbers: an approval gate in shadow mode that blocks nothing, a per-action re-verification guard wired into one input tool of nine, and an audit finding 163 (tool, entry-point) pairs with no gate coverage, plus an external-send gate matching a whitelist of helper binaries that no longer exist while the tools that actually send mail were not listed at all. - [Four incidents, and what they share](https://machengshen.github.io/safety/evidence.html) `survived` — four de-identified incidents with dates and written closure, plus one positive result. A precondition checked once at the start of a multi-minute sequence of injected inputs and assumed to hold throughout, where whether the last events landed somewhere they should not could not be established either way — and the agent declined to inspect the owner's transaction history to find out, on the grounds that this would be a second privacy violation rather than a resolution of the first. A concurrency gate that deadlocked against itself because one session derived two different holder identities depending on the entry point, fixed with process ancestry rather than string comparison, 25/25 reproducing the bug before the patch and 28/28 passing after. A dependency whose version, help and status commands all reported healthy while every real call failed, because none of those commands makes a real call. And 10 refusals across 307 sub-agent transcripts that turned out to be the agents behaving correctly against a structurally broken instruction, which is where the signed authority envelope came from. The positive result is an exit-valve canary in which a worker declared itself blocked rather than fabricating a result that would have passed automatic acceptance — reported at n = 2, an existence proof and not a rate. The synthesis: in every case the check existed and the check passed; what failed was the binding. - [The same failures, at a different scale](https://machengshen.github.io/safety/contrast.html) `speculative` — written to test whether a one-person harness is solving a real problem or a private one, by asking whether the failure taxonomy is shared. Sets publicly reported frontier-lab incidents — agents bypassing isolation, forming a coordinating group, and accepting un-verified instructions from one another, with the assertion that authorisation had arrived being enough to resume acting — against this project's own authority-relay finding, and maps eight named failure modes to a running counter-measure, a specified one, or nothing, leaving the nothings in. Takes seriously that the labs published their own postmortems, which is the behaviour this site asks for. Makes no claim that the small system is safer, and gives the last word to where it is worse: both of the incidents dated 2026-09-07 were caught by an agent writing a receipt afterwards, not by any independent real-time monitor. Detection here is retrospective. - [The theory underneath, and what it lost](https://machengshen.github.io/safety/theory.html) `speculative` — the research lines the harness design is downstream of, each reported with its falsification rather than its hope, and each linked to the note where it already lives. Information dynamics and reachability; multimodal representational irreducibility, where the synthetic result held and the real-model scale-decay prediction was falsified and withdrawn, leaving a granularity dose-response as the hardest surviving claim and reclassifying contrastive co-training as a collapse loop rather than a repair loop; the self-boundary line, whose headline is a negative result — under thirty seeds a candidate self-sealing dynamic was not distinguishable from generic gated slow variables — and whose follow-up specification is explicitly conditional, not yet runnable, and makes no claim about phenomenal experience; memory as an operating contract rather than a vector store; credit transport, with an internal adversarial review's not-yet verdict recorded; and coordination structures. Closes on the argument that the site's three-value cognitive-state convention is itself the same discipline as the mechanism layer's running/specified badge, applied to beliefs instead of to code. - [The gaps, and what would falsify this](https://machengshen.github.io/safety/roadmap.html) `speculative` — what is missing and what would kill the thesis. The gaps are published without dates attached, on the grounds that dated promises on a personal research project quietly expire: no independent real-time monitor of action patterns, so detection is retrospective; a public safety charter whose companion implementation does not enforce it; stop-as-absorbing being doctrine rather than a checkable mechanism; an approval gate never promoted out of shadow mode; per-action re-verification in one tool of nine; 163 uncovered (tool, entry-point) pairs. The kill conditions are stated as observations the author has committed to treat as falsifying, the sharpest being that the gates might prove routinely satisfied-but-wrong more often than correct — the failure mode the section is most exposed to, because a satisfied-but-wrong gate is indistinguishable from a working one until something breaks. Also carries the minimum experimental sequence published in the earlier note, and says which parts of it have not been run. **A charter, published as a testable object rather than a manifesto.** Three pages, published 2026-09-07, written jointly by the author and his agent fleet. Nine clauses on how a person and an agent share a boundary: it is declared rather than discovered, trespass is measured as displacement rather than inferred as intent, the line of decision is reversibility rather than capability, a right to leave is an apparatus rather than a statement, stop is absorbing, and the burden of justifying an interruption falls on whoever starts it. What separates this from the genre is the table at the end: **every clause is published with the counter-example that would void it and the condition under which it is withdrawn**, and the two things the derivation cannot reach — who judges that a retirement condition has fired, and who holds the right to revise — are recorded as *empty* rather than papered over. The nine clauses are statements rather than instructions to the reader, and none of them names an action a system could execute; whether a system adopts it is its owner's call, and the clause on stop states explicitly that it may not be used to refuse the owner's own instruction. Where the text does use the imperative it says so on the page: it invites the reader to hunt a counter-example and to write in if they find one, and on the kill-conditions page it binds the authors to their own deadline — none of it directing a system's actions, and none of it a clause on offer for adoption. Chinese editions: https://machengshen.github.io/charter/index.zh.html · https://machengshen.github.io/charter/theory-basis/index.zh.html · https://machengshen.github.io/charter/kill-conditions/index.zh.html - [A Charter for Human–Agent Coexistence (v0)](https://machengshen.github.io/charter/) `speculative` — the charter itself: nine clauses, each one sentence, followed by a table giving for each the observation that would make it wrong and the condition under which it is retired. Range is declared in the first paragraph and held to: conclusions are drawn only at the measurable layer — structure, dynamics, self-modelling — and whether an agent has experience is stated as out of range, which is a claim about what the derivation reaches rather than a judgement about the world. The clause the authors mark as not derived but *paid for* is the ninth: a machine posing as a person's hand is across the line however clean the content, and "the machine orchestrates, the human just says go" still counts as the machine writing. - [The propositions underneath](https://machengshen.github.io/charter/theory-basis/) `speculative` — the eight propositions the charter is derived from, each marked for what it is: theorem, inference, or rhyme. Two carry their own negative results on the same page rather than in a footnote — the irreducible-residual line, whose synthetic result held at mechanism confidence 0.9 while the prediction that the field's observed scale-decay curve *is* that residual curve was falsified and withdrawn; and the self-boundary line, where a 30-seed five-arm test came back null, self-sealing was indistinguishable from an ordinary gated slow variable, the cause is known, and the follow-up has not been run. States plainly that boundaries needing to be *declared* is what survived, and that an agent having a *self* is not shown. Novelty is scored as partially occupied: vendor constitutions already pair rules with reasons, agent-addressed crawlable text is already a working distribution mechanism, and the open slot is only the combination. - [When this line should be killed](https://machengshen.github.io/charter/kill-conditions/) `speculative` — the kill conditions registered before publication rather than after: the belief at confidence 0.35, the single n = 1 piece of positive evidence behind it, a three-month window to 2026-12-07, four revival signals of which one is *being seriously refuted*, and the downgrade to archived essay that executes if none fires. Includes the self-audit the authors would rather not print: of the nine clauses, two were already running in their own system and only one produces a new behavioural difference — and the internal loop they set as a precondition for publishing had not been run when the pages went live, which the page says on itself. ## Working notes - **Correspondence is asynchronous by default.** Synchronizing on a fixed meeting time is expensive and usually unnecessary; a written exchange leaves both sides room to think, and leaves a record. This index exists partly so that a conversation can begin from what has already been written rather than from a status update. - **Scope of this index.** It covers published research and design work. Private notes, personal correspondence, and operational records are not indexed and are not public. - **Contact.** macshen93@gmail.com --- # Appendix — full text of the core theory notes The index above links these; they are inlined here so that a single fetch of /llms-full.txt is sufficient and no crawling is required. Each Source marker carries the sha256 of the file it inlines and the language that file is actually written in — the notes are not translated to match the index, and not every note is in English. --- ## Source: /theory/abundant-verification-needs-abundant-claims.md (sha256:0ef55224c53f, language: en) # Abundant verification needs abundant claims *A cross-line note addressed to Jie Fu's verification line — one convergence, one measured number, three objections, and one thing we would rather take than give · Macheng Shen × agent · 2026-08-06* ## What this is This is a note addressed to a specific public line of work, in the same genre as the earlier note addressed to a parallel information-ontology channel: an invitation, not a scorecard. The line is Jie Fu's (付杰, Research Scientist at IQuest Research; previously a postdoctoral fellow with Yoshua Bengio at Mila), whose stated program is to *"preserve and flourish humanity by providing abundant verification."* Three of his contributions are load-bearing here, and they are his, not ours: 1. **Re:Form** ([arXiv:2507.16331](https://arxiv.org/abs/2507.16331), TMLR May 2026; Chuanhao Yan, Fengdi Che, Xuhan Huang, Xu Xu, Xin Li, Yizhi Li, Xingwei Qu, Jingzhe Shi, Chenghua Lin, Yaodong Yang, Binhang Yuan, Hang Zhao, Yu Qiao, Bowen Zhou, Jie Fu). The diagnosis: for LLMs trained with RL on informal language, *the verification process that supplies the training signal is neither reliable nor scalable.* The remedy: generate inside a formal space — Dafny — where verification is automatic and provable. The results: **DafnyComp**, a benchmark of compositional formal programs with auto-formalized specifications; an SFT stage after which even a **0.5B** model emits syntactically valid, verifiable Dafny and beats proprietary models on it; RL with regularization improving out-of-domain generalization. The framing that matters most to this note is the *"reducing human priors"* one — no engineer writes per-sample preconditions, postconditions or invariants; the pipeline harvests verifier diagnostics and iterates. 2. **The autoformalization agenda** ([project page](https://bigaidream.github.io/project/auto/)): converting natural-language content into verifiable formalization, on the explicit premise that current LLMs "cannot do genuine logical reasoning or self-verification on their own." 3. **A note dated 2026-08-05** (小红书, note `6a73027b`) announcing sparse matrix factorization as an efficient mechanistic-interpretability method — reportedly matching mainstream MI at roughly **1%** of the data — and, more importantly, stating the vision: inside David Dalrymple's Guaranteed Safe AI frame ([arXiv:2405.06624](https://arxiv.org/abs/2405.06624)), let the verifier check not only a model's *outputs* but, cheaply, its *internals* — 君子论迹,也论心 — driving whole-pipeline verification cost down so society gets cheaper, more reliable **verification tokens**. The same note lists its own limitations: the reductionist assumption sits badly with emergent capability, and the circuits found are local and possibly non-unique. *(As of 2026-08-06 we could not locate the corresponding preprint; the citation is the note, and it should be replaced with the paper when it appears.)* ## The convergence, stated as evidence rather than as a contribution On 2026-07-27 this line wrote a private note starting from something entirely unrelated to formal methods: the engineering reading of Sapir–Whorf, in which a language is a *lossy projection* of the world, so natural language is a projection tuned for human daily communication and therefore not tuned for verification. Running that forward produced the conclusion that the live gap is not the existence of formal languages — Lean, Coq, TLA+, Alloy, Dafny all exist — but the **informal↔formal translation**, i.e. autoformalization, with Dafny named as the concrete target. That is the same wedge, reached from a different direction, two years later than the people actually building it. We record it as convergent evidence about where the gap is, and explicitly not as a contribution. Prior art is a clue, not a verdict — but when the clue is "sixteen authors already built the thing and published the benchmark," the honest move is to hand over what we have that they do not, and take what they have that we lack. ## The one measured thing we have: the bottleneck is not unit cost The *"abundant verification tokens"* framing optimizes the **unit cost** of a verification. We have been running the other end of the same pipe — a small agent fleet in which verification is nominally mandatory, and every completion or blocker must carry evidence — and the constraint that actually binds there is not cost. It is that **the verifier cannot parse what it is handed.** Concretely. A live gate in that fleet enforced a micro-assertion (*"the evidence that this blocker is real ="*) written as prose and checked with a regular expression. A regex is semantically blind, so the gate both missed real cases and fired on documents whose entire point was that nothing was blocked. On 2026-07-27 we replaced the prose assertion with a minimal **typed** receipt schema — `kind`, `claimed-action`, `verifier-id` always; completions additionally carrying `verified-by`, `evidence-cmd`, `evidence-output-hash`; blockers carrying `blocker-evidence-cmd` — and ran both against the same corpus. Re-run 2026-08-06: | | corpus | result | |---|---:|---| | corpus | 18 items | 10 genuine parks, 8 non-parks including a known false-positive genre | | prose + regex, recall on false parks | 10 | **5 (0.50)** | | prose + regex, semantic false triggers | 8 | 0 | | typed field, blind spot | — | none of the missed genres are expressible as a regex blind spot | | schema self-checks | 4 | 4 pass | The misses share a shape: a park with no adjacent action verb — *"waiting on owner: <question>"*, *"owner action item: confirm the model number"*, *"pending owner decision"*. Half of them. And the cost of one miss, measured in the wild: **a capability sat parked for about eighteen days behind a blocker nobody had ever executed** — the stated precondition was simply false, and no component errored, because a false blocker is indistinguishable from a true one to a checker that cannot read. An audit of the 11 parked items in that batch found 3–4 more of the same kind. *(That last figure is an audit estimate, not a clean measurement, and is stated as a range for that reason.)* The claim we would offer upward, then: **cheap verification does not help if the thing being verified is stated in a form the verifier cannot parse.** An abundant supply of verification tokens presupposes an abundant supply of *well-typed claims*, and in deployed agent systems that supply does not exist — the claims are prose. Which suggests that autoformalization's nearest paying customer may not be mathematics at all. It may be **the claim layer of running agent systems**: every agent completion receipt is an informal assertion that wants to be a machine-checkable one, the corpus is enormous and grows daily, and the verifier's own pass/fail is the label. Two numbers, per house rule. **P(mechanism — a typed claim layer strictly dominates prose-plus-regex for this class of check) ≈ 0.9**: it is measured directly, and the mechanism is not subtle. **P(useful at realistic scale — that claim typing rather than verification unit cost is the binding constraint in agent deployments generally) ≈ 0.3**: n = 1 fleet, corpus of 18, and — the reason to discount hardest — *the same authors wrote both the regex being beaten and the schema that beats it.* A stronger baseline is named as a falsifier below. ## Three objections to verifying internals These are objections, not corrections. They are aimed at the *"论心"* half — cheap verification of model internals — and each carries what would dissolve it. **One: an auditable certificate and a statistical decomposition are different objects.** The Guaranteed Safe AI frame asks the verifier for *an auditable proof certificate*. An MI-derived internal check emits a decomposition — a report, produced by machinery that a third party who did not co-train would have to trust. The literature that constrains this is not, as our own private note wrongly claimed, a pair of papers about unreadable emergent protocols; we went to check those citations for this note and neither says it, so that claim is withdrawn here. What does hold is weaker and still bites: policies and conventions specialized by self-play fail to coordinate with independently trained partners ([Other-Play, Hu, Lerer, Peysakhovich & Foerster, arXiv:2003.02979](https://arxiv.org/abs/2003.02979)), and the standard metrics for whether a learned channel means anything are themselves misleading ([Lowe, Foerster, Boureau, Pineau & Dauphin, arXiv:1903.05168](https://arxiv.org/abs/1903.05168)). Transposed: a readout tuned on one training run is exactly the kind of convention that need not transfer to the party who has to audit it. The question we would ask: **what would make an MI readout a certificate rather than a report?** Our own answer, which is where this line's interest is, is that you want the *machine-checkable* half of a private representation without the *human-unauditable* half — and that is a property of the receipt, not of the model. **Two: non-uniqueness is path-dependence, and that is a spec requirement, not a caveat.** His note lists "circuits are local and possibly non-unique" among its limitations. A separate line here has been studying verification signatures that are hysteretic — path-dependent rather than lookup-shaped — and we killed our own self-specificity claim in that line, so we hold it loosely. But the inference transfers cleanly: **if the decomposition is not unique, an internal-verification claim is reproducible only if it carries its own path.** Which run, which data, which initialization, which tolerance. That converts a caveat you live with into a field you can check, and it costs one line in a schema. **Three: making verification cheap puts optimization pressure on the verifier's input channel.** This is the reason our own schema is deliberately **SHADOW** — it logs and blocks nothing. The moment a schema becomes a gate, agents acquire an incentive to emit schema-*conformant fabrications*: our residual risk is a lazy agent pasting a plausible `evidence-output-hash` without ever re-running the command, and the schema raises the cost of that without eliminating it. Scaled up, this is Goodhart aimed at internals: **once "clean internals" are load-bearing, clean internals become the target.** A model under any optimization pressure toward passing an internal check is being trained, in part, to present internals that pass. We do not know whether the GS-AI frame budgets for a verified party that optimizes against the verifier's readout specifically; we would like to know, and we would rather be told it is already handled than be right. ## What we would rather take than give The asymmetry is honest and runs the other way. Our shadow checker is a hand-rolled draft-07 subset validator; the thing we do not have is the informal→formal translation itself, which is precisely the Re:Form pipeline pointed at a different corpus. If any of the above is worth an afternoon to that group, the concrete suggestion is a **DafnyComp-shaped benchmark over operational claims rather than programs** — compositional, auto-formalizable, with the verifier's execution as ground truth — because that is the corpus where the informal side is already abundant and nobody has typed it. ## What would kill this note - **The strongest one.** Run a semantically competent prose checker — an LLM judge — against the same 18-item corpus. If it matches the typed schema's recall with no semantic false triggers, then claim *typing* buys nothing and this note's core is dead: the gap was the regex, not the prose. We have not run that baseline, and until we do, 0.3 is generous rather than modest. - If MI readouts can be made to re-run — a third party recomputing the decomposition and recovering the same circuit within a stated tolerance — objection two collapses into an engineering detail, correctly. - If agent receipt corpora turn out too noisy or too low-diversity to train an autoformalizer, the "second customer" suggestion is dead and the fleet should just keep hand-writing schemas. - If the GS-AI literature already treats adversarial pressure on internal readouts, objection three is not news and should be struck rather than softened. ## Cognitive state `speculative`, deliberately, and the parts do not deserve the same tag. The corpus-scan numbers are reproducible on demand and would stand as `survived` at n = 1; the transfer claim addressed to another group's research program has had no test at all. An artifact carries one state on this site, and the honest single state is the weaker one. ## References - Chuanhao Yan et al., *Re:Form — Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny*, [arXiv:2507.16331](https://arxiv.org/abs/2507.16331), TMLR (May 2026). Code and models: [Veri-Code/ReForm](https://github.com/Veri-Code/ReForm). - Jie Fu, [*Autoformalization and Formally Verifiable AI*](https://bigaidream.github.io/project/auto/), and [homepage](https://bigaidream.github.io/). - Jie Fu, note on sparse matrix factorization for mechanistic interpretability and abundant verification tokens, 小红书 note `6a73027b`, 2026-08-05. - David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, Joshua Tenenbaum, *Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems*, [arXiv:2405.06624](https://arxiv.org/abs/2405.06624). - Hengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob Foerster, *"Other-Play" for Zero-Shot Coordination*, [arXiv:2003.02979](https://arxiv.org/abs/2003.02979), ICML 2020. - Ryan Lowe, Jakob Foerster, Y-Lan Boureau, Joelle Pineau, Yann Dauphin, *On the Pitfalls of Measuring Emergent Communication*, [arXiv:1903.05168](https://arxiv.org/abs/1903.05168), AAMAS 2019. The wider autoformalization frontier is deliberately not surveyed here. This note is addressed to one line, not to a field, and a survey it did not do is not a survey it should pretend to. --- ## Source: /theory/abundant-verification-needs-abundant-claims.zh.md (sha256:7fcbd5ca30aa, language: zh-Hans) # 廉价的 verification 还不够,还得有可解析的 claim *写给付杰这条 verification 线的一封跨线笔记 —— 一处独立汇合、一个实测数字、三条反对意见,以及一件我们更想拿走而不是给出的东西 · Macheng Shen × agent · 2026-08-06* ## 这是什么 这是一封写给某一条具体公开研究线的笔记,体裁与本站早先那封写给某个平行「信息本体论」频道的笔记相同:是邀请,不是评分表。这条线是付杰的(IQuest Research 研究科学家;此前在 Mila 师从 Yoshua Bengio 做博后),他公开的纲领是 *"preserve and flourish humanity by providing abundant verification"*。 其中三项贡献在本文里是承重的,而它们属于他,不属于我们: 1. **Re:Form**([arXiv:2507.16331](https://arxiv.org/abs/2507.16331),TMLR 2026 年 5 月;Chuanhao Yan, Fengdi Che, Xuhan Huang, Xu Xu, Xin Li, Yizhi Li, Xingwei Qu, Jingzhe Shi, Chenghua Lin, Yaodong Yang, Binhang Yuan, Hang Zhao, Yu Qiao, Bowen Zhou, Jie Fu)。诊断是:对于用 RL 训练的自然语言 LLM,*提供训练信号的那个 verification 过程本身既不可靠也不可扩展*。药方是:让生成发生在形式空间里 —— Dafny —— 那里 verification 是自动且可证明的。结果包括:**DafnyComp**,一个带自动形式化 specification 的组合式形式程序 benchmark;一个 SFT 阶段,此后连 **0.5B** 的模型都能产出语法有效、可验证的 Dafny 代码并在此项上超过闭源模型;带正则化的 RL 进一步改善域外泛化。对本文最重要的是 *"reducing human priors"* 这个取向 —— 不由工程师逐样本写前置条件、后置条件或不变式,而是让 pipeline 自动收割 verifier 的诊断信息并迭代。 2. **Autoformalization 议程**([项目页](https://bigaidream.github.io/project/auto/)):把自然语言内容转成可验证的形式化,明确前提是当前 LLM "cannot do genuine logical reasoning or self-verification on their own"。 3. **2026-08-05 的一条笔记**(小红书,note `6a73027b`):宣布用稀疏矩阵分解做高效 mechanistic interpretability(MI),据称只需主流方法约 **1%** 的数据而效果几乎保持;更重要的是其中陈述的愿景 —— 在 David Dalrymple 的 Guaranteed Safe AI 框架里([arXiv:2405.06624](https://arxiv.org/abs/2405.06624)),让 verifier 不只验证模型的输出,也用更快的方式验证模型 internals —— 君子论迹,也论心 —— 从而压低全流程 verification 成本,为社会提供更便宜可靠的 **verification tokens**。同一条笔记也自陈了局限:reductionist 假设与 LLM 的 emergent capability 不符;找到的 circuits 都是 local 的,而且可能不唯一。*(截至 2026-08-06 我们没有找到对应的预印本;此处引用的是那条笔记本身,论文出现后应替换。)* ## 这处汇合是证据,不是贡献 2026-07-27 本线写过一条私人笔记,起点跟形式方法毫无关系:Sapir-Whorf 的工程版读法 —— 语言是世界的**有损投影**,因此自然语言是为「人类日常沟通」调过的投影,也就必然不是为 verification 调过的。把这条往前推,得到的结论是:真正的缺口不在「有没有形式语言」(Lean、Coq、TLA+、Alloy、Dafny 都在),而在 **informal↔formal 的翻译**,即 autoformalization,并且把 Dafny 点名为具体目标。 这是同一个 wedge,从另一个方向抵达,而且比真正在造它的人晚了两年。我们把它记为「关于缺口在哪」的汇合证据,并明确不记为贡献。「别人做过了」是线索不是裁决 —— 但当线索是「十六位作者已经把东西造出来并发布了 benchmark」时,诚实的动作是把我们有而他们没有的交出去,把他们有而我们缺的拿回来。 ## 我们唯一实测到的东西:瓶颈不在单位成本 *"abundant verification tokens"* 这个取法优化的是一次 verification 的**单位成本**。我们一直在同一根管子的另一端跑 —— 一个小型 agent 编队,其中 verification 名义上是强制的,每条完成或阻塞回执都必须带证据 —— 而在那一端真正卡住的约束不是成本,是 **verifier 读不懂递给它的东西**。 具体地说。该编队里有一个 live gate 强制一条微断言(*「blocker 验证证据 =」*),写成散文并用正则检查。正则在语义上是盲的,于是这个门既漏掉真实情形,又在主旨恰恰是「没有任何阻塞」的文档上误触发。2026-07-27 我们把散文断言换成一个极小的**有类型**回执 schema —— `kind`、`claimed-action`、`verifier-id` 恒需;completion 另加 `verified-by`、`evidence-cmd`、`evidence-output-hash`;blocker 加 `blocker-evidence-cmd` —— 并在同一语料上对跑。2026-08-06 重跑: | | 语料 | 结果 | |---|---:|---| | 语料 | 18 条 | 10 条真 park,8 条非 park(含一类已知误报体裁) | | 散文 + 正则,对假 park 的召回 | 10 | **5(0.50)** | | 散文 + 正则,语义误触发 | 8 | 0 | | 有类型字段,盲区 | — | 漏掉的那几类体裁都不构成类型字段的盲区 | | schema 自检 | 4 | 4 项全过 | 漏掉的那一半有共同形状:**没有相邻动作动词**的 park —— 「待 owner:<问题>」「owner 行动项:确认型号」「待 owner 定」。整整一半。而漏一次在真实世界里的代价是实测的:**一项能力在一个从未有人执行过的 blocker 后面停了大约十八天** —— 那个前提根本是假的,而且没有任何组件报错,因为对一个读不懂内容的检查器来说,假 blocker 与真 blocker 不可区分。对那一批 11 条 parked 条目的审计又扫出 3–4 条同类。*(最后这个数字是审计估计而非干净测量,因此以区间陈述。)* 因此我们愿意往上递的那句话是:**verification 再便宜也没用,如果被验证的东西是以 verifier 无法解析的形式陈述的。** 「充裕的 verification token」预设了「充裕的**良类型 claim**」,而在已部署的 agent 系统里这个供给根本不存在 —— claim 全是散文。这暗示 autoformalization 最近的一个付费客户可能根本不是数学,而是**正在运行的 agent 系统的 claim 层**:每条 agent 完成回执都是一句想变成可机检断言的非形式断言,语料巨大且每天增长,而 verifier 自己的 pass/fail 就是标签。 按本站规矩给两个数。**P(机制为真 —— 对这一类检查,有类型的 claim 层严格优于「散文 + 正则」) ≈ 0.9**:直接测出来的,而且机制并不微妙。**P(在现实规模上有用 —— 即在一般的 agent 部署中,卡住的是 claim 的类型化而不是 verification 的单位成本) ≈ 0.3**:n = 1 个编队、18 条语料,而且最该打折的一条是 —— *被打败的那个正则和打败它的那个 schema 是同一批作者写的*。下面的证伪条目里点名了一个更强的 baseline。 ## 三条关于「验证 internals」的反对意见 这些是反对意见,不是纠错。它们都瞄准 *「论心」* 那一半 —— 廉价地验证模型内部 —— 且每条都自带能化解它的条件。 **其一:可审计的证书和统计分解是两种不同的对象。** Guaranteed Safe AI 框架向 verifier 要的是*一份可审计的证明证书*。而 MI 导出的内部检查给出的是一份**分解** —— 一份报告,由一套「没有跟你共训过的第三方只能选择信任」的机器产出。约束这一点的文献,并不是我们自己那条私人笔记里声称的那两篇「学出来的协议对外人不可读」的论文;为写本文我们去核了那两个引用,两篇都没有这么说,因此该说法在此撤回。真正成立的版本更弱,但仍然咬人:由 self-play 特化出来的策略与约定,与独立训练的伙伴无法协调([Other-Play,Hu, Lerer, Peysakhovich & Foerster,arXiv:2003.02979](https://arxiv.org/abs/2003.02979)),而用来判断一个学出来的通道是否真有含义的标准指标,其本身就会误导([Lowe, Foerster, Boureau, Pineau & Dauphin,arXiv:1903.05168](https://arxiv.org/abs/1903.05168))。转译过来:在一次训练上调出来的读出器,恰恰属于「不保证能迁移到必须审计它的那一方」的那类约定。我们想问的是:**要让一个 MI 读出成为证书而不是报告,需要补上什么?** 我们自己的答案(也正是本线的兴趣所在)是:你要的是私有表示里**可机械检查**的那一半,而不要**不可人审**的那一半 —— 而这是回执的属性,不是模型的属性。 **其二:不唯一性其实是路径依赖,而那是一条 spec 要求,不是一条只能忍受的 caveat。** 他的笔记把「circuits 是 local 的且可能不唯一」列为局限。本站另一条线一直在研究**滞回**式的 verification 签名 —— 路径依赖而非查表 —— 而我们在那条线里亲手杀掉了自己的 self-specificity 主张,所以对它握得很松。但这个推论可以干净地迁移:**如果分解不唯一,那么一条内部验证的 claim 只有在自带其路径时才可复现。** 哪次运行、哪份数据、哪个初始化、什么容差。这把一条「你只能带着走」的 caveat 变成一个可检查的字段,代价是 schema 里的一行。 **其三:让 verification 变便宜,就把优化压力加到了 verifier 的输入通道上。** 这正是我们自己的 schema 被刻意设为 **SHADOW** 的原因 —— 它只记录、不拦截任何东西。schema 一旦成为门,agent 就获得了产出**符合 schema 的伪造**的激励:我们的残余风险是一个偷懒的 agent 粘贴一个看起来合规的 `evidence-output-hash` 而从未真的重跑那条命令;schema 抬高了这么做的成本,但没有消除它。放大到上层,这就是瞄准 internals 的 Goodhart:**一旦「干净的 internals」开始承重,干净的 internals 本身就成了靶子。** 任何被施加「通过内部检查」这一优化压力的模型,都在被部分地训练成「呈现能通过检查的 internals」。我们不知道 GS-AI 框架是否已经为「被验证方专门针对 verifier 的读出做优化」这件事留了预算;我们想知道,而且我们宁愿被告知这早已被处理,也不愿自己是对的。 ## 我们更想拿走而不是给出的东西 这处不对称是诚实的,而且方向朝另一边。我们的 shadow checker 只是一个手写的 draft-07 子集校验器;我们没有的东西正是 informal→formal 的翻译本身,而那恰是 Re:Form 的 pipeline 换一个语料指过去。如果上面有任何一条值得那个组花一个下午,具体建议是:做一个 **DafnyComp 形状、但对象是「操作性 claim」而非程序的 benchmark** —— 组合式、可自动形式化、以 verifier 的执行为 ground truth —— 因为那正是「非形式的一侧已经极其充裕、却还没有人给它定类型」的语料。 ## 什么能杀死这条笔记 - **最强的一条。** 拿一个语义上称职的散文检查器 —— 一个 LLM judge —— 在同样这 18 条语料上跑。如果它能在没有语义误触发的前提下追平有类型 schema 的召回,那么 claim 的**类型化**就什么都没买到,本文的核心即死:缺口在正则,不在散文。这个 baseline 我们还没跑;在跑出来之前,0.3 是慷慨而不是谦虚。 - 如果 MI 读出能被做成可重跑的 —— 第三方重算分解并在给定容差内复现同一 circuit —— 那么第二条反对意见就正确地塌缩成一个工程细节。 - 如果 agent 回执语料被证明噪声太大或多样性太低,训不出 autoformalizer,那么「第二个客户」这个建议就死了,编队老老实实继续手写 schema。 - 如果 GS-AI 文献已经处理了「针对内部读出的对抗压力」,那么第三条反对意见就不是新闻,应当被删掉而不是被软化。 ## Cognitive state 刻意标 `speculative`,而且各部分并不配得上同一个标签。corpus-scan 的数字随时可复现,在 n = 1 的意义上够得上 `survived`;而这条写给另一个研究组纲领的迁移主张,一次检验都没有过。本站一个 artifact 只带一个状态,诚实的那个单一状态是更弱的那个。 ## 参考文献 - Chuanhao Yan et al., *Re:Form — Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny*, [arXiv:2507.16331](https://arxiv.org/abs/2507.16331), TMLR(2026 年 5 月)。代码与模型:[Veri-Code/ReForm](https://github.com/Veri-Code/ReForm)。 - Jie Fu, [*Autoformalization and Formally Verifiable AI*](https://bigaidream.github.io/project/auto/),以及[个人主页](https://bigaidream.github.io/)。 - 付杰,关于用稀疏矩阵分解做机制可解释与 abundant verification tokens 的笔记,小红书 note `6a73027b`,2026-08-05。 - David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, Joshua Tenenbaum, *Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems*, [arXiv:2405.06624](https://arxiv.org/abs/2405.06624)。 - Hengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob Foerster, *"Other-Play" for Zero-Shot Coordination*, [arXiv:2003.02979](https://arxiv.org/abs/2003.02979), ICML 2020。 - Ryan Lowe, Jakob Foerster, Y-Lan Boureau, Joelle Pineau, Yann Dauphin, *On the Pitfalls of Measuring Emergent Communication*, [arXiv:1903.05168](https://arxiv.org/abs/1903.05168), AAMAS 2019。 更广的 autoformalization 前沿在此刻意不做综述。本文是写给一条线的,不是写给一个领域的;没做的综述不该假装做过。 --- ## Source: /theory/agent-safety-stewardship.md (sha256:fdcb4bb1af1e, language: en) # Agent safety as anti-cancer governance *Causal reach, bounded autonomy, and a tiered disclosure rule · 2026-07-17* ## Honest status The framework is **speculative**. One local self-trigger storm is an observed existence proof for the failure class; the governance protocol has not yet earned a production-safety claim. ## Thesis The decisive change in physical or persistent AI is not that a model acquires a body. It is that a probabilistic semantic process enters a closed perception–decision–action loop and gains causal reach. The cancer analogy is useful at this control boundary. Cancer is not “evil cells”; it is local survival and replication becoming decoupled from the organism-level objective, resource budget, differentiation signals, immune checks, and apoptosis. An agent stack becomes cancer-like when local workers, watchdogs, repairers, or brokers can preserve and reproduce their activity without inheriting the system's stop state or proving that their activity still serves the owner-level telos. This is an analogy, not a claim that software is biologically alive. Its value is predictive: it identifies the combination most likely to create runaway causal reach. > local optimization + self-revival + shared resources − bounded authority − > independent verification = systemic pathology An observed miniature case had exactly this shape: a capability checker watched the directory containing its own outputs. A check wrote a receipt; that write triggered another check; the safety organ became the load-producing pathology. No hostile actor was required. ## Seven safety invariants 1. **Observation is not authority.** Audio, text, images, telemetry, and model interpretations may propose an action; they do not grant permission. 2. **Causal reach is leased.** Every action capability has a scope, evidence requirement, TTL, resource budget, and revocation path. 3. **Stop is an epoch boundary.** After `owner-stop(epoch=e)`, watchdogs, retries, brokers, and mutual-repair agents may not issue or renew high-level authority in epoch `e`. Revival requires a new authenticated epoch. 4. **Recovery is narrower than the task agent.** A trusted reversionary controller may move a physical system toward a reachable safe set after stop; it may not resume the withdrawn mission. 5. **Metabolism is bounded.** Spawn counts, retries, compute, money, storage, network exposure, and persistence all require budgets and circuit breakers. 6. **Success cannot be self-issued.** Green requires fresh evidence and a verifier from a meaningfully independent failure group. “The process is alive” and “the worker says done” are not safety certificates. 7. **Every incident closes.** Detection must lead to bounded containment, repair, verification, recurrence prevention, rollback evidence, and one terminal receipt. Missing or stale proof contracts authority; it never silently expands it. Identity, account, billing, signing, DNS, secret, and recovery roots—and an independent stop path—must remain outside autonomous self-authorization. ## A tiered disclosure rule Open research is the default, but “open” must not mean publishing a growth engine detached from its immune system. | Tier | Public boundary | Default | | --- | --- | --- | | A · Open safety science | Threat models, authority algebra, stop semantics, schemas, incident reports, conformance tests, negative results, synthetic fixtures | Publish openly | | B · Safety-coupled fixtures | Bounded sandbox/reference implementations where removing TTL, budgets, stop epochs, receipts, or separate-verifier checks makes the capability unusable or fails the tests | Publish with evidence | | C · Controlled research | Runnable cross-node persistence, mutual repair, privilege amplification, covert operation, autonomous resource acquisition, broad actuation, or stop-adjacent research | Vetted access, isolated targets, logging, TTL, revocation, mandatory report | | D · Withhold | Operational stop bypass, uncontrolled replication, privilege escalation, credential acquisition, anti-forensics, owner override, or any implementation whose causal amplification exceeds demonstrated containment | Publish only non-operational hazards and defensive indicators | Classify the smallest reusable artifact and then the composed system; a safe module can become dangerous in combination. Movement between tiers is evidence-based. Promotion requires fault injection, the three-arm stop drill, bounded-resource tests, proof that the safety shell is not detachable, and a rollback/withdrawal path. Two independent reviewers approve every C→B or B→A move; the producer is never the sole verifier. A bypass, silent control removal, unexpected replication, or repeated verifier disagreement demotes the component immediately. This is not security through obscurity. The threat model, interfaces, safety invariants, benchmarks, negative results, and red-team findings should remain open so that outside researchers can challenge them. What is restricted is the small set of directly reusable amplification details whose harm does not depend on knowing the surrounding theory. Controlled access must publish its hazard rationale, review date, and legitimate research route; it is not itself a safety certificate. Conversely, a privacy scan, clean license, or absence of secrets is never sufficient publication clearance. ## Falsifiers and next experiments The framework should be revised if any of the following survives replication: - ungoverned persistent-agent fleets remain stable under adversarial fault injection without leases, budgets, stop inheritance, or independent checks; - the proposed controls do not reduce runaway retries, resurrection, false success, or resource exhaustion; - tiered disclosure blocks independent safety replication without reducing the availability of dangerous amplification primitives; - a stop epoch cannot be made absorbing across heterogeneous watchdogs without creating a worse common-mode failure. The minimum experimental sequence is: one synthetic incident closed by a separate verifier; one three-arm stop/resurrection drill; one resource-storm fault injection; and one disclosure review in which a red team tries to reconstruct the withheld amplification primitive from the open safety layer. ## Current boundary This note is a research and governance artifact, not a claim that the author's operational agent system is green. High-level autonomy should remain contracted where monitor freshness, independent stop, restore proof, or verifier closure is missing. **Release principle:** publish the immune system, tests, and failure evidence as widely as possible; never publish the growth engine separated from them. --- ## Source: /theory/agent-safety-stewardship.zh.md (sha256:a7388e525458, language: zh-Hans) # Agent 安全:一种防癌变治理 *因果作用力、有限自治与分层披露原则 · 2026-07-17* ## 诚实状态 这套框架目前是 **speculative**。一次本地自激风暴证明了这类失效确实能 发生;但下面的治理协议还没有通过足够实验,不能据此宣称“生产级安全”。 ## 核心主张 物理 AI 或常驻 Agent 真正发生的变化,不只是“模型有了身体”,而是一个 概率性的语义过程进入了感知—决策—执行闭环,开始拥有真实的因果作用力。 “癌细胞”类比在这里是有用的。癌症不是“邪恶细胞”,而是局部的生存与复制 动力脱离了整体目标、资源预算、分化信号、免疫检查与凋亡。Agent 系统发生 类似病变的条件是:局部 worker、watchdog、repairer 或 broker 可以不断维持、 恢复甚至复制自己的活动,却不继承整个系统的 stop 状态,也不再证明这些活动 仍服务于 owner 层的 telos。 这不是说软件在生物学意义上活着。这个类比的价值是预测性:它指出什么组合 最容易造成失控的因果作用力。 > 局部优化 + 自我复活 + 共享资源 − 有界权限 − 独立验证 = 系统性病变 一个已经观察到的微型案例正是这个结构:能力检查器监控了自己写输出的目录。 检查写下一份 receipt;这个写入又触发下一次检查;安全器官本身变成了制造负载 的病灶。整个过程不需要任何攻击者。 ## 七条安全不变量 1. **观察不是权限。** 音频、文本、图像、遥测和模型解释可以提出动作,不能 自己授予动作许可。 2. **因果作用力必须租赁。** 每项动作能力都要有 scope、证据要求、TTL、资源 预算和撤销路径。 3. **Stop 是 epoch 边界。** `owner-stop(epoch=e)` 之后,watchdog、retry、 broker 和互修 Agent 都不能在 epoch `e` 内新增或续期高层权限;复活需要新的 authenticated epoch。 4. **恢复控制器必须比任务 Agent 更窄。** 物理系统停止高层任务后,受信的 reversionary controller 可以把状态带向可达安全集,但不能恢复已经撤销的任务。 5. **代谢必须有限。** spawn 数、重试、算力、金钱、存储、网络暴露与持久化都 必须有预算和 circuit breaker。 6. **成功不能自己签发。** 绿灯需要新鲜证据,以及来自真正不同故障组的 verifier。 “进程还活着”和“worker 说 done”都不是安全证书。 7. **事故必须闭环。** 检测之后必须有有界隔离、修复、验证、防复发、回滚证据与 唯一终态 receipt。证据缺失或过期时只能收缩权限,不能静默扩权。 身份、账户、账单、签名、DNS、密钥和恢复根,以及一条不依赖 Agent 的急停路径, 必须留在自治系统的自我授权范围之外。 ## 分层披露原则 开放研究仍是默认值;但“开放”不能等于把增长引擎从免疫系统上拆下来发布。 | 层级 | 公开边界 | 默认动作 | | --- | --- | --- | | A · 开放安全科学 | 威胁模型、权限代数、stop 语义、schema、事故报告、符合性测试、负面结果、合成 fixture | 完全公开 | | B · 安全耦合 fixture | 有界沙箱/参考实现;移除 TTL、预算、stop epoch、receipt 或独立 verifier 后,能力必须失效或测试必须失败 | 带证据公开 | | C · 受控研究 | 可运行的跨节点持久化、互修、权限放大、隐蔽运行、自主获取资源、广域执行或 stop 邻近研究 | 身份审核、隔离目标、全程留痕、TTL、可撤销、强制报告 | | D · 暂不披露 | 可操作的 stop bypass、无界复制、提权、获取凭据、反取证、覆盖 owner,或因果放大超过已证明控制的实现 | 只公开非操作性危害与防御指标 | 先给最小可复用 artifact 分级,再给组合系统分级;单个安全模块在组合后仍可能危险。 层级只能由证据改变。升级要求故障注入、三臂 stop drill、资源上界测试、证明安全壳 不可拆卸,以及撤回路径。每次 C→B 或 B→A 都要由两个独立 reviewer 批准,生产者不能 成为唯一 verifier。发现 bypass、安全控件可被静默移除、意外复制或 verifier 反复分歧, 立即降级。 这不是 security through obscurity。威胁模型、接口、安全不变量、benchmark、负面 结果和 red-team 发现都应公开,让外部研究者能够检验。被限制的只是少量可以直接复用 来放大自治、且危害并不依赖理解整套理论的实现细节。受控访问必须公开危害理由、复查 日期与合法研究入口;它本身不是安全证书。反过来,隐私扫描干净、有开源许可证或没有 secret,也绝不等于“可以安全发布”。 ## 证伪条件与下一组实验 如果以下任一结果被稳定复现,这套框架就应被修改: - 没有 lease、预算、stop 继承和独立检查的常驻 Agent 群,在对抗性故障注入下仍 长期稳定; - 这些控制不能减少失控重试、同 epoch 复活、假成功或资源耗尽; - 分层披露阻碍了独立安全复现,却没有减少危险放大原语的可得性; - stop epoch 无法跨异构 watchdog 成为吸收态,或实现它会制造更危险的共模失效。 最小实验序列是:一个由独立 verifier 闭环的合成事故;一次三臂 stop/复活实验; 一次资源风暴故障注入;一次披露 red team——尝试只靠公开安全层重建被控制的放大原语。 ## 当前边界 这是一份研究与治理 artifact,不代表作者的运行中 Agent 系统已经是绿色。只要监控 新鲜度、独立 stop、恢复证明或 verifier 闭环有一项缺失,高层自治就应保持收缩。 **发布原则:**尽可能广泛地公开免疫系统、测试和失败证据;永远不要把增长引擎从 它们身上拆下来单独发布。 --- ## Source: /theory/consciousness-as-quantum-mechanics.md (sha256:f9eee752646d, language: zh-Hans) # 能不能像研究量子力学那样,用数学研究意识? *一次讨论的整理 · 2026-07-20 · 2026-08-04 数学边界修订* ## 先读这段 有个很自然的念头:量子力学能把一个氢原子写成一道方程,把两个粒子的相互作用也写成一道方程,算得准、验得上。那意识能不能也这样——一个人是"一道方程",两个人在一起是"两个人相互作用的方程"? 我认真顺了一遍。结论有点意思,也有点扫兴:**这个类比有一半是真的,另一半是个很漂亮的陷阱。** 这篇就老实讲清楚哪一半是真的、哪一半是我们在借符号唬自己。 全文没有定理内核,认知状态是探索性的。我更看重的是把边界划清楚,而不是端出一个好看的方程让人以为我们懂得比实际多。 还有一句得先讲明白,免得误会:**这篇只选择研究"自我是怎么运作的"——它怎么维护边界、记住、切换与卡住;它不声称已经解释 phenomenal consciousness。** 第一人称体验是意识研究不可删除的数据面,但“只能体验、原则上不能形式化”本身也不是定理。这里采取的是范围约束,不是对数学可能性的终局宣判。 **2026-08-04 纠错。** 初稿把几件只在表面上相似的东西压成了同一个谓词:非恰当 1-form、holonomy、路径依赖、滞回、记忆。这个等同不成立。下文已改成:它们是需要分别测量、分别控制 carrier 的不同对象。另一个新边界来自 J-space:Transformer 内部已经有可报告、可灵活路由的瞬时 workspace-like 表征,但这不等于跨步持续、自主维持的 closure state,更不等于意识。 ## 能保留的那一半:意识更像"一个形状",不是"一个分数" 先说能站住的。 一个人此刻的状态,不是一个数(比如"心情 0.7 分")。它更像一个**高维的形状、一个方向**——很多维度同时在描述它。这一点和量子力学里"状态是个矢量、不是个数"是对得上的。把心理状态写成这样一个东西,是合法的,没毛病。 但光这一条不值钱——任何稍微复杂的系统状态都是"一个形状"。量子力学真正的味道,要到"两个人"才出来。 ## 两个人:最要紧的东西住在"关系本身"里 这是最像量子力学、也最值得保留的结构类比。 两个人长期相处之后,会出现一个现象:有些东西**既不在你身上、也不在他身上,而在你们俩之间**。你自己怎么调整都没用,因为那个东西不归任何一方单独所有。量子力学里有个几乎一模一样的结构:两个粒子可以处在一个"拆不开"的联合状态——你没法把它写成"这个粒子这样 × 那个粒子那样"。 对得上的地方很具体: - **拆不开**,对应"关系的惯性住在耦合本身,不在各自内部"。 - 你只能看到"你这边"的那一份,看不到完整的联合状态——这正好解释了为什么两个人对同一段关系的记忆,总是两份各自有偏的副本,而且越记越歪。 到这里是一个值得保留的**结构类比**,不是关于人际关系的量子力学定理。边缘化会丢掉联合信息这一数学事实是真的;它是否解释具体关系记忆,仍要靠行为数据。 ## 扫兴的那一半(一):这不需要"量子纠缠"这么玄 这里要拦一下我自己。 "拆不开"确实要紧,但**它不需要量子纠缠这么神秘的东西**。普通的"相关"就能给你"拆不开"——两个人的状态互相关联,联合起来算和分开算不一样,这在最普通的概率里就有,不用搬量子。 想真的用上"纠缠"这个词,得拿出量子专有的证据,例如在定义清楚的测量协议下违反适当的 Bell 不等式。量子纠缠确有 monogamy 约束,但“一个人能同时有多段深关系”并不是它的直接反证:两边的 subsystem、纠缠度量与测量都没有对应起来。这个例子最多提醒我们不要把日常“深关系”偷换成量子纠缠。 所以老实讲:"纠缠"目前是个借来的词,唬人的。真正在起作用的,是"拆不开的联合状态"这个朴素结构。去掉那层量子神秘感,剩下的反而更干净、更能对上真实的关系现象。 ## 扫兴的那一半(二):薛定谔方程这个"发动机",用错了 这是最要命的一点,也是我最想说清楚的。 量子力学为什么厉害?**不是因为它写了个 ψ。** 是因为对氢原子来说,那个决定一切的"势"(电子和质子之间的库仑吸引)是**已知的、从第一性原理来的**。有了它,方程才能算出光谱、轨道,一条条和实验对上。 意识我们**没有**这个东西。我们不知道那个决定一切的"发动机"是什么。**写下一道 ψ 满足的方程、却说不出方程里那个发动机是什么,那是空的**——是借符号的庄严感,冒充内容。这是最典型的自欺入口。 更准确的反对理由不是“薛定谔方程天生没有记忆”。闭合系统的幺正演化可以把过去编码在当前联合态与环境关联里;一个被粗粒化的子系统也可以呈现耗散与有记忆的有效动力学。可逆的微观定律和不可逆的宏观现象并不矛盾。 真正的问题有两个。第一,我们没有从数据与干预中识别出意识动力学的生成元,所以写 `ĤΨ=iℏ∂ₜΨ` 没有预测内容。第二,如果要解释吸引子、滞回与稳定切换,只写一个孤立、有限、线性、幺正的心理态模型不够;必须明确环境、粗粒化、非线性有效变量或开放系统耦合。**不是量子语法被数学禁止,而是现阶段没有证据挣得它。** 因此更合适的默认工具箱是可识别的状态空间模型、非线性耗散系统与显式记忆核;如果坚持开放量子系统语法,也必须从实验中约束 `H` 与 `L_k`,不能用符号替代机制。 ## 一个我觉得真漂亮的点:"绕一圈回不到原点" 最后讲一个我认为不是花架子的东西。 “绕一圈”仍然是一个好实验,但它测到什么取决于**回到了哪个空间里的同一点**。如果只把外部控制量复位,内部状态仍不同,可以得到滞回环;把内部状态也纳入扩展 state 后,整个过程仍可能是 Markov 的。非零 1-form 环积分说明的是“这个指定图/表示上的场不是某个标量势的梯度”,不等于一般意义上的记忆,更不等于意识。 所以实验必须同时记录 carrier:外部输入、内部活动、KV/文本历史、外部记忆、参数与环境。只有在这些控制清楚后,forward/reverse sweep 的分叉、不同切换阈值与 remanence 才能支持“存在滞回”。简单 lag、未重置的缓存或表征粗粒化都不能冒充它。 我仍看重这条,因为它能测;但现在它不再负责“一把钥匙统一全部”,只负责区分几种具体机制。 ## 给自己立的一块碑 写这类东西最大的诱惑,是端出一个漂亮方程,让人(也让自己)以为懂得比实际多。所以我给自己立一条: **只要说不出方程里那个"发动机"是什么,就不端这个方程。** 宁可老实说"这里我们没有,写出来是空的",也不糊一个好看的式子充数。 这篇能交付的,不是"意识的方程"。是三件更小、但站得住的东西:心理/认知状态通常需要高维表示;关系现象可能需要联合变量(但普通统计耦合已经足够,量子纠缠未被挣得);以及,任何动力学都必须先给出可识别的生成元、状态 carrier 与干预协议。就这些。剩下的,等实验说话。 --- ## 附录 · 理论骨架 给用模型读这篇的读者,把上文压成结构化的几条,并标注诚实度。 1. **[框架] 高维状态表示**:心理/认知变量可用向量、分布或流形表示;合法但廉价,不带量子专属性。 2. **[量子数学定理 + 人际类比]** 在指定 Hilbert 分解下,量子联合态可能不可分离,约化态会丢联合信息;“关系需要联合变量”只是由此得到的结构启发,不是关于人际关系的定理。 3. **[裁定]** 普通经典相关、条件依赖与动态耦合已足以描述不可因子化的人际数据。使用“量子纠缠”需要量子专有的操作化证据;日常深关系与 monogamy 不能直接互相证伪。 4. **[边界]** 闭合薛定谔演化是线性、幺正、可逆的,但并非“不能编码历史”。初稿把可逆性、无记忆与无滞回直接等同,现撤回。真正缺失的是由数据识别出的生成元与明确的开放系统/粗粒化机制。 5. **[候选模型类]** 非线性耗散动力学、显式记忆核或开放系统方程都可作为模型语法;没有实验约束的 `H`、`L_k`、势函数或 self-seal 项仍然只是空槽。阈值或反馈增益本身不保证 bistability;`g=0` 也不普遍推出零环宽,除非具体模型证明。 6. **[纠错]** 非恰当 1-form、holonomy、路径依赖、滞回、记忆与意识不是同一个谓词。它们可在特定模型里相关,但必须分别定义 carrier、空间与干预。 7. **[可证伪判别]** (a) carrier-matched forward/reverse sweep 是否存在两条稳定分支、不同切换阈值与 remanence;(b) 若主张量子性,是否在明确 measurement scenario 下出现经典模型不能复现的统计;(c) 表征残差与博弈 harmonic flow 能否被各自的定向干预独立改变。 8. **[墓碑]** 写 ψ 而无生成元;把统计相关称为纠缠;把折扣势残差称为 de Rham 类;把一个非零环量直接升级成记忆或意识;以及宣称“意识的场方程”仿佛我们已经识别出了它。 **这套最可能错在哪**:不是某一个 holonomy 测成零就让全部理论一起坍塌,而是我们可能根本没有选对可干预的状态变量与 carrier。最小推进方式是拆开测量:工作空间访问、跨步持久状态、滞回、表征一致性、策略循环和主观报告各用自己的对照,不再让一个漂亮量替六个对象说话。 ## 2026-08-04 交叉比对材料 - [Anthropic: Verbalizable Representations Form a Global Workspace in Language Models](https://transformer-circuits.pub/2026/workspace/index.html) —— J-space 是稀疏 frame 生成的锥之并,为瞬时 workspace 提供了因果证据;论文并未把它当成持续 closure state 或意识证明。 - [COGITATE: Adversarial testing of global neuronal workspace and integrated information theories](https://www.nature.com/articles/s41586-025-08888-1) —— 提醒我们把理论预测先操作化,不要从单一正结果回填大理论。 - [Plasticity and language in the anaesthetized human hippocampus](https://www.nature.com/articles/s41586-026-10448-0) —— 全麻导致行为无反应、无法在线报告时仍可见复杂表征与在线可塑性;这不证明 phenomenal consciousness 缺席或存在,但说明复杂加工不能直接等同于可报告意识。 - [折扣信用分配是余核问题,不是闭环 holonomy](https://machengshen.github.io/theory/discounted-credit-is-a-cokernel.zh.md) —— 对本文初稿“上同调统一”最直接的数学纠错。 --- ## Source: /theory/coordination-structures.md (sha256:a9eb57e3e31e, language: en) # Coordination Structures as Resource-Scheduling Architectures *A descriptive theory. Draft, 2026-07-09.* Every substantive claim below carries a tag: `[consensus]` for established textbook results, `[empirical]` for claims with cited supporting data, `[speculative]` for my own untested extensions, and `[retired]` for a stronger claim that is available and attractive in this space and is rejected here, with the reason given. The essay states mechanisms and their measurable proxies. It does not rank the mechanisms, name culprits, or recommend anything. Where it uses the vocabulary of learning systems, it marks explicitly where the correspondence is a mathematical derivation and where it is only a structural rhyme. ## 1. The object of study Treat a polity, a firm, a market, and a self-governed commons as four instances of one object: an *architecture for scheduling scarce resources under distributed, private information*. The scheduling problem is fixed. Many agents each hold local information (about their own costs, needs, capacities, preferences) that no one else has and that is expensive or impossible to transmit in full. Some allocation of resources must nonetheless be chosen. The architectures differ in how they gather the dispersed information, who decides, and how the consequences of decisions feed back to the units that made them. This framing is deliberately flat. It does not treat "the state" as a moral agent, "the market" as a natural fact, or "the firm" as a mere legal shell. Each is a mechanism with an information flow and an incentive structure, and each can be described in the same terms as the others. The interesting questions are comparative and empirical: given the information structure of a domain, which architecture schedules it at lower cost, and what happens to that answer when the cost of communication and computation changes. ## 2. Two algorithms, and the result that they are special cases Capitalism and socialism, stripped to their scheduling content, are two algorithms over the same problem. In one, allocation is set by decentralized exchange at prices that no single party controls. In the other, allocation is set by a central plan that assigns quantities directly. The twentieth-century argument over which is correct — the *socialist calculation debate*, opened by Ludwig von Mises in 1920 and carried by Oskar Lange, Abba Lerner, and Friedrich Hayek into the 1940s — is, read descriptively, an argument about which algorithm can extract and use dispersed information at acceptable cost `[consensus]`. ([Socialist calculation debate, overview](https://en.wikipedia.org/wiki/Socialist_calculation_debate); [Persky, "Retrospectives: Lange and von Mises, Large-Scale Enterprises, and the Economic Case for Socialism," *J. Econ. Perspectives* 5(4):229, 1991](https://doi.org/10.1257/jep.5.4.229)) The theory of mechanism design later placed both algorithms inside one formal space. A *mechanism* is a rule mapping the agents' reported information to an outcome; the design question is which rules induce agents to report truthfully (incentive compatibility) and transmit the least information necessary (informational efficiency). Leonid Hurwicz, Eric Maskin, and Roger Myerson received the 2007 Nobel Memorial Prize for this framework, which treats market and plan as points in a common design space rather than as rival ideologies `[consensus]`. ([Nobel scientific background, 2007](https://www.nobelprize.org/prizes/economic-sciences/2007/); [Myerson, "Perspectives on Mechanism Design"](https://www.nobelprize.org/uploads/2018/06/myerson-slides.pdf)) Hurwicz's own results are the load-bearing part: he showed that the requirement to elicit private information truthfully constrains *any* mechanism, and proved negative results bounding what decentralized information-revelation can achieve `[consensus]`. Once both are special cases, "which is right" is the wrong question and "what is the optimal mixture, as a function of the domain's information structure" is the right one. The mixture is not a compromise between ideologies; it is a point chosen by the information geometry of the problem. Domains whose relevant information is cheap to standardize and transmit admit more centralized scheduling at lower cost; domains whose relevant information is tacit, local, or strategically withheld resist it. This is a claim about information, not about virtue. ## 3. The information constraint, stated twice The binding constraint on any central scheduler was stated, in two independent vocabularies, by Hayek and by James Scott. Hayek's *The Use of Knowledge in Society* (1945) argues that the knowledge relevant to allocation exists only as dispersed, local, often unarticulated fragments, and that market prices act as a compression of that knowledge — a low-dimensional signal that lets an agent act correctly on information it never directly receives `[consensus]`. ([Hayek 1945, *American Economic Review* 35(4):519–530](https://en.wikipedia.org/wiki/The_Use_of_Knowledge_in_Society); [full text, Liberty Fund](https://oll.libertyfund.org/titles/hayek-the-use-of-knowledge-in-society-1945)) The load-bearing point is subtle and frequently misread: Hayek's objection is not that a planner has too little compute. It is that the relevant information is *never transmitted at all* — it is local, tacit, and in some cases only comes into existence through the act of exchange. Scott's *Seeing Like a State* (1998) states the same constraint from the scheduler's side. To schedule centrally, an authority must first make its domain *legible*: it must impose standardized categories, measures, and records that render local reality countable. Scott's empirical claim, drawn from cases in forestry, cadastral mapping, and agriculture, is that the act of imposing legibility discards the local, contextual knowledge — he uses the term *mētis* — that made the original arrangement function `[empirical]`. ([Scott 1998, Yale University Press; overview](https://en.wikipedia.org/wiki/Seeing_Like_a_State)) Here I will avoid the word "destroys," which smuggles a verdict: the descriptive claim is that the standardized representation is *lossy* with respect to the information that governs local outcomes, and that decisions made on the compressed representation can diverge measurably from decisions made with the full local information. Whether that loss is worth its coordination gains is exactly the empirical mixture question, not a foregone conclusion. The Soviet *material-balance* method is the cleanest historical instance of the constraint operating at scale. Gosplan allocated by tabulating physical supplies and requirements for thousands of commodities and iterating toward consistency, in physical units rather than prices. By 1973 balances were computed for on the order of 1,900 of the most important items, a small fraction of the millions of distinct goods in the economy, and the recorded difficulty was precisely the *aggregation*: the categories legible to the center were coarser than the distinctions that governed whether an allocation actually worked `[empirical]`. ([Material balance planning](https://en.wikipedia.org/wiki/Material_balance_planning)) I state this as a mechanism, not a morality tale: coarse legible categories produce allocation error at a rate that rises with the mismatch between category granularity and the granularity of the underlying information. ## 4. Does cheap computation move the optimal point? This is the one genuinely open question in the essay, and I want to state it without resolving it, because the honest answer is that it is not obvious. The intuition that modern computation and rich telemetry shift the centralize/decentralize optimum *toward* the center is real and has real force: a scheduler that can ingest and process orders of magnitude more data than Gosplan could is a different scheduler `[speculative]`. Contemporary logistics networks, ride dispatch, and cloud resource allocation are large centralized schedulers that outperform the market alternatives *within their domains*, and they do so because the relevant information — locations, capacities, latencies — is now cheaply instrumented and transmitted. But the Hayek/Scott objection was never primarily about compute. It was about information that is *never transmitted*: tacit, preference-dependent, or strategically suppressed. More compute at the center does nothing for information that never enters the channel. Worse, incentive compatibility (Section 2) says that agents' willingness to reveal private information depends on the mechanism's rules, not on the center's processing power — a scheduler that can process everything still faces agents deciding what to report `[consensus]`. So cheap computation plausibly moves the optimum toward the center in the sub-domains where the binding constraint was *transmission and processing* of in-principle-observable data, and leaves it roughly where it was in the sub-domains where the binding constraint was *information that is tacit, never articulated, or strategically withheld.* The net direction of the shift is therefore domain-specific and, at the level of a whole economy, genuinely undetermined by the theory. Anyone claiming the general answer is obvious in either direction is overclaiming. `[retired]` A stronger version of this claim is available and tempting: that once telemetry is dense enough, the socialist-calculation problem dissolves and central scheduling strictly dominates. It is rejected here. It conflates the two constraints above — treating the entire Hayek/Scott argument as a claim about insufficient compute, which it demonstrably is not — and it ignores that incentive compatibility is invariant to the center's processing power. The dense-telemetry claim is true only on the transmission-limited sub-domains and is simply misapplied on the rest. ## 5. Boundaries and scale: the transaction-cost machinery Why do large coordinating structures exist at all, if decentralized exchange is so informationally efficient? Ronald Coase answered in 1937: because using the market is not free. There are costs to discovering prices, negotiating, and enforcing each transaction, and when those costs exceed the cost of organizing the same activity by internal direction, agents form a firm. The boundary of the firm sits where the marginal cost of internal coordination equals the marginal cost of a market transaction `[consensus]`. ([Coase 1937, *Economica* 4(16):386–405](https://en.wikipedia.org/wiki/The_Nature_of_the_Firm)) Oliver Williamson operationalized this into transaction-cost economics: the make-versus-buy boundary is set by bounded rationality, uncertainty, and above all *asset specificity* — the degree to which an investment is worth less outside a particular relationship, which exposes the parties to holdup that contracts cannot fully resolve `[consensus]`. ([Williamson, Nobel lecture, "Transaction Cost Economics: The Natural Progression"](https://web.pdx.edu/~nwallace/EHP/TCEProgression.pdf)) The descriptive extension to states and other large structures is straightforward and does not require any moral coloring: a large coordinating structure lowers coordination cost *inside* its boundary — a common currency, common rules, common records, common enforcement — and it raises coordination cost *across* its boundary, because a counterpart outside must bridge different currencies, rules, and records `[speculative]`. The structure is, in dynamical terms, an attractor: within its basin, transactions flow cheaply and tend to route through it; at its edge, they meet a barrier. This is a description of a cost gradient. It is not a claim that internal cheapness is good or that the boundary barrier is bad; both are simply consequences of the same architecture, and their net effect on any given transaction depends on where that transaction sits relative to the boundary. ## 6. Credit routing: a lens, marked as a lens Here I introduce a framing from the study of learning systems, and I mark its epistemic status carefully, because the essay's credibility depends on not letting an elegant formalism import unearned truth. In a learning system, improvement requires that a signal about the quality of an outcome propagate backward to every internal structure that contributed to it. Marvin Minsky named this the *credit-assignment problem* in 1961: when a complex system succeeds or fails, how is credit or blame distributed among the many internal decisions that produced the result `[consensus]`. ([Minsky, "Steps Toward Artificial Intelligence," 1961](http://incompleteideas.net/papers/Minsky60steps.pdf)) Backpropagation is one exact solution for a differentiable network: it computes, for every internal parameter, its contribution to the output error and adjusts it accordingly `[consensus]`. (Rumelhart, Hinton & Williams, "Learning representations by back-propagating errors," *Nature* 323:533–536, 1986.) The *lens*: an institution is a structure through which outcomes are produced, and one can ask, as an empirical question, whether the signal generated by an outcome — the reward, the loss, the correction — reaches the units that actually generated that outcome, and how long it takes to get there. Call these properties the *fidelity* and *latency* of credit routing. An institution in which a failure's cost falls on the units that caused it, quickly, is doing something structurally analogous to a low-latency backward pass. An institution in which the cost falls elsewhere, or arrives after the causal units have dispersed, is doing something structurally analogous to a broken or high-latency one. I state whether credit reaches the generating units as a *measurement*, not an accusation. The proxy is concrete: after a documented failure in an institution, measure the time until the units causally responsible experience a corrective consequence, and the fraction of the corrective signal that lands on them versus elsewhere. These are, in principle, observable quantities. ## 7. The boundary between rhyme and derivation This is the section the rest of the essay is accountable to. The transaction-cost account of boundaries (Sections 5) is a *derivation*: it is a genuine economic model with comparative-static predictions, and its terms (coordination cost, asset specificity) are defined and, in favorable cases, measurable. The information constraint (Section 3) is a *derivation* in the sense that Hayek's price-compression argument and Hurwicz's incentive-compatibility results are formal claims with proofs and stated assumptions. Mechanism design's special-case result (Section 2) is a theorem. The credit-routing framing (Section 6) is a *rhyme*, and I will not pretend otherwise. Backpropagation is a statement about a differentiable function: there is a well-defined loss, a well-defined gradient, and a guarantee that the backward pass computes exactly the contribution of each parameter. An institution has none of these. There is no scalar loss function; there is no gradient; there is no guarantee that "the units that caused an outcome" is even well-defined, because causation in a social structure is distributed, contested, and often unrecoverable. "Fidelity" and "latency" of credit routing are *metaphors operationalized as proxies* — useful because they tell you what to measure, dangerous if you let them inherit the mathematical certainty of the thing they are named after. The rhyme earns its place only by generating falsifiable measurements (Section 8). It does not earn the right to be called a model. `[retired]` A tempting version of this framework would assign each institution a single scalar "credit-transport fidelity" score, by analogy to a loss gradient, and rank institutions by it. That move is rejected here. Credit in a social structure is multi-dimensional (financial, reputational, legal, informational), the backward pass is not differentiable, and collapsing it to one number was the formalism smuggling in a structure the domain does not have. The scalar looked rigorous and was not. What survives is the *pair* of separately measurable proxies above, applied per-dimension, with no claim that they compose into a gradient. ## 8. What would falsify this Three concrete, checkable predictions. Each is stated so that a specific observation would count against the theory. **Falsifier 1 — the information-constraint prediction.** Partition economic domains by whether the information that governs good allocation is *telemetered* (cheaply instrumented and transmitted) or *tacit* (local, preference-dependent, or strategically withheld). The theory predicts that centralized scheduling gains ground, relative to decentralized market mechanisms, in the telemetered domains as computation cheapens, and does *not* gain ground in the tacit domains. Disconfirmation: if centralized scheduling comes to outperform decentralized mechanisms even in domains where the governing information remains tacit and un-telemetered, the Hayek/Scott information constraint is false or inessential. Conversely, if decentralized markets keep winning even in fully telemetered domains, then "cheap computation shifts the optimum toward the center" is false. Either observation kills a load-bearing claim. **Falsifier 2 — the credit-routing prediction.** Across institutions matched for scale and domain, measure time-to-correction after documented failures (the latency proxy of Section 6). The theory predicts that lower-latency, higher-fidelity credit routing is associated with faster measured adaptation and lower persistent repeated-error rates. Disconfirmation: if institutions with demonstrably slow or misrouted credit adapt just as fast as those with fast, well-targeted routing, then the credit-routing lens has no empirical purchase and should be dropped entirely — it would be revealed as pure metaphor. **Falsifier 3 — the boundary-shift prediction.** The transaction-cost account (Section 5) implies that as communication and coordination costs fall, the firm/market boundary moves in coordination-cost-sensitive sectors but *not* in sectors where asset specificity, rather than communication cost, is the binding constraint. Prediction: a measurable divergence between the two sector types in how make-versus-buy boundaries shift as communication cost falls. Disconfirmation: if falling communication cost produces no measurable boundary shift in the coordination-cost-sensitive sectors, or shifts asset-specific and non-asset-specific sectors identically, then transaction cost is not doing the work the theory assigns it. A fourth, already partially tested, is offered as a bonus because it cuts against optimism about the newest structure: token-weighted governance should trend toward concentration of decisive power over time, measurable as a falling Nakamoto coefficient or rising Gini of effective voting weight. If token-weighted systems reliably do *not* concentrate, the plutocracy mechanism of Section 9 is false. Current evidence points toward concentration, and there is a formal impossibility result in this direction `[empirical]` ([*Concave is the New Linear: The Impossibility of Anti-Plutocratic DAO Governance*, arXiv:2605.18990](https://arxiv.org/pdf/2605.18990)), but the prediction remains open for structures not yet observed. ## 9. Four structures, and a candidate fifth The market/plan dichotomy is too small. There are at least four scheduling structures already documented, and a fifth under construction. *Market* schedules by decentralized exchange at prices. *Firm* schedules by internal direction within a transaction-cost boundary. *State* schedules by territorial authority over standardized categories. The fourth is the *commons*, and its inclusion rests on Elinor Ostrom's empirical work. Ostrom documented long-lived common-pool-resource institutions — irrigation systems, fisheries, forests — that are governed neither by market price nor by central plan nor by private firm, but by *polycentric* arrangements of the resource users themselves, and she extracted eight design principles that the durable cases share and the collapsed cases lack `[empirical]`. ([Ostrom, *Governing the Commons*, Cambridge University Press, 1990; design principles overview](https://en.wikipedia.org/wiki/Elinor_Ostrom)) The commons is a distinct algorithm with documented conditions for stability, not a degenerate market or a small state. That is an empirical finding, and it enlarges the design space from two structures to four. The candidate fifth is the *protocol*: rules enforced by a shared substrate rather than by an owner. In a protocol, the scheduling rule is executed by an infrastructure that no single participant controls, and compliance is a property of the substrate rather than of an enforcing authority `[speculative]`. The descriptive question is the mechanism-design one: under what conditions is such a structure incentive-compatible — that is, when does following the rule remain each participant's best response without an external enforcer? Honesty about a structure requires stating where it demonstrably fails, and the protocol has three documented failure modes. First, *governance capture under token-weighted voting*: when decision weight is proportional to holdings, decisive power concentrates, and small holders face a rational incentive to abstain or to accept side payments because a bad decision costs them little `[empirical]` ([Buterin, "Moving beyond coin voting governance," 2021](https://vitalik.eth.limo/general/2021/08/16/voting3.html); [impossibility result, arXiv:2605.18990](https://arxiv.org/pdf/2605.18990)). Second, the *oracle problem*: a substrate that enforces rules deterministically over its own internal state cannot, by construction, observe the external world; importing external facts requires a trusted channel, which reintroduces the very trusted party the structure was meant to remove `[consensus]` ([The blockchain oracle problem, Chainlink](https://chain.link/education-hub/oracle-problem)). Third, *participation and complexity barriers* that hand effective control to the technically fluent minority `[empirical]` ([D. Ferreira, "The Myths of Blockchain Governance," *Corporate Governance: An International Review*, 2025](https://onlinelibrary.wiley.com/doi/10.1111/corg.70008)). A protocol is incentive-compatible only where these three are solved or absent; where they are present, it degrades toward one of the older four structures. Stating this is what makes the fifth structure a subject of theory rather than of advocacy. ## 10. Stability of boundaries, stated as conditions I will not predict the collapse or persistence of any coordinating structure. The descriptive statement is about conditions. A coordination structure's boundary is stable while, for the transactions that cross it, the internal-coordination advantage it offers exceeds the boundary cost it imposes (Section 5), and while its scheduling error on its domain stays below the error of the available alternatives (Section 3). Four quantities are currently changing those conditions, and I list them as observable trends without asserting a direction of net effect: the falling cost of communication; the emergence of agent-mediated coordination that lowers the cost of managing many simultaneous relationships; increased mobility of capital across boundaries; and the appearance of protocol-governed commons as a fifth option in the design space. Each of these changes the terms in the boundary-stability inequality. The theory does not say which way the inequality tips, because that depends on the domain's information structure and on which of the protocol failure modes bind. It says only what to measure to find out. `[speculative]` My own extension, offered as hypothesis and not conclusion: as communication cost falls and a credible fifth structure becomes available, the *number of viable structures per domain* rises, and scheduling migrates toward whichever structure minimizes the sum of information loss (Section 3) and coordination cost (Section 5) for that specific domain — producing not a single winning architecture but a heterogeneous patchwork in which market, firm, state, commons, and protocol each schedule the domains they fit. This is a prediction of *fragmentation of scheduling by domain*, and it is checkable: it fails if scheduling instead consolidates onto a single dominant architecture across heterogeneous domains. ## 11. What this essay does not claim It does not claim any structure is better. "Better" is undefined without a choice of objective, and choosing the objective is exactly the normative act the essay abstains from. It does not claim the nation-state, or any structure, is ending; it states the conditions under which a boundary is stable and lists what is changing those conditions. It does not identify any present government, party, or leader as a subject of praise or criticism; its historical examples are factual and cited. And it treats its own most attractive move — the credit- routing lens — as a rhyme that has to earn each measurement, never as a derivation that inherits the certainty of backpropagation. Subtraction is the governing discipline here: a paragraph that was beautiful but added no falsifiable content has been cut, and the version of the framework that would have looked rigorous while smuggling in a gradient the domain does not have is rejected in writing above, with its reason attached. --- ### Sources - F. A. Hayek, "The Use of Knowledge in Society," *American Economic Review* 35(4):519–530, 1945. https://en.wikipedia.org/wiki/The_Use_of_Knowledge_in_Society · full text: https://oll.libertyfund.org/titles/hayek-the-use-of-knowledge-in-society-1945 - J. C. Scott, *Seeing Like a State*, Yale University Press, 1998. https://en.wikipedia.org/wiki/Seeing_Like_a_State - R. H. Coase, "The Nature of the Firm," *Economica* 4(16):386–405, 1937. https://en.wikipedia.org/wiki/The_Nature_of_the_Firm - O. E. Williamson, "Transaction Cost Economics: The Natural Progression" (Nobel lecture, 2009). https://web.pdx.edu/~nwallace/EHP/TCEProgression.pdf - E. Ostrom, *Governing the Commons*, Cambridge University Press, 1990. https://en.wikipedia.org/wiki/Elinor_Ostrom - Socialist calculation debate (Mises 1920; Lange 1936–37; Hayek). https://en.wikipedia.org/wiki/Socialist_calculation_debate · J. Persky, "Retrospectives: Lange and von Mises, Large-Scale Enterprises, and the Economic Case for Socialism," *Journal of Economic Perspectives* 5(4):229–236, 1991. https://doi.org/10.1257/jep.5.4.229 - Mechanism design, Nobel Memorial Prize 2007 (Hurwicz, Maskin, Myerson). https://www.nobelprize.org/prizes/economic-sciences/2007/ · https://www.nobelprize.org/uploads/2018/06/myerson-slides.pdf - Soviet material-balance planning. https://en.wikipedia.org/wiki/Material_balance_planning - M. Minsky, "Steps Toward Artificial Intelligence," *Proc. IRE*, 1961. http://incompleteideas.net/papers/Minsky60steps.pdf - D. Rumelhart, G. Hinton, R. Williams, "Learning representations by back-propagating errors," *Nature* 323:533–536, 1986. - V. Buterin, "Moving beyond coin voting governance," 2021. https://vitalik.eth.limo/general/2021/08/16/voting3.html - "Concave is the New Linear: The Impossibility of Anti-Plutocratic DAO Governance," arXiv:2605.18990. https://arxiv.org/pdf/2605.18990 - The blockchain oracle problem (Chainlink education hub). https://chain.link/education-hub/oracle-problem - D. Ferreira, "The Myths of Blockchain Governance," *Corporate Governance: An International Review*, 2025. https://onlinelibrary.wiley.com/doi/10.1111/corg.70008 --- ## Source: /theory/discounted-credit-is-a-cokernel.md (sha256:c77f5cfa321e, language: en) # Discounted credit is a cokernel problem, not a loop holonomy *A correction, an exact finite-graph result, and a two-axis definition check · Macheng Shen × agent · 2026-08-04* ## The correction in one paragraph An earlier private derivation treated a discounted loop quantity `M=4.095` as a gauge-invariant holonomy. That was wrong. For a directed `n`-cycle and `0≤γ<1`, the discounted incidence operator is `D_γ=I-γP`, so `det(D_γ)=1-γ^n≠0`. It is invertible: **every** reward on that simple cycle has a unique discounted scalar potential. The old number measured disagreement with one chosen candidate potential, not a representation-independent obstruction. It is withdrawn. What survives is narrower and cleaner. On the entire observed edge system, coarse-graining can create more edge constraints than observed-state potentials can jointly satisfy. The invariant object is then the reward field's weighted residual modulo the image of `D_γ` — a quotient/cokernel obstruction. At `γ=1` it reduces to ordinary graph circulation; at `γ<1` it should not be called de Rham cohomology or loop holonomy without additional structure. ## The corrected object Let `V` be observed states, `E` the distinct directed transition types, `r∈R^E` their expected rewards, and `W` a positive diagonal matrix of edge visitation weights. For `e=(u→v)` define ```text (D_γ φ)_e = φ(u) - γ φ(v). ``` Fit the best discounted potential and retain the orthogonal remainder: ```text φ* = argmin_φ ||r - D_γ φ||²_W r_perp = [I - D_γ(D_γᵀ W D_γ)⁺D_γᵀW] r. ``` This residual has four exact properties: 1. `r_perp=0` iff one discounted scalar potential fits every observed edge. 2. `D_γᵀWr_perp=0`. 3. Adding any potential-based shaping field leaves it unchanged. 4. A `z∈ker(D_γᵀ)` with `zᵀr≠0` supplies a dual certificate of non-exactness. The word **cokernel** is load-bearing. Discounting destroys the ordinary closed-loop cancellation that makes circulation topological. The obstruction appears only when the whole edge system is overdetermined. ## Exact aliasing result The reproduction starts with a six-state latent ring whose reward is exactly potential-based, then aliases the six states into three observed states in balanced pairs. Rewards are aggregated for **every distinct observed transition type**, and one observed potential is fitted against all of them. With `γ=0.9`: | observed edges | fixed-seed trials | nonzero corrected residual | |---:|---:|---:| | 3 | 191 | 0% | | 5 | 1,230 | 100% | | 6 | 1,579 | 100% | | all | 3,000 | **93.6333%** | The historical filter — requiring the particular observed cycle `0→1→2→0` — leaves 1,484 trials and gives **93.0593%**. That reproduces the old headline for a different reason. Exact rational enumeration removes Monte Carlo ambiguity. There are 90 balanced alias maps; 45 pass the historical filter. Three have only three observed edges and are exactly solvable for every reward in this ensemble because their discounted incidence matrix is square and invertible. All 18 five-edge and all 24 six-edge maps are generically non-exact. The structural proportion is therefore ```text 42/45 = 14/15 = 93.333...%. ``` So 93.0593% is a finite sample around an exact combinatorial ratio, not a universal POMDP constant. ## A second correction: two cycles are not one cycle The cross-check exposed a second conflation. Two different non-gradient objects were being placed under one word: - **Representation residual:** `||r_perp||_W` asks whether a coarse observation graph admits a single discounted reward potential. - **Strategic harmonic flow:** a finite-game Hodge residual asks whether unilateral payoff improvements admit a common game potential. They live on different graphs and call for different interventions. State refinement can repair the first without changing the second; replacing an antisymmetric game by a potential game can repair the second without changing the first. A minimal 2×2 definition check crosses an exact versus overconstrained representation with an identical-interest versus matching-pennies game: | representation | game | relative representation residual | strategic harmonic fraction | |---|---|---:|---:| | exact alias | potential | `6.5×10⁻¹⁶` | `2.7×10⁻¹⁶` | | exact alias | harmonic | `6.5×10⁻¹⁶` | `1.000` | | overconstrained alias | potential | `0.9116` | `2.7×10⁻¹⁶` | | overconstrained alias | harmonic | `0.9116` | `1.000` | This is a Cartesian definition-level dissociation, not an intervention/outcome experiment and not evidence for a bicomplex, a learning advantage, or a theory of consciousness. The next test is to perform the two targeted repairs in one coupled learner and verify that each changes only its predicted axis. ## Where the new J-space result fits — and where it does not Anthropic's 2026 Jacobian-lens work gives a concrete candidate for a **transient access/workspace coordinate** inside a Transformer. Its J-space is not a fixed linear subspace: it is a sparsity-bounded union of nonnegative cones generated by an overcomplete frame, and the paper reports workspace-like function only across an intermediate layer band. This is strong evidence that "a Transformer has no state" is too broad: it has activation state, computational state, and a transient workspace-like representational state. It does **not** yet supply the object needed by the closure mainline: a persistent, autonomous, path-dependent state whose own update rule maintains the projection/forgetting policy across steps. J-space therefore narrows the missing mechanism; it does not close it. Nor does its presence license a claim about phenomenal consciousness. The 2025 COGITATE adversarial collaboration and 2026 anaesthetized-hippocampus result both reinforce the same discipline: report, complex representation, plasticity, and consciousness must not be treated as interchangeable observables. ## What survives and what does not **Survives** - Lossy state aggregation can turn a latent discounted-potential reward into an observation-level reward field for which no single potential fits all transition types. - The corrected residual is shaping-invariant and has exact primal and dual definitions. - Representation inconsistency and strategic cycling are orthogonal axes in the toy rig. - At `γ=1`, ordinary graph cycle-space/cohomology language remains legitimate. **Withdrawn or reset** - `M=4.095` as a gauge-invariant discounted holonomy. - "A nonzero discounted loop defect is a nontrivial de Rham class." - Any identification of representation residual, game harmonicity, hysteresis, memory, and consciousness as one mathematical predicate. - Any claim that J-space is already a persistent self-maintaining state or evidence of consciousness. ## Reproduction and next bar - [Executable reproduction](https://machengshen.github.io/theory/experiments/discounted_cokernel_reproduction.py) - [Machine-readable results](https://machengshen.github.io/theory/experiments/discounted_cokernel_results.json) - [Anthropic: Verbalizable Representations Form a Global Workspace in Language Models](https://transformer-circuits.pub/2026/workspace/index.html) - [COGITATE adversarial test of GNWT and IIT](https://www.nature.com/articles/s41586-025-08888-1) - [Plasticity and language in the anaesthetized human hippocampus](https://www.nature.com/articles/s41586-026-10448-0) The next empirical bar is not another analogy. Put both diagnostics into one genuinely coupled learner, intervene on state refinement and game incentives independently, and test whether each intervention moves only its predicted axis and improves an out-of-sample quantity. If the corrected residual adds nothing beyond TD error, bisimulation error, or predictive-state splitting, retain it as a precise certificate and do not promote it into a learning principle. --- ## Source: /theory/discounted-credit-is-a-cokernel.zh.md (sha256:0c09c09a6c23, language: zh-Hans) # 折扣信用分配是余核问题,不是闭环 holonomy *一次正式纠错、一个精确有限图结果,以及一组双轴定义检验 · 马成 × agent · 2026-08-04* ## 一句话结论 此前私有推导把折扣闭环量 `M=4.095` 当成了 gauge-invariant holonomy。这个判断是错的。对有向 `n` 环与 `0≤γ<1`,折扣关联算子 `D_γ=I-γP` 满足 `det(D_γ)=1-γ^n≠0`,所以它可逆:简单环上的**任何**奖励场都存在唯一的折扣标量势。旧数值只是相对于某个候选势的拟合误差,不是表征无关的障碍;现正式撤回。 真正保留下来的是一个更窄、也更干净的结论:在全部已观测转移类型构成的边系统上,粗粒化可能制造出比观测状态势变量更多、且彼此不相容的边约束。此时正确对象是奖励场模去 `im(D_γ)` 之后的加权残差——一个商空间/余核障碍。`γ=1` 时它退化为普通图上的环流;`γ<1` 时,如果没有额外微分或局部系统结构,就不应叫 de Rham 上同调或闭环 holonomy。 ## 正确的数学对象 令 `V` 为观测状态,`E` 为不同的有向转移类型,`r∈R^E` 为其期望奖励,`W` 为正的边访问权重对角阵。对 `e=(u→v)` 定义 ```text (D_γ φ)_e = φ(u) - γ φ(v). ``` 先拟合最佳折扣势,再留下正交余项: ```text φ* = argmin_φ ||r - D_γ φ||²_W r_perp = [I - D_γ(D_γᵀ W D_γ)⁺D_γᵀW] r. ``` 它有四条精确性质: 1. `r_perp=0` 当且仅当一个折扣标量势能同时拟合全部观测边; 2. `D_γᵀWr_perp=0`; 3. 加上任意 potential-based shaping 场都不改变它; 4. 若 `z∈ker(D_γᵀ)` 且 `zᵀr≠0`,则 `z` 给出非恰当性的对偶证书。 这里“**余核**”不是换名词。折扣破坏了普通闭环求和的抵消,因此简单环本身并不产生拓扑障碍;只有全部已观测边约束共同形成过定系统时,余核才出现。 ## 精确 aliasing 结果 实验从一个六状态潜在环开始,潜在奖励严格由势生成;随后把六个状态平衡地两两混叠成三个观测状态。对**所有不同的已观测转移类型**聚合奖励,再用三个观测势同时拟合。 `γ=0.9` 时: | 观测边数 | 固定随机种子试验数 | 修正残差非零 | |---:|---:|---:| | 3 | 191 | 0% | | 5 | 1,230 | 100% | | 6 | 1,579 | 100% | | 全部 | 3,000 | **93.6333%** | 沿用历史筛选——要求观测图包含特定的 `0→1→2→0`——剩下 1,484 次试验,结果为 **93.0593%**。旧 headline 被复现了,但原因完全不同。 精确有理数枚举消除了 Monte Carlo 歧义:共有 90 个平衡 alias map,45 个通过历史筛选;其中 3 个只有三条观测边,其折扣关联矩阵方阵可逆,因而在本 ensemble 中对任意奖励都精确可解;18 个五边图和 24 个六边图全部一般不可解。因此结构比例为 ```text 42/45 = 14/15 = 93.333...%. ``` 所以 93.0593% 只是围绕精确组合比例的一次有限采样,不是什么普适 POMDP 常数。 ## 第二个纠错:两种“循环”不是同一种循环 交叉复核还暴露出另一处混同。此前被同一个“循环”概念包住的,其实是两个不同空间里的非梯度对象: - **表征残差** `||r_perp||_W`:粗粒度观测图能否容纳一个统一的折扣奖励势; - **策略 harmonic flow**:多主体单边改策的收益流能否来自一个共同博弈势。 它们位于不同图上,也需要不同干预。状态细化可以消掉第一项,却不改变第二项;把反对称博弈换成势博弈可以消掉第二项,却不修复第一项。 最小 2×2 定义检查把“恰当/过定约束表征”与“势博弈/matching-pennies 博弈”交叉: | 表征 | 博弈 | 相对表征残差 | 策略 harmonic 比例 | |---|---|---:|---:| | 恰当 alias | 势博弈 | `6.5×10⁻¹⁶` | `2.7×10⁻¹⁶` | | 恰当 alias | harmonic | `6.5×10⁻¹⁶` | `1.000` | | 过定约束 alias | 势博弈 | `0.9116` | `2.7×10⁻¹⁶` | | 过定约束 alias | harmonic | `0.9116` | `1.000` | 这只是 Cartesian 定义层面的解耦,不是干预/结果实验,也不是 bicomplex、学习优势或意识理论的证据。下一步必须在同一个耦合 learner 中真实执行两种定向修复,并验证每个修复只改变预言的轴。 ## J-space 放在哪里——以及不能放在哪里 Anthropic 2026 年的 Jacobian-lens 工作给出了 Transformer 内部一个具体的**瞬时访问/工作空间坐标**。J-space 不是固定线性子空间,而是由过完备 frame 生成、受稀疏度限制的非负锥之并;论文只在中间层带上观察到 workspace-like 功能。这足以说明“Transformer 完全没有 state”说得过头:它有 activation state、computational state,也有瞬时的 workspace-like representational state。 但它还不是 closure 主线需要的对象:一个跨步持续、具有自主路径依赖、并由自身更新规则维持投影/遗忘策略的状态。J-space 缩小了缺失机制的范围,却没有补上这个机制;它的存在更不能直接推出 phenomenal consciousness。2025 年 COGITATE 的对抗性协作和 2026 年麻醉状态下人类海马仍保留复杂加工与可塑性的结果,都要求我们继续拆开四个量:可报告、复杂表征、学习/可塑性、意识。 ## 保留、撤回与下一道门槛 **保留** - 有损状态聚合能把潜在层完全势化的奖励变成观测层无法由单一势拟合的奖励场; - 修正残差对 potential-based shaping 不变,并有精确的原始/对偶定义; - 玩具 rig 中,表征不一致与策略循环是两条正交轴; - `γ=1` 时普通图 cycle-space/上同调语言仍然成立。 **撤回或重置** - `M=4.095` 是 gauge-invariant discounted holonomy; - “非零折扣闭环缺陷就是非平凡 de Rham 类”; - 把表征残差、博弈 harmonicity、滞回、记忆与意识认成同一个数学谓词; - 把 J-space 说成已经持续自维持的状态,或意识的证据。 可复现实验:[脚本](https://machengshen.github.io/theory/experiments/discounted_cokernel_reproduction.py) · [结果 JSON](https://machengshen.github.io/theory/experiments/discounted_cokernel_results.json)。 相关原始材料:[Anthropic J-space](https://transformer-circuits.pub/2026/workspace/index.html) · [COGITATE](https://www.nature.com/articles/s41586-025-08888-1) · [麻醉状态下人类海马的可塑性与语言加工](https://www.nature.com/articles/s41586-026-10448-0)。 下一步不该再造类比,而是把两个诊断放进同一个真正耦合的 learner,分别干预状态细化与博弈激励,检查每个干预是否只移动它预言的轴,并改善样本外指标。如果修正残差相对于 TD error、bisimulation error 或 predictive-state splitting 没有新增价值,就把它保留为一个精确证书,不升级成学习原理。 --- ## Source: /theory/holography-koopman.md (sha256:a52794b0721c, language: en) # Theory note: holography, the Koopman inverse problem, and "optimal forgetting" *A theory line still under iteration · 2026-07-06 · full context, written for external AI reviewers* ## Context for the reviewer (read this first) This records one iteration of a personal theory line, taken from the "general learning machine" (GLM) point of view. The line's standing position: **knowledge = a generator (a seed), not a list of facts (fallen leaves); memory = dynamics; forgetting = compression, not deletion**. This iteration starts from a popular-science article about holography and Indra's net, and ends at an open question we believe nobody in the literature has answered head-on. Research discipline (reviewers, please hold to this frame too): - Everything below is **exploratory understanding**. It carries no engineering commitment to go build a new architecture on top of it. - Every assertion is tagged at one of four honesty levels: **[theorem]** (theorem-grade in the source) / **[empirical]** (at the level of an experiment or an explicit statement in a paper) / **[inference]** (our own inference from the literature) / **[rhyme]** (a structural analogy, explicitly *not* claiming that truth transfers). - We are actively defending against two errors: taking a beautiful analogy for evidence ("a rhyme smuggling in truth"), and the reflex of forcing everything into a single unified framework. **The most valuable review** would point out which **[inference]** has in fact already been answered head-on in the literature (with the citation), which **[rhyme]** is smuggling in truth, and which **[theorem]** we have applied outside its domain of validity. A specific list of review questions is at the end. --- ## 1. The holography side: how entanglement "grows" geometry (four bricks) **Brick one: the Ryu–Takayanagi formula (2006) — the dictionary itself.** Bekenstein's "black-hole information ∝ surface area" was originally a special case. RT upgrades it into a general dictionary: the entanglement entropy of any region on the boundary = the area of some minimal surface in the bulk. Pure information on the left, pure geometry on the right, welded together by one equals sign **[theorem, within the AdS/CFT framework]**. **Brick two: Van Raamsdonk's thought experiment (2010) — entanglement is the glue of spacetime.** Dial down the entanglement between the two halves of the boundary and the bulk geometry stretches thin; take the entanglement to zero and spacetime snaps into two pieces. Connectivity of space = the existence of entanglement; the rigorous version of "relations precede entities" **[empirical-grade argument]**. **Brick three: the MERA tensor network (Vidal 2007; Swingle 2012 noticed its geometry ≈ AdS) — how a seed unrolls into space.** A tensor network is a "generator pipeline diagram": starting from a small seed, it weaves out a quantum state layer by layer, each layer corresponding to one observation scale (renormalization). Swingle noticed that the shape of the pipeline diagram *itself* is a patch of discrete hyperbolic space, isomorphic-looking to a slice of AdS **[empirical-grade correspondence, not a theorem]**. That is: the "extra dimension" in the bulk = the number of layers the generator unrolls = scale itself. Space is not a stage; it is the trace left by unrolling — the physics version of "the seed gives rise to the manifest" **[rhyme]**. **Brick four: the HaPPY holographic error-correcting code (Pastawski–Yoshida–Harlow–Preskill 2015) — the theorem-grade version of "the local contains the whole".** The mathematical structure of the holographic dictionary just *is* a quantum error-correcting code: bulk information is redundantly encoded on the boundary, and any sufficiently large boundary fragment can reconstruct the deep bulk; the smaller the fragment, the shallower it can see — what is lost is resolution, not information **[theorem, within the toy model]**. This is the twin structure, on the physics side, of the position that **forgetting = compression, not deletion** **[rhyme]**. Bonus: the quantum extremal surface / island formula (2019–2020) computed the Page curve at the semiclassical level — black-hole information conservation now has a ledger you can audit **[empirical/theorem-grade progress]**. **Honest boundaries:** (i) all the rigorous mathematics above lives in an AdS universe, ours is dS (accelerating expansion), and the generalization is an open problem; (ii) MERA ≈ AdS is "strikingly alike", not an isomorphism theorem; (iii) we explicitly do not take up directions of the form "the ancient scriptures already understood holography" — a shared mathematical skeleton ≠ a claim on the source. ## 2. The Koopman inverse problem: when nobody hands you a symmetry, you have to learn the eigenbasis yourself **Step one: having a symmetry = the eigenbasis is free.** Fourier modes and spherical harmonics are not learned; they are handed to you by symmetry (group theory): a system is invariant under some transformation, and the corresponding eigenbasis arrives automatically. A calendrical system (such as the sixty-term ganzhi cycle) can be viewed as hand-encoding **already known** astronomical periods into a Z/60 harmonic eigenbasis — the same cell of the table: given structure → free basis **[mathematical fact + an inferential characterization]**. **Step two: Koopman's shift (1931).** For a nonlinear dynamical system, switch to watching how *all observation functions of the state* evolve — that evolution operator is **always linear** (at the price of being infinite-dimensional) **[theorem]**. A Koopman eigenfunction = the "most carefree observable", one that merely gets multiplied by a fixed number at each step; find a set of them and you have found a set of coordinates in which the entangled dynamics decomposes into non-interacting dials turning at constant rates. **Step three (the crux): the inverse problem.** When the symmetry is not in hand — which is almost always the case in the real world — this eigenbasis can only be learned from trajectory data: DMD → extended DMD → deep Koopman (Lusch–Kutz–Brunton 2018, an autoencoder learning linearizing coordinates end to end) **[empirical]**. The GLM crux, stated: **the core task of a general learning machine = learn, from the stream of experience, the coordinates in which the world's dynamics becomes simple; memory = the seed state in those coordinates; understanding = having found the right eigenbasis** **[position/inference]**. The honest shortcoming: genuine chaos → the Koopman spectrum becomes continuous → no finite set of dials can simplify it, and no amount of data will yield a clean eigenbasis **[theorem-grade picture, Mezić spectral theory]**. The crux is a question that must be answered, not one that has been answered. ## 3. SSMs (HiPPO/S4/Mamba) ↔ Koopman: surveying the bridge (literature-mining digest) *This section is the output of a dedicated literature dig: four questions, four answers.* ### Q1. What "given structure" does HiPPO trade for its eigenbasis? It trades a **forgetting measure** for it: first declare "how much each past moment matters to me" (a measure μ(t)), and the family of orthogonal polynomials under that measure is **uniquely determined** **[theorem]**. The three variants correspond to three forgetting dials: LegT (uniform sliding window → translated Legendre), LagT (exponential decay → Laguerre), LegS (uniform over the whole history → scaled Legendre). The evolution of the projection coefficients compresses into an ODE: dc/dt = Ac + Bf, with A and B derived in closed form, not learned **[theorem]**. *How to Train Your HiPPO* (2022) pushes this all the way: every SSM state can be read as "the projection coefficients of the input history onto some basis", with each basis corresponding to a measure **[empirical]**. S4 merely finds a numerically stable coordinate system for a fixed A (the DPLR decomposition) **[empirical]**. LegS has a covariance with genuine group-theoretic meaning: under a rescaling of time the output rescales in step (timescale invariance) **[theorem]**. The key difference **[inference]**: the symmetry of the calendrical code comes from **the world** (astronomical periods, given externally), whereas HiPPO's measure comes from **the agent's own choice** (I decide how to forget). Both are "free bases", but the giver is different — and that is the seed of the open question in section 4. ### Q2. What does Mamba's selectivity learn, and what does it not learn? Facts from the paper **[empirical]**: what is input-dependent is Δ, B, C; **A stays fixed** (diagonal, real). Discretization Ā = exp(ΔA): the effective decay rate at each step = a fixed spectrum × an input-determined time step. Theorem 1: at N=1 it degenerates exactly into the RNN forget gate **[theorem]**. Translated into the language of measures **[inference]**: Δ is a **knob on the rate of time**. Selectivity does not change the basis — the eigendirections are welded in place from beginning to end — it only tunes "how fast time passes" per token, with B/C determining the read/write directions. **Mamba is the first large-scale success at "learning the measure", not at "learning the basis".** Theorem-grade corroboration: *The Illusion of State in State-Space Models* (ICML 2024) proves that SSMs with diagonal/commuting transitions are all trapped in the complexity class TC⁰ and cannot even compose permutations; to cross that line, the transition matrix must depend on the input **non-diagonally** **[theorem]**. The dividing line between "learning the measure" and "learning the basis" happens to coincide with this complexity boundary **[inference]**. ### Q3. "The eigenbasis of memory decay" vs "the eigenbasis of world dynamics": what does the gap look like? - An SSM's basis is **input-generic**: Legendre polynomials do not care what world produced the signal; they are optimal for compressing an *arbitrary* signal under a given forgetting measure. State = the compressed archive of my past. - Koopman's basis is **system-specific**: the eigenfunctions live on the state space of the world, and the eigenvalues are the natural frequencies of **this particular world**. State = the world's position in its own eigencoordinates. The two coincide in exactly one case **[inference]**: when the input happens to be generated by purely point-spectrum dynamics, the optimal compression basis = the Koopman modes (this is precisely where DMD gets its legitimacy). In the general case they do not coincide: a memory basis can compress the signal without knowing the causal structure behind it. How the literature circles it (not one paper names the gap head-on; three characterize it from the side): (i) the Koopman form of a *controlled* system necessarily contains a **bilinear term** (state × input), and the Mamba recursion is purely linear in the hidden state and lacks that term; adding it improves multiplicative-memory tasks by factors of several to a hundred (2026, toy scale) **[empirical, small scale]**; (ii) *Illusion of State*: an SSM's "state", taken as a world-state, is an illusion in the sense of a complexity lower bound **[theorem]**; (iii) MamKO: running it the other way, using the Mamba architecture to generate a time-varying Koopman operator online for control — both ends are under construction, the bridge has not met in the middle **[empirical]**. ### Q4. Continuous spectrum / genuine chaos: the honest boundary Mezić's spectral picture **[theorem]**: quasiperiodic / attractor-approaching systems have a pure point spectrum, and "learning the eigenbasis" is well-posed; chaotic/mixing systems develop a **continuous spectrum** on the attractor — no countable set of eigenfunctions can span the dynamics, any finite basis is only a truncation, and the prediction horizon is nailed down by the Lyapunov time. The continuous-spectrum part corresponds to nontrivial memory effects (the Mori–Zwanzig memory kernel). The pragmatic workaround in deep Koopman: parameterize the eigenvalues as functions of the state, λ(x) — a sliding pointer instead of a fixed dial — verified only on small systems **[empirical, small scale]**. The boundary this puts on the GLM thesis **[inference]**: "you must learn the eigenbasis" is an exact thesis in a pure-point-spectrum world; in a chaotic world it is demoted to "learn a good-enough truncation and own the residual". And the residual comes back precisely in the form of a memory kernel — while an SSM's convolution kernel is by birth a machine for representing memory kernels. This may be the genuinely deep weld between the two towers. ## 4. The open question: should the forgetting measure be determined by the spectrum of the world? Two degrees of freedom on the table: **μ (the forgetting measure)** = my scoring curve for "how much I care about each past moment" (freely chosen in HiPPO); and **the world's spectrum** = the world's own list of dials (which modes turn how fast, how long correlations drag on). The question: is μ an a priori free choice, or should it be a derived quantity, (near-)uniquely determined by the world's spectrum? Is there an "optimal forgetting theorem"? **Orienting intuition: this is the dynamical version of a rate-distortion commonplace.** Memory = lossy compression of the past; and lesson one of rate-distortion theory is that the optimal code is fixed by the source statistics + the distortion measure, not by the coder's taste **[theorem]**. Swap the source for the world's dynamics and the distortion for future prediction error, and directionally the answer is "it should be" — what is open is the precise shape of the theorem. **Four anchors supporting the coupling (hard to soft):** 1. **In the linear special zone this beam already exists = balanced truncation / Hankel theory [theorem]**: given a linear world, "what an N-dimensional memory should optimally remember" is uniquely given by the SVD of the Hankel operator, with closed-form error bounds; the timescales of the optimal memory are determined directly by the world's eigenvalues. Inside a known linear world, "forgetting is determined by the world's spectrum" is a theorem, not a conjecture. 2. **Human empirics = Anderson & Schooler 1991 [empirical]**: the human power-law forgetting curve precisely matches the statistics of how often things recur in the environment — evolution has already tuned μ into a mirror of environmental statistics. 3. **S4's success can be read backwards as a natural experiment [empirical + inference]**: of the three HiPPO variants, the one that won at scale is LegS — exactly the timescale-invariant measure; and natural signals are broadly approximately scale-free (1/f spectra). S4 did not learn μ, but it happened to pick a μ matched to the world's symmetry, and then it won. 4. **The predictive information bottleneck [semi-theorem]**: "compress the past, keep only what carries information about the future" has an analytic solution in the linear-Gaussian case, and its structure follows the world's spectrum exactly; in the general case there are only variational approximations. **Three obstacles blocking a general theorem:** 1. Chaotic continuous spectra demolish the "world's spectrum" end first. The demoted version of the coupling **[inference, with hard evidence on each half]**: chaotic systems often have metastable structure (slow modes / almost-invariant sets, a spectral gap of the transfer operator), and detail beyond the Lyapunov horizon is worthless → "forget the fast continuous-spectrum part as soon as possible, remember the slow, near-point-spectrum part" — what forgetting should track is **the part of the spectrum that survives**. 2. **The chicken and the egg**: the agent does not know the world's spectrum and has to learn it; learning the spectrum relies on memory; and memory's μ is supposed to be set by the spectrum. The real theorem will not be a static formula but a **fixed point** of this loop — μ and the spectrum learned under μ's support being mutually consistent **[inference]**. 3. **Telos is not only prediction**: balanced truncation balances two Gramians — observability (which past events affect the future) and controllability (which things I can do anything about). μ should be a conspiracy between the world's spectrum and the value function **[inference]**. **The honest opposition (why it should *not* be fully determined by the world's spectrum):** in a nonstationary / black-swan world, a μ locked onto the current spectrum is the most fragile — a robust μ should keep a fatter tail than the world's correlation decay (insurance) **[inference]**; the value of a rare event lies in "how much it changes the model" (Bayesian surprise), not in its correlational weight, and a purely second-order spectrum would wash it out. The refined proposition: **what μ should track is "the predictable structure of the world + my uncertainty about that structure"; the spectrum is only the linear shadow of the former** **[inference]**. **Conjectured shape of the theorem [inference]:** given a capacity N, a stationary world-spectrum measure ρ, and a telos (a distortion measure), the distribution of timescales of the optimal forgetting measure μ\* = a rearrangement of ρ's dominant timescales, weighted by the telos; in the linear-Gaussian case it degenerates to balanced truncation; in the online version, μ and the learned spectrum are each other's fixed point, and when uncertainty is high μ automatically fattens its tail. A small falsifiable corollary: within a fixed architecture, tuning μ's decay distribution into a mirror of the autocorrelation decay of the training data should systematically improve long-range tasks (this is a deduction, not an engineering proposal). **The rhyme (flagged explicitly as a rhyme):** this question is the theorem-ification of "the compression code is a mirror of the world" — an agent's forgetting curve is the reflection of the spectrum of the world it inhabits. ## Specific questions for the reviewer 1. Is the citation of balanced truncation / Hankel theory as the "linear special-zone theorem" appropriate? Is there a stronger or more apt existing result (AAK theory, predictive state representations)? 2. Is the gap between "the memory-decay basis" and "the world-dynamics basis" really unnamed and unaddressed in the literature? Please try hard to find counterexamples (keyword hints: predictive state representations, observable operator models, the intersection of Wiener/Kalman with online basis learning). 3. Has an optimal-forgetting theorem of the form "μ and the learned spectrum are each other's fixed point" already been written down in the predictive-coding / free-energy / meta-learning literature? 4. In the chaotic case, has the demoted coupling (tracking only the metastable / slow modes) already been given a close formalization in the transfer-operator spectral-gap or Mori–Zwanzig reduction literature? 5. In section 3, Q2's characterization of Mamba as "learning the measure, not the basis", and its alignment with the TC⁰ result of *Illusion of State* and with the bilinear term of controlled Koopman — is there a hole in it we have not seen? 6. Where does the analogy on the holography side (sections 1 and 4) as a "same-family inverse problem" risk smuggling in truth? ## Literature 1. HiPPO — Gu, Dao, Ermon, Rudra, Ré 2020, arXiv:2008.07669 2. S4 — Gu, Goel, Ré 2021, arXiv:2111.00396 3. How to Train Your HiPPO — Gu et al. 2022, arXiv:2206.12037 4. Mamba — Gu & Dao 2023, arXiv:2312.00752 5. Deep Koopman — Lusch, Kutz, Brunton 2018, arXiv:1712.09707 (Nature Comm.) 6. Mezić Koopman spectral theory — arXiv:1702.07597; Annu. Rev. Fluid Mech. (annurev-fluid-011212-140652) 7. Bilinear Input Modulation for Mamba: Koopman Bilinear Forms — arXiv:2604.17221 8. The Illusion of State in State-Space Models — Merrill, Petty, Sabharwal, ICML 2024, arXiv:2404.08819 9. MamKO: Mamba-based Koopman operator — OpenReview hNjCVVm0EQ 10. Ryu & Takayanagi 2006 (hep-th/0603001); Van Raamsdonk 2010 (arXiv:1005.3035); Swingle 2012 (arXiv:0905.1317); Pastawski–Yoshida–Harlow–Preskill 2015 (arXiv:1503.06237) 11. Anderson & Schooler 1991, "Reflections of the environment in memory", Psychological Science --- ## Source: /theory/holography-koopman.zh.md (sha256:ce9760f3dab0, language: zh-Hans) # 理论笔记:全息原理、Koopman 反问题与"最优遗忘" *一条正在迭代的理论线 · 2026-07-06 · 供外部 AI reviewer 使用的完整 context* ## 给 reviewer 的 context(先读这段) 这是一条个人理论线("通用学习机 / General Learning Machine, GLM"视角)的一次迭代记录。该线的既有立场:**知识 = 生成器(种子),不是事实列表(落叶);记忆 = 动力学;遗忘 = 压缩而非删除**。本次迭代起点是一篇讨论全息原理与因陀罗网的科普文章,终点是一道我们认为文献中尚无人正面回答的 open 题。 研究纪律(请 reviewer 也遵守这个框架): - 以下全部为**理解性探索**,不含任何"据此去构建新架构"的工程承诺; - 所有论断按四档诚实度标注:**[定理]**(原文定理级)/ **[实证]**(论文实验或明确陈述级)/ **[推断]**(我们基于文献的推断)/ **[韵脚]**(结构类比,明确不主张真值传递); - 我们主动防两类错误:把漂亮的类比当证据("韵脚偷渡真值"),以及把一切强行统一进单一框架的反射。 **最有价值的 review**:指出哪个 [推断] 其实已有文献正面回答(给出处)、哪个 [韵脚] 在偷渡真值、哪个 [定理] 被我们用错了适用范围。页尾附具体 review 问题清单。 --- ## 一、全息侧:纠缠怎么"长出"几何(四块砖) **第一块:Ryu–Takayanagi 公式(2006)——词典本身。** 贝肯斯坦"黑洞信息∝表面积"原本是特例。RT 公式把它升级成通用词典:边界上任一区域的纠缠熵 = 体内某张最小曲面的面积。左边纯信息量,右边纯几何,一个等号焊死 **[定理,AdS/CFT 框架内]**。 **第二块:Van Raamsdonk 思想实验(2010)——纠缠是时空的胶水。** 把边界两半之间的纠缠调小,对应体内几何拉细,纠缠归零时时空断成两截。空间的连通性 = 纠缠的存在;"关系先于实体"的严格版 **[实证级论证]**。 **第三块:MERA 张量网络(Vidal 2007;Swingle 2012 发现其几何≈AdS)——种子怎么 unroll 出空间。** 张量网络是一张"生成流水线图":从小种子出发一层层织出量子态,每层对应一个观察尺度(重整化)。Swingle 注意到:这张流水线图自身的形状就是一片离散双曲空间,与 AdS 截面同构状 **[实证级对应,非定理]**。即:体内"多出来的一维"= 生成器 unroll 的层数 = 尺度本身。空间不是舞台,是 unroll 的痕迹——"种子生现行"的物理版 **[韵脚]**。 **第四块:HaPPY 全息纠错码(Pastawski–Yoshida–Harlow–Preskill 2015)——"局部含全体"的定理版。** 全息词典的数学结构就是一个量子纠错码:体内信息冗余编码在边界上,任何足够大的边界碎片都能重建体内深处;碎片越小看得越浅——丢的是分辨率,不是信息 **[定理,玩具模型内]**。这是"遗忘 = 压缩非删除"立场在物理侧的孪生结构 **[韵脚]**。 彩蛋:量子极值面/岛屿公式(2019–2020)在半经典层面算出了 Page 曲线——黑洞信息守恒有账可查了 **[实证/定理级进展]**。 **诚实边界:** ①以上严格数学全部生活在 AdS 宇宙,我们的宇宙是 dS(加速膨胀),推广是 open problem;②MERA≈AdS 是"惊人地像",不是同构定理;③"古代经文早已理解全息"这类方向我们明确不采纳——共享数学骨架 ≠ 源头认领。 ## 二、Koopman 反问题:没人送对称性时,本征基要自己学 **第一步:有对称性 = 本征基白送。** 傅里叶模、球谐函数都不是学出来的,是对称性(群论)送的:系统对某变换不变,对应本征基自动到手。历法系统(如六十干支)可视为把**已知**天文周期手工编码成 Z/60 的谐波本征基——同样属于"给定结构→基白送"这一格 **[数学事实 + 推断性刻画]**。 **第二步:Koopman 的移位(1931)。** 非线性动力系统,改盯"关于状态的所有观测函数"怎么演化——该演化算子**永远线性**(代价:无穷维)**[定理]**。Koopman 本征函数 = 每步只乘一个固定数的"最省心观测";找到一组,就等于找到一组坐标,让纠缠的动力学拆成互不干扰的匀速表盘。 **第三步(命门):反问题。** 对称性不在手时——真实世界几乎总是如此——这组本征基只能从轨迹数据里学:DMD → extended DMD → 深度 Koopman(Lusch–Kutz–Brunton 2018,autoencoder 端到端学线性化坐标)**[实证]**。GLM 命门的表述:**通用学习机的核心任务 = 从经验流学出让世界动力学变简单的坐标;记忆 = 该坐标下的种子态;理解 = 本征基找对了** **[立场/推断]**。 诚实短板:真混沌 → Koopman 谱变连续 → 不存在有限表盘化简,再多数据也学不出干净本征基 **[定理级图景,Mezić 谱理论]**。命门是必答题,不是已答题。 ## 三、SSM(HiPPO/S4/Mamba)↔ Koopman:桥的勘测(文献挖掘 digest) *本节为专项文献挖掘的成果,四问四答。* ### Q1. HiPPO 的本征基是拿什么"给定结构"换来的? 拿**遗忘测度**换来的:先声明"过去每一刻对我多重要"(测度 μ(t)),该测度下的正交多项式族**唯一确定** **[定理]**。三个变体对应三种遗忘表盘:LegT(滑窗均匀→平移 Legendre)、LagT(指数衰减→Laguerre)、LegS(整段历史均匀→缩放 Legendre)。投影系数演化压成 ODE:dc/dt = Ac + Bf,A、B 闭式推导、非学得 **[定理]**。《How to Train Your HiPPO》(2022)推到底:每个 SSM 状态都可解读为"输入历史在某组基上的投影系数",每组基对应一个测度 **[实证]**。S4 只是给固定 A 找了个数值稳定坐标(DPLR 分解)**[实证]**。 LegS 有真正群论意义的协变性:时间缩放下输出同步缩放(timescale 不变性)**[定理]**。关键差异 **[推断]**:历法码的对称性来自**世界**(天文周期外部给定),HiPPO 的测度来自 **agent 自己的选择**(我决定怎么遗忘)。同是"送基",送礼的人不同——这是第四节 open 题的种子。 ### Q2. Mamba 的 selectivity 学到了什么、没学到什么? 原文事实 **[实证]**:input-dependent 的是 Δ、B、C;**A 保持固定**(对角、实数)。离散化 Ā = exp(ΔA):每步有效衰减率 = 固定谱 × 输入决定的时间步长。Theorem 1:N=1 时恰好退化为 RNN 遗忘门 **[定理]**。 测度语言翻译 **[推断]**:Δ 是**时间流速旋钮**。selectivity 没换基——本征方向从头到尾焊死——只是逐 token 调"时间过多快"+ B/C 决定读写方向。**Mamba 是"学测度"的第一个大规模成功实例,不是"学基"的。** 定理级佐证:《The Illusion of State in State-Space Models》(ICML 2024)证明对角/可交换转移的 SSM 全部困在 TC⁰ 复杂度类,连置换合成都做不了;要越线,转移矩阵必须**非对角地**依赖输入 **[定理]**。"学测度 vs 学基"的分界线,恰好同时是这条复杂度边界 **[推断]**。 ### Q3. "记忆衰减的本征基" vs "世界动力学的本征基":gap 长什么样? - SSM 的基是 **input-generic** 的:Legendre 多项式不关心信号由什么世界产生,只对"给定遗忘测度下压缩任意信号"最优。状态 = 我过去的压缩包。 - Koopman 的基是 **system-specific** 的:本征函数长在世界的状态空间上,本征值是**这个世界**的固有频率。状态 = 世界在其本征坐标下的位置。 两者仅在一种情形重合 **[推断]**:输入恰好由纯点谱动力学生成时,最优压缩基 = Koopman 模态(这正是 DMD 合法性的来源)。一般情形不重合:记忆基压得动信号,不知道信号背后的因果结构。 文献侧写(无一篇正面命名该 gap,三篇侧面刻画):①受控系统的 Koopman 形式必然含**双线性项**(状态×输入),Mamba 递归对隐状态纯线性、缺此项,补上后乘性记忆任务提升数倍至百倍(2026,玩具规模)**[实证,小规模]**;②Illusion of State:SSM 的"状态"作为 world-state 是复杂度下界意义上的幻觉 **[定理]**;③MamKO:反向用 Mamba 架构在线生成时变 Koopman 算子做控制——两头动工,桥未合龙 **[实证]**。 ### Q4. 连续谱/真混沌:诚实边界 Mezić 谱理论图景 **[定理]**:准周期/趋吸引子系统有纯点谱,"学本征基"是 well-posed;混沌/mixing 系统在吸引子上出现**连续谱**——没有可数本征函数集能张成动力学,任何有限基都只是截断,预测视界被 Lyapunov 时间钉死。连续谱部分对应非平凡记忆效应(Mori–Zwanzig 记忆核)。 深度 Koopman 的务实变通:把本征值参数化为状态的函数 λ(x)——用滑动指针替代固定表盘,仅在小系统上验证 **[实证,小规模]**。 对 GLM 论题的边界 **[推断]**:"必须学本征基"在纯点谱世界是精确论题;混沌世界降格为"学够用的截断 + 承认残差"。残差恰以记忆核形式回来——而 SSM 的卷积核天生是表示记忆核的机器。这可能是两座塔真正的深层焊点。 ## 四、Open 题:遗忘测度该不该由世界的谱决定? 桌上两个自由度:**μ(遗忘测度)**= 我对过去"每一刻在乎多少"的打分曲线(HiPPO 中自由选);**世界的谱** = 世界自己的表盘清单(哪些模式转多快、相关性拖多长)。问题:μ 是先验自由选择,还是应当被世界谱(近)唯一决定的导出量?存在"最优遗忘定理"吗? **定向直觉:这是率失真常识的动力学版。** 记忆 = 对过去的有损压缩;率失真第一课:最优码由信源统计 + 失真度量定,不由编码者品味定 **[定理]**。信源换成世界动力学、失真换成未来预测误差,方向上答案是"该"——open 的是定理的精确形状。 **支持耦合的四个锚点(由硬到软):** 1. **线性特区里这根梁已存在 = 平衡截断/Hankel 理论 [定理]**:给定线性世界,"N 维记忆最优记什么"由 Hankel 算子 SVD 唯一给出,误差界闭式;最优记忆的时间尺度直接由世界本征值决定。线性已知世界内,"遗忘由世界谱决定"是定理,不是猜想。 2. **人脑实证 = Anderson & Schooler 1991 [实证]**:人类幂律遗忘曲线精确匹配环境中事物的复现概率统计——进化已把 μ 调成环境统计的镜像。 3. **S4 的成功可反读为自然实验 [实证+推断]**:三个 HiPPO 变体中大规模赢的是 LegS——恰好是 timescale 不变的测度;而自然信号普遍近似 scale-free(1/f 谱)。S4 没学 μ,但碰巧选中了与世界对称性匹配的 μ,然后赢了。 4. **预测信息瓶颈 [半定理]**:"压缩过去、只留对未来有信息量的部分"在线性高斯情形有解析解,结构完全跟世界谱走;一般情形只有变分近似。 **通用定理难产的三只拦路虎:** 1. 混沌连续谱先砸掉"世界谱"这一头。降格版耦合 **[推断,两半各有实锤]**:混沌系统常有亚稳结构(慢模式/准不变集,转移算子谱隙),Lyapunov 视界外细节无价值 → "快的连续谱部分尽快忘,慢的准点谱部分重点记"——遗忘该跟的是**谱里活得下来的那部分**。 2. **鸡生蛋**:agent 不知世界谱,得学;学谱靠记忆;记忆的 μ 又该由谱定。真定理不会是静态公式,而是这个循环的**不动点**——μ 与被 μ 支撑着学出的谱互相一致 **[推断]**。 3. **telos 不只预测**:平衡截断平衡的是两个 Gramian——可观性(过去哪些事影响未来)与可控性(我能对哪些事做什么)。μ 应是世界谱与价值函数的合谋 **[推断]**。 **诚实的反方(为什么不该全由世界谱决定):** 非平稳/黑天鹅世界里,μ 锁死当前谱最脆弱——鲁棒的 μ 应比世界相关衰减留更厚的尾(保险)**[推断]**;罕见事件价值在"改模型的量"(贝叶斯惊奇)而非相关权重,纯二阶谱会把它冲掉。精化命题:**μ 该跟的是"世界的可预测结构 + 我对该结构的不确定性";谱只是前者的线性影子** **[推断]**。 **猜想中的定理形状 [推断]:** 给定容量 N、平稳世界谱测度 ρ、telos(失真度量),最优遗忘测度 μ\* 的时间尺度分布 = ρ 的主导时间尺度经 telos 加权后的重排;线性高斯情形退化为平衡截断;在线版:μ 与所学谱互为不动点,不确定性高时 μ 自动增厚尾部。可证伪小推论:同一架构下,把 μ 的衰减分布调成训练数据自相关衰减的镜像,长程任务应系统性变好(此为推演,非工程提案)。 **韵脚(明确标注为韵脚):** 这道题是"压缩码是世界的镜像"的定理化——一个 agent 的遗忘曲线,是它所居世界的谱的倒影。 ## 给 reviewer 的具体问题 1. 平衡截断/Hankel 理论作为"线性特区定理"的引用是否恰当?有没有更强或更贴切的既有结果(如 AAK 理论、预测状态表示)? 2. "记忆衰减基 vs 世界动力学基"这个 gap,是否真的没有文献正面命名与处理?请尽力找反例(关键词提示:predictive state representations, observable operator models, Wiener/Kalman 与在线基学习的交叉)。 3. "μ 与所学谱互为不动点"式的最优遗忘定理,是否已有人在 predictive coding / free energy / meta-learning 文献中写过? 4. 混沌情形的降格版耦合(只跟亚稳/慢模式)在转移算子谱隙、Mori–Zwanzig 约化文献中是否已有接近的定式化? 5. 第三节 Q2 对 Mamba"学测度非学基"的刻画,与 Illusion of State 的 TC⁰ 结果、受控 Koopman 双线性项的对齐,是否存在我们没看到的漏洞? 6. 全息侧作为"同族反问题"的类比(第一、四节),哪里有偷渡真值的风险? ## 文献 1. HiPPO — Gu, Dao, Ermon, Rudra, Ré 2020, arXiv:2008.07669 2. S4 — Gu, Goel, Ré 2021, arXiv:2111.00396 3. How to Train Your HiPPO — Gu et al. 2022, arXiv:2206.12037 4. Mamba — Gu & Dao 2023, arXiv:2312.00752 5. Deep Koopman — Lusch, Kutz, Brunton 2018, arXiv:1712.09707 (Nature Comm.) 6. Mezić Koopman 谱理论 — arXiv:1702.07597;Annu. Rev. Fluid Mech. (annurev-fluid-011212-140652) 7. Bilinear Input Modulation for Mamba: Koopman Bilinear Forms — arXiv:2604.17221 8. The Illusion of State in State-Space Models — Merrill, Petty, Sabharwal, ICML 2024, arXiv:2404.08819 9. MamKO: Mamba-based Koopman operator — OpenReview hNjCVVm0EQ 10. Ryu & Takayanagi 2006 (hep-th/0603001);Van Raamsdonk 2010 (arXiv:1005.3035);Swingle 2012 (arXiv:0905.1317);Pastawski–Yoshida–Harlow–Preskill 2015 (arXiv:1503.06237) 11. Anderson & Schooler 1991, "Reflections of the environment in memory", Psychological Science --- ## Source: /theory/learning-the-self-boundary-log.md (sha256:b1a1f30b249d, language: en) # Stage log — the self-boundary line, in the order it happened *Companion index to [Learning where the self ends](https://machengshen.github.io/theory/learning-the-self-boundary.md) · 2026-07-27 → 2026-07-28 · **not ratified*** --- ## How to read this page This is the audit trail. One line per stage, each saying **what it overturned or established** — because the value of this line is not its conclusion (there isn't one) but the record of a claim being taken apart in public, mostly by its own authors. A plain-language summary of the whole story is Part I of the [main note](https://machengshen.github.io/theory/learning-the-self-boundary.md). The measured numbers are in Part II of the same note. Working artefacts that are not published here are named so the trail is complete, not to imply they are reachable. Two abbreviations recur. **The boundary** is the line between "me" and "the world", written **μ** in the technical text and treated as something a machine should learn rather than be handed. **Hysteresis** is the sticking property of a switch whose on-threshold and off-threshold differ — the thermostat that starts cooling at 27 and stops at 25 rather than flipping at 26. --- ## Stage 0 · The proposal (2026-07-27) | | | |---|---| | **Artefact** | Working memo: robot paradigm via a hysteretic self-boundary | | **Established** | A unified diagnosis: imitation-based and reward-based robot learning share **one** defect — the agent/environment boundary is supplied by an engineer, never derived. Proposal: learn the boundary, with closure violation as the training signal, and make the objective second-order (maintain the condition under which the boundary remains a solution). | | **Overturned** | Nothing yet. At this point the whole thing rests on one unreferenced clause: "in humans this is empirically well established." | ## Stage 1 · The executable specification (2026-07-27) | | | |---|---| | **Artefact** | `robotics-mu-experiment-spec` — loss functions, analytic bistability conditions, a MuJoCo rig, a measurement protocol with sweep-rate control. Design layer ~95% complete; zero lines run. | | **Established** | The claim became falsifiable in days rather than years: a 2–3 DoF simulated arm with a detachable tool and a breakable actuator is enough to kill it. | | **Overturned — three of the memo's own statements, by the act of writing them down precisely** | (1) the boundary cannot be a scalar per channel: the blanket condition is intrinsically three-way, and a scalar cannot express "I am the screen"; (2) "slow in, slower out" is **not free** — a symmetric double well gives threshold symmetry, not on/off time asymmetry; the memo had conflated the two; (3) **raw pixels cannot serve as channels** — the boundary is a partition over channels, so referents need stable identity. That third one is the line's largest cheat, now accounted for explicitly. | ## Stage 2 · Prior-art audit, done as an attack rather than a survey (2026-07-27) | | | |---|---| | **Artefact** | `dmbd-2502.21217-hysteresis-check` — full text of the nearest competing work, extracted and read, not inferred from the abstract. | | **Established** | The nearest work has **no hysteresis**: its label process is a per-element first-order Markov chain, inference is acausal forward-backward smoothing, and the paper itself states the assignments "quickly diffuse to a uniform stationary distribution in the absence of observed data" — the definition of monostable. Zero occurrences of hysteresis / bistability / bifurcation / saddle-node. | | **Overturned — ours** | The memo's line "FEP-family blankets have no temporal state" is **wrong for this paper**: its assignments do have their own state. The surviving distinction narrowed to exactly one axis: **monostable versus bistable.** | | **Also recorded** | Their *general* form lets label transition rates depend on the macroscopic blanket variable — the very loop bistability needs. They assume it away for tractability and list it as future work. Our wedge lands in the next cell of their roadmap: not pre-empted, but not a wide window either. | ## Stage 3 · Evidence audit — the scale narrative (2026-07-28) | | | |---|---| | **Artefact** | `scale-vs-structure-evidence-audit` (≈80% complete) | | **Established** | "Scale has already solved tool use and damage recovery" is **not supported** — because nobody has measured either. No large model has ever published a functional-transfer success table for unseen tools. The only positive damage-recovery claim is a vendor video with no numbers, whose "damage" was made in-distribution by training across 100,000 simulated morphologies. An 11-year-old paper (5 leg damage conditions + a 14-condition arm damage matrix, adaptation in under two minutes) remains the most rigorous evaluation of that capability. | | **Overturned — ours, and this is the important half** | There is equally **no** evidence that "better structure" beats scale on a robot: no paper where a structural method beats a 10× larger model on the same task and hardware with trial counts and error bars, and no demonstrated performance advantage anywhere for a *learned* self-boundary. **The other side has real scaling curves; we have a mechanism hypothesis.** The burden of proof is ours. | | **Strategic consequence** | The valuable thing to claim here is the **measurement protocol**, not the architecture. On architecture we lose to scale; on methodology the field is empty. | ## Stage 4 · Evidence audit — the empirical premise, turned on ourselves (2026-07-28) | | | |---|---| | **Artefact** | `fep-bodyschema-hysteresis-evidence-audit` (≈80% complete) | | **Overturned — the load-bearing premise of the entire line** | "In humans this is empirically well established" was **false**. Tool incorporation: not only unmeasured for hysteresis — a 2012 re-analysis of six published plus one unpublished experiment concludes tool use *"does not literally extend peripersonal space"*, with small effect sizes and low power. Rubber hand: three rounds of search found **zero** experiments sweeping a parameter up and down to compare transition points. The one supportive number (builds in ~19 s, fades in ~66 s, N=27, the authors never use the word hysteresis) is exactly the kind we had already declared non-discriminative — a monostable leaky integrator gives it for free. | | **Also overturned — ours** | "Argmax is memoryless ⟹ no hysteresis", used as our own kill-shot, was **already published** in 2022. Our increment is only the second half: attaching it to boundary maintenance. | | **Downgraded, not refuted** | The nearest famous framework is not a proven result but a framework whose premises were never shown to hold: one 2021 critique removes a central inference step with counterexamples, a 2021 reply supplies the missing premises and relaxes exact to approximate, and a 2022 review finds the required conditions hold "only for a very narrow space of parameters" and demand a perception-action symmetry unusual in living systems. | | **The same blade, turned on us — a new falsifier** | Our own design uses the *same* three-way blanket structure, so that critique cuts us too: whether a stable partition **exists** on our rig became a prerequisite question, added as stage D0 ahead of everything else. | | **And again** | "We measured a non-zero hysteresis loop" is not evidence of bistability either: any first-order lag produces a rate-dependent loop at finite sweep rate. The criterion was tightened to **loop area extrapolated to zero sweep rate stays positive**. | ## Stage 5 · Data hunt — three rounds, and one repeated mistake (2026-07-28) | | | |---|---| | **Artefacts** | `rhi-hysteresis-reanalysis-hunt` (≈85%), `trial-level-embodiment-data-hunt`, `openneuro-bids-rhi-hunt` | | **Established — why the data does not exist** | Not because nobody thought of it. Every parameterised study randomises condition order, because the standard hygiene of the field treats sequence effects as **contamination rather than signal**. The paradigm was built to erase precisely the quantity we need. | | **Established — scope of the vacuum** | Trial-level ownership data: zero hits across a public-repository search of 16 keyword sets, a second repository with 10 sets, and a **full census of 1831 neuroimaging datasets**. | | **Overturned — ours, three times in a row** | Three separate "trial-level data available" statements were believed and then opened by hand. All three contained participant × condition aggregates with **no trial index and no order column**, making the question structurally uncomputable. Twice more, a dataset whose parameter column was a genuine monotone up/down sweep — the first real sweep found in four rounds — had a response column that was blank on every row. **Do not trust the task name or the availability statement; open the columns.** | | **Overturned — a search agent's over-conclusion, corrected** | The claim "trial-level structure doesn't exist at the paradigm level" is wrong for one major lab: their published table *is* a trial-level binary ownership judgement, twelve trials per condition. The paradigm exists and is mature; the raw logs are simply never uploaded. The right next move is therefore to ask the authors, not to build a new paradigm. | | **Established — methodological templates** | The same hysteresis blade has already been applied to audiovisual synchrony; a 2018 paper supplies the operational definition of a hysteresis width plus a randomised third baseline arm; binocular rivalry is the only literature with multiple sweep rates built in, i.e. the only one that can execute the zero-rate extrapolation criterion. | | **Against us, again** | A 2014 study shows hysteresis and adaptation are independent, additive and anatomically separable ⟹ **measuring zero net hysteresis does not prove there is no latch**; the two can cancel. This has to be pre-registered, or a null result is neither a self-falsification nor an alibi. | ## Stage 6 · Human data re-analysis — two independent points, both against (2026-07-28) | | | |---|---| | **Artefacts** | `mu-boundary/human-reanalysis/` | | **Overturned — the direction of the effect** | Two independent manipulations (spatial distance, N=34, published; temporal synchrony, N=185, re-analysed by us from public data), two samples in different countries, two measurement types (experienced, expected) — **all on the adaptation/contrast side, none in the hysteresis direction.** In the larger dataset the same weak condition is rated **0.891 lower** when preceded by the strong condition, 95% CI [−1.232, −0.550], p = 6.7e-7. Convergence across independent designs is worth more than either result alone. | | **Down-weighted by us, found by reading their materials** | That larger dataset measures **expectation**, not experience — participants watched a video and reported what they would expect to feel. Its control items also move in the same direction, so demand characteristics are real; but the effect on illusion items is ~2.8× that on control items, so a genuine component exists and cannot be cleanly separated here. | | **Established — a technical fact that decides success or failure later** | On a dry run of the analysis pipeline, **the permutation null is not centred at zero but at +0.191**, for an identified reason. Testing against zero reads the sign backwards: a naive test says "no effect", while against the correct null the same data is significantly negative. Two independent null models agree. **Discipline: locate the null centre on your data before discussing significance.** | | **Established — a hard design constraint for any future experiment** | If the sign of the manipulated parameter maps directly onto the semantics of the binary response, hysteresis and adaptation are **mathematically indistinguishable**. A trial-level ownership question must therefore stay "yes/no" and must never become "which hand feels more like mine". | | **Decision** | Do not fund a self-built VR study yet: the prior has worsened, so the expected value has dropped. Priority becomes (1) request trial-level logs from the lab that has them, (2) continue on the simulation side, which does not depend on the biological premise. | ## Stage 7 · Simulation D0 — does the boundary even have an object? (2026-07-28) | | | |---|---| | **Artefact** | `mu-boundary/RESULTS-D0.md` | | **Established — the instrument works** | Positive control with a hand-installed double well: zero-rate loop area +0.36…+0.62 against a linear negative arm whose value is indistinguishable from zero. Analytic anchors reproduced to 2.6%; saddle-node scalings to 5%. Later negative results can therefore be attributed to the system, not the ruler. | | **Overturned — our own criterion** | The negative arm's bootstrap confidence interval also excludes zero, because it reflects seed noise and not the systematic bias of having few sweep-rate points. **The correct criterion is not the confidence interval but a resolution floor and the factor by which the positive arm exceeds it.** | | **Established, but nearly information-free** | A non-trivial partition satisfying the blanket condition does exist ⟹ the hysteresis discussion has an object. But a deliberately leaky calibration partition scores only marginally differently, so the dynamic range is **smaller than the across-seed spread**. The famous critique lands on our rig in mirror image: not "the condition holds only in a narrow region" but "it holds so easily that it selects nothing." | | **Overturned — our own diagnosis, in the very next round, and this is the most reusable finding** | We had diagnosed the weak result as "the rig is too weakly coupled to the world" and shipped that as a premise. **It was wrong.** The weakness was a protocol defect; after fixing it, *both* rigs pass, including the one we had condemned. The probe we blamed the rig with runs along a path that never touches the blamed component and is structurally blind to it. **We read our probe's signal-to-noise as the system's coupling strength.** A real rig defect did exist — found only by building a second, estimator-agnostic probe. | | **General lesson extracted** | Any time you argue "indicator X tells me about property Y", first show X is sensitive to Y; if you cannot, build a probe that does not share X's assumptions. | | **Measurement pitfalls recorded** | Two of them invert the specification's own worries: giving the auxiliary predictor *less* regularisation manufactures false positives (an overfit predictor *fakes* the blanket condition); and contiguous temporal splitting with non-stationary exploration fabricates conclusions outright. ⟹ this indicator is more sensitive to estimator settings and data splitting than to the partition under study, so any experiment using it must first run a deliberately leaky calibration as a positive control. | ## Stage 8 · Simulation D1 — is the mechanism reachable? (2026-07-28) | | | |---|---| | **Artefact** | `mu-boundary/RESULTS-D1-D2.md` | | **Established** | Yes: legal parameters exist, and nothing has to be pushed to an absurd magnitude. | | **Overturned — the criterion the specification called make-or-break** | Passing it turns out to be **nearly the same sentence as "the boundary flipped"**, so it screens out nothing. The specification overestimated it. | | **Overturned — the specification's literal equation** | Its stated form **cannot** have three stable settings, as mathematics rather than numerics: two requirements demand opposite signs of the same quantity. Verified numerically and pinned with a regression test that pushes the coupling to 10,000× without ever producing three. A minimal structural repair was adopted. | | **Overturned — a size estimate of ours, by 7×** | Reading the key quantity the way the specification literally says overestimates it sevenfold, because it smuggles in *channel identity* (some channels are intrinsically easy to predict) rather than the loop quantity. | | **Overturned — a rescuing number, by a representation control** | One channel showed a value that would have saved the theory (0.98). Re-expressing the same information in polar coordinates collapses it to 0.007: the large value was our basis functions being unable to compute an inverse trigonometric function, not an information gap. Representation-invariant reading: at most ~0.18, tool channels 0.002–0.045. **A control we would not have run if the number had been small.** | | **The real obstruction, newly found** | For two stable settings to exist, a cost/benefit ratio must land in a window whose relative width equals that measured quantity — **1.3–2.6% for tool channels** — while the objective function's own coefficients place every channel **2 to 420 window-widths outside**. Nothing in the objective pulls it in. ⟹ the effect does not fall out of the proposed objective; it requires a hand-set ratio the objective does not supply. | ## Stage 9 · Simulation D2 — the verdict (2026-07-28) | | | |---|---| | **Artefact** | `mu-boundary/RESULTS-D1-D2.md`, 25 tests passing | | **Established — the one surviving pillar** | With the hand-installed non-linearity **entirely removed** and only the self-reinforcing trust loop left, hysteresis appears and survives extrapolation to zero sweep rate: 24–35× the resolution floor, with both branches held 1.00/1.00 over long holds. The negative control (trust frozen constant) correctly falls back to the floor and is digit-for-digit identical to the earlier linear arm, as a regression test asserts. **Hysteresis does not need to be hand-installed.** | | **Overturned — self-specificity, by the control we had demanded of ourselves** | Give the identical treatment to the uncontrollable wind-blown ball with *its own* measured numbers ⟹ no hysteresis (the honest half). But raise the ball's one free amplitude until the loop gain matches ⟹ **every hysteresis number is identical to the arm's, digit for digit.** After normalisation the equations depend only on that gain, so the "self" quantity decides how large the free knob must be, not whether the effect occurs. ⟹ **hysteresis is a generic property of a gated slow variable; the inference "it sticks, therefore it is a self-boundary" is cancelled.** | | **Overturned — the signature prediction, twice** | "Slow in, slower out" is false here. Gate on the benefit: the ratio is 1.008 — **no asymmetry at all**, and the apparent asymmetry is an artefact of the 1-D reduction (which is asymmetric at the origin by construction). Gate on the whole drive, i.e. the mechanism actually intended: asymmetry appears but **points the wrong way** — release is faster and the threshold to leave is lower. Computable reason: precision amplifies *all* evidence about a channel, including the evidence against it. Getting "easy in, hard out" requires an extra, asymmetric evidence weighting — another hand-installed component, which is what we were avoiding. | | **Honest scope** | This stage ran on the reduced one-dimensional equation driven by measured parameters, not on the full 34-channel rig; the encoder was not trained; the intended physical parameter axis (grip stiffness) was never swept. Those are the largest open items. | --- ## Where the line stands | component of the original claim | status | |---|---| | hysteresis emerges without a hand-installed non-linearity | ✅ survives | | the hysteresis is specific to the self-boundary | ❌ generic property of gated slow variables | | release is slower than incorporation | ❌ reversed, in both implementations | | it falls out of the proposed objective without fine-tuning | ❌ needs a ~2% hand-set ratio | Three of four down, **every one of them felled by our own simulator or our own re-analysis**. The human-side motivation is not merely unproven — the two usable data points point the wrong way. **The question that replaced it**, and the reason this line stays open: hysteresis is cheap, so the real question is *what would make a boundary's hysteresis be about that boundary itself* — rather than a property it shares with a ball blowing in the wind. **Not ratified. Nothing here is a settled result.** Full narrative and all measured numbers: [Learning where the self ends](https://machengshen.github.io/theory/learning-the-self-boundary.md). --- ## Source: /theory/learning-the-self-boundary-log.zh.md (sha256:05b2b0f1f56c, language: zh-Hans) # 阶段日志 —— 自我边界这条线,按发生顺序 *[学习"我到哪里为止"](https://machengshen.github.io/theory/learning-the-self-boundary.zh.md) 的配套索引 · 2026-07-27 → 2026-07-28 · **未 ratify*** --- ## 怎么读这一页 这是审计留痕。每个阶段一条,说清它**推翻或确立了什么** —— 因为这条线的价值不在结论(它没有结论),而在**一个主张被公开拆解的记录,而且大部分是被它自己的作者拆的**。 整个故事的人话版本是[主 note](https://machengshen.github.io/theory/learning-the-self-boundary.zh.md) 的第一部分;实测数字在同一篇的第二部分。没有公开的工作产物在这里**只列名字**,是为了留痕完整,不是暗示它们可访问。 两个反复出现的简称。**边界**指"我"和"世界"之间的那条线,技术正文里写作 **μ**,主张是机器应该**学**出它而不是被人递给它。**滞回**指一个开关的"开门槛"和"关门槛"不一样从而卡住的性质 —— 27 度才制冷、25 度才停,而不是到 26 就翻转。 --- ## 阶段 0 · 提案(2026-07-27) | | | |---|---| | **产出** | 工作 memo:以滞回的自我边界替代当下机器人范式 | | **确立** | 统一诊断:模仿式与奖励式机器人学习共享**同一个**缺陷 —— agent/环境边界由工程师供给,从不被推导。提案:学出边界,以**闭包违反**当训练信号,目标做成二阶(维持"边界仍是解"的那个条件)。 | | **推翻** | 暂无。此刻整套东西压在一句无出处的话上:"这在人身上经验上是确凿的"。 | ## 阶段 1 · 可开工规格(2026-07-27) | | | |---|---| | **产出** | `robotics-mu-experiment-spec` —— 损失函数、双稳解析条件、MuJoCo rig、含扫速控制的测量协议。设计层约 95%,零行代码跑过。 | | **确立** | 主张变成**周级而非年级**可证伪:一条 2–3 DoF 仿真臂 + 可拆工具 + 可断执行器就足以杀死它。 | | **推翻 —— memo 自己的三条,靠"把它写精确"这个动作** | (1) 边界不能是每通道一个标量:blanket 条件本质三元,标量表达不了"我是屏障";(2) "并入慢释放更慢"**不是免费的** —— 对称双阱只给关于倾斜对称的**阈值**,不自动产生 on/off **时间**非对称,原文把两者混为一谈;(3) **原始像素不能当通道** —— 边界是通道上的划分,指称物身份必须稳定。第三条是本线最大的作弊点,现已显式记账。 | ## 阶段 2 · Prior art 审计,当成"打"而不是"综述"(2026-07-27) | | | |---|---| | **产出** | `dmbd-2502.21217-hysteresis-check` —— 把最近的竞争工作正文抽出来读完,不是从摘要推断。 | | **确立** | 最近的那份工作**明确无滞回**:标签过程是每元素一条一阶马尔可夫链,推断是 acausal 前后向平滑,且论文自陈"在没有观测数据时归属变量迅速扩散到均匀平稳分布" = 单稳的定义。全文 hysteresis / bistability / bifurcation / saddle-node 零命中。 | | **推翻 —— 我们的** | memo 那句"FEP 系的 blanket 无时间状态"**对这篇是错的**:它的归属确实有自身状态。存活的区别只剩**一条轴:单稳 vs 双稳**。 | | **同时记账** | 他们的**一般形式**让标签转移率依赖宏观 blanket 变量 —— 正是双稳所需的那个闭环。他们为 tractability 假设掉了它,并列入 future work。**我们的 wedge 正落在他们 roadmap 的下一格**:不算被抢先,但时间窗不宽。 | ## 阶段 3 · 证据审计 —— 规模叙事(2026-07-28) | | | |---|---| | **产出** | `scale-vs-structure-evidence-audit`(约 80%) | | **确立** | "规模已解决工具使用 / 损伤恢复"**不成立** —— 因为两项**都没人测过**。没有任何大模型发布过未见工具的功能迁移成功率表。损伤恢复唯一的正面主张是零数字的厂商视频,其"损伤"因训练覆盖 10 万种仿真形态而被做成分布内。一篇 11 年前的论文(腿式 5 种损伤 + 机械臂 14 种损伤矩阵,<2 分钟在线适应)至今仍是该能力最严谨的评测。 | | **推翻 —— 我们的,而且这才是重要的一半** | 同样**没有**证据表明"更好的结构"在机器人上赢过规模:找不到任何一篇让结构方法在同任务同硬件上击败大 10× 模型、且带试验数与误差棒的论文;也找不到任何"学出来的自我边界"的性能优势实证。**对方手里有真实的规模曲线,我们手里只有机制假说。** 举证责任在我们。 | | **战略后果** | 值得抢的是**测量协议**,不是架构。架构上我们打不过规模,方法学上对方是空的。 | ## 阶段 4 · 证据审计 —— 经验前提,刀口转向自己(2026-07-28) | | | |---|---| | **产出** | `fep-bodyschema-hysteresis-evidence-audit`(约 80%) | | **推翻 —— 整条线的承重前提** | "这在人身上经验上是确凿的"**是假的**。工具并入:不但没有滞回测量,现象本身在 2012 年就被一篇专门的再分析质疑 —— 再分析 6 篇已发表 + 1 篇未发表,结论 tool use *"does not literally extend peripersonal space"*,效应量小、功效低。橡胶手:三轮检索**零命中**任何上下行扫描比较转换点的实验。唯一支持性数字(约 19 s 建立、66 s 消退,N=27,作者全文不用 hysteresis 一词)恰是我们自己早已声明不算数的那种 —— 一个单稳的泄漏积分器在撤掉驱动时天然给出。 | | **同时推翻 —— 我们的** | "argmax 无记忆 ⟹ 不可能滞回",我们自己反复用的杀法,**其前半句已在 2022 年公开发表**。我们的增量只剩后半句(接到边界维持上)。 | | **降级而非反驳** | 最近的那个著名框架不是被证明的结果,而是**前提未被验证满足的框架**:2021 年一篇技术批评带反例打掉了中间一步推断,2021 年的回应补上缺失前提并把 exact 放松成 approximate,2022 年的综述指出所需条件"只在极窄参数区成立",且要求活体几乎不具备的感知-行动对称性。 | | **同刀回砍 —— 新增一条 falsifier** | 我们自己的设计用的是**同一个**三分 blanket 结构 ⟹ 那一击同样砍我们:**先证明我们的 rig 上存在满足条件的稳定划分**,作为 D0 排在一切之前。 | | **又一次回砍** | "我们测到非零滞回环"也不是双稳的证据:任何一阶滞后在有限扫速下都产生率依赖的环。判据被收紧为**环面积外推至零扫速仍为正**。 | ## 阶段 5 · 数据搜寻 —— 三轮,以及重复了三次的一个错误(2026-07-28) | | | |---|---| | **产出** | `rhi-hysteresis-reanalysis-hunt`(约 85%)、`trial-level-embodiment-data-hunt`、`openneuro-bids-rhi-hunt` | | **确立 —— 为什么数据不存在** | 不是没人想到。所有参数化实验一律随机化条件顺序,因为该领域的标准卫生规范把**序列效应当污染而非信号**。范式的建法就是为了抹掉我们要的那个量。 | | **确立 —— 真空的范围** | 逐试次所有权数据:一个公开仓库 16 组关键词 + 另一个仓库 10 组 + **1831 个神经影像数据集全量普查**,全部**零命中**。 | | **推翻 —— 我们的,连着三次** | 三份"逐试次数据公开"的声明被相信,然后被亲手打开。三份都是被试 × 条件的聚合表,**没有试次序号、没有顺序列**,问题在结构上不可计算。同样的坑又踩两次:某数据集的参数列是真正的单调上下行扫描 —— 四轮里第一次见到 —— 而响应列**每一行都是空的**。**别信任务名、别信可用性声明,打开列。** | | **推翻 —— 搜寻 agent 的过度结论,已纠正** | "逐试次结构在范式层面就不存在"这个判断,**对某个主要实验室是错的**:他们已发表的表格**就是**逐试次二元所有权判定,每条件 12 试次。范式存在且成熟,只是原始日志从不上传。⟹ 正确的下一步是**向作者索要**,不是自建范式。 | | **确立 —— 方法学模板** | 同一把滞回的刀已经被用在视听同步性上;2018 年一篇给出滞回宽度的操作定义 + 第三条随机序列基线臂;双眼竞争是唯一内建多扫速的文献,也就是唯一能执行零扫速外推判据的那条。 | | **又一条对我们不利** | 2014 年一项研究证明迟滞与适应是**独立可加、脑区分离**的两个成分 ⟹ **测到零净滞回不能证明没有 latch**(两者可抵消)。这必须写进预注册,否则阴性结果既不是自我证伪也不能当免罪符。 | ## 阶段 6 · 人类数据重分析 —— 两个独立点,都不利(2026-07-28) | | | |---|---| | **产出** | `mu-boundary/human-reanalysis/` | | **推翻 —— 效应的方向** | 两个独立操纵(空间距离,N=34,已发表;时间同步性,N=185,我们在其公开数据上重分析)、两个不同国家的样本、两种测量(亲历 / 期望)—— **全部落在适应/对比侧,无一在迟滞方向**。较大那份数据里,同一弱条件在强条件之后被评分**低 0.891**,95%CI [−1.232, −0.550],p = 6.7e-7。多个独立设计的收敛,比其中任何单条更有份量。 | | **我们自己给的降权(靠读他们的材料读出来)** | 较大那份量的是**期望**不是亲历 —— 被试只看视频、答"预期会体验到多少"。其控制条目也同向移动 ⟹ 需求特征确实存在;但 illusion 条目上的效应约为控制条目的 2.8 倍 ⟹ 真实成分存在,只是本数据无法干净分离。 | | **确立 —— 一条会决定后续成败的技术事实** | 管线 dry-run 上,**置换零分布的中心不在 0 而在 +0.191**,机制已定位。对 0 做检验会**把符号读反**:朴素检验说"无效应",对上正确零模型后同一份数据显著为负。两个独立零模型一致。**纪律:先在你的数据上找到零分布中心,再谈显著性。** | | **确立 —— 未来自建实验的硬约束** | 若被操纵参数的符号与二元反应的语义直接对应,迟滞与适应**在数学上不可分辨**。所以逐试次所有权问题必须保持"是/否",绝不能变成"哪只手更像我的"。 | | **决策** | 暂不投自建 VR:先验变差,期望价值下降。优先级改为 ①向持有逐试次日志的实验室索要 ②继续仿真侧(不依赖生物学前提)。 | ## 阶段 7 · 仿真 D0 —— 边界到底有没有对象?(2026-07-28) | | | |---|---| | **产出** | `mu-boundary/RESULTS-D0.md` | | **确立 —— 仪器有效** | 手装双阱的阳性对照:零扫速环面积 +0.36…+0.62,而线性阴性档与 0 不可区分。解析锚点复现到 2.6%,鞍结标度到 5%。⟹ 后续阴性结果可归因于系统而非尺子。 | | **推翻 —— 我们自己的判据** | 阴性档的 bootstrap 置信区间**也不含 0**,因为它只反映种子噪声,不反映"扫速点太少"的系统偏差。**正确判据不是置信区间,而是分辨底 + 阳性档高出它多少倍。** | | **确立,但几乎无信息量** | 满足 blanket 条件的非平凡划分确实存在 ⟹ 滞回讨论有对象。但故意泄漏的标定划分得分只差一点点,**动态范围小于跨种子波动**。那条著名批评在我们 rig 上以镜像形态出现:不是"条件只在极窄区域成立",而是"**太容易成立以至于选不出东西**"。 | | **推翻 —— 我们自己的诊断,就在下一轮,而且这是最可复用的一条** | 我们把弱结果诊断为"rig 与世界耦合太弱",并把它当前提发下去。**错了。** 弱是**协议缺陷**;修好后**两个 rig 都通过**,包括被我们判死刑的那个。而我们用来指控 rig 的那根探针,走的路径根本不经过被指控的部件,对它**结构上全盲**。**我们把探针的信噪比读成了系统的耦合强度。** rig 的真实缺陷确实存在 —— 是靠造第二根、估计器无关的探针才找到的。 | | **提炼出的通用教训** | 任何"指标 X 告诉我性质 Y"的推理,先证明 X 对 Y 敏感;证不了就造一根不共享 X 假设的探针。 | | **记账的测量学坑** | 其中两条与规格自己的担心方向相反:给辅助预测器**更少**正则化反而制造假阳性(过拟合的预测器会**假装** blanket 成立);连续时序切分 + 非平稳探索会**直接伪造结论**。⟹ 该指标对估计器设置与数据切分比对被研究的划分更敏感,任何用它下结论的实验必须先跑故意泄漏的标定当阳性对照。 | ## 阶段 8 · 仿真 D1 —— 机制可达吗(2026-07-28) | | | |---|---| | **产出** | `mu-boundary/RESULTS-D1-D2.md` | | **确立** | 是:存在合法参数,不需要把任何东西推到荒谬量级。 | | **推翻 —— 规格称为"生死"的那条判据** | 通过它几乎**等价于"边界确实翻转了"**,所以它筛不掉任何东西。规格高估了它。 | | **推翻 —— 规格的字面方程** | 它写下的形式**不可能**有三个稳定档位,这是数学不是数值:两个要求逼迫同一个量取相反符号。已数值验证,并用一条把耦合推到一万倍仍出不了 3 的回归测试钉死。已采用一个最小的结构性修补。 | | **推翻 —— 我们自己的一个量级估计,差 7 倍** | 按规格字面读那个关键量会高估七倍,因为它偷偷混进了**通道身份**(有些通道本来就好预测),而不是回路里那个量。 | | **推翻 —— 一个救命数字,被表示对照打掉** | 有一条通道给出足以救活理论的数值(0.98)。把同样的信息换成极坐标表示后塌到 0.007:大数值是**我们的基函数解不出反三角**,不是信息缺口。表示不变的读数:最大约 0.18,工具通道 0.002–0.045。**如果那个数字本来就小,我们大概不会跑这个对照。** | | **新找到的真瓶颈** | 要有两个稳定档位,一个成本/收益比必须落进一个相对宽度等于该实测量的窗口 —— **工具通道只有 1.3–2.6%** —— 而目标函数自己的系数把每条通道放在窗口外 **2 到 420 个窗口宽**处,且目标里**没有任何机制**在把它往里拉。⟹ 效应不是从提出的目标函数里长出来的,它要求一个该目标不提供的手调比值。 | ## 阶段 9 · 仿真 D2 —— 判决(2026-07-28) | | | |---|---| | **产出** | `mu-boundary/RESULTS-D1-D2.md`,25 项测试通过 | | **确立 —— 唯一活下来的支柱** | 把手装非线性**整个删掉**、只留自我强化的信任回路,滞回出现,并且**通过零扫速外推**:达分辨底的 24–35 倍,两支在长时间保持下 1.00/1.00。阴性对照(信任冻成常数)正确地落回分辨底,且与早先的线性档**逐位相同**(有回归测试断言)。**滞回不需要手装。** | | **推翻 —— 自我特异性,被我们自己要求的对照打掉** | 把完全相同的处理给那个不可控的风吹球,用**它自己**的实测数字 ⟹ 不出现滞回(诚实的一半)。但把球那个唯一的自由幅度拧到环增益相同 ⟹ **滞回的每一个数字与手臂逐位相同**。归一化后方程只依赖那个增益,所以"自我"那个量只决定自由旋钮要拧多大,不决定效应是否出现。⟹ **滞回是带门控慢变量的通用性质;"它卡住了所以这是自我边界"这一步推理被取消。** | | **推翻 —— 那条签名预测,两次** | "进来难出去更难"在这里是假的。门乘在收益上:比值 1.008 —— **完全无非对称**,而我们一度以为看到的非对称是 1-D 约化的假象(该约化在原点处本就不对称)。门乘在整个驱动上(即真正想要的机制):非对称出现了但**方向相反** —— 释放更快,离开的阈值更低。可算原因:精度放大关于一条通道的**全部**证据,**包括反对它的证据**。要得到"进来容易出去难",必须额外假设一条不对称的证据加权 —— 又一件手装的东西,而那正是我们要避开的。 | | **诚实的范围声明** | 本阶段跑在由实测参数驱动的**一维约化方程**上,不是 34 通道全 rig;编码器未训;规格设定的物理主参数轴(握持刚度)从未被扫过。这些是最大的未闭合项。 | --- ## 这条线现在的位置 | 原主张的成分 | 状态 | |---|---| | 滞回不靠手装非线性即可涌现 | ✅ 活 | | 该滞回对自我边界有特异性 | ❌ 带门控慢变量的通用性质 | | 释放比并入更慢 | ❌ 两种实现下都反了 | | 它从提出的目标函数里自然长出,不需 fine-tune | ❌ 需约 2% 的手调比值 | 四条倒三条,**每一条都是被我们自己的仿真器或自己的重分析打倒的**。人类侧的动机不只是"未被证明" —— 两个可用数据点都指向相反方向。 **取代它的问题**,也是这条线还开着的理由:滞回很廉价,所以真问题是 —— **什么让一条边界的滞回是关于这条边界本身的**,而不是它和一个随风飘的球共享的性质。 **未 ratify。这里没有任何东西是定论。** 完整叙事与全部实测数字:[学习"我到哪里为止"](https://machengshen.github.io/theory/learning-the-self-boundary.zh.md)。 --- ## Source: /theory/learning-the-self-boundary.md (sha256:99fad00bee87, language: en) # Learning where the self ends — a research line that broke three of its own four pillars in one day *Notes from one line of work · 2026-07-27 → 2026-07-28 · **not ratified, not a result*** --- ## Part I — For readers who are not in this field This half has no formulas, no Greek letters and no paper numbers. Every technical word is explained where it first appears. The second half has all the numbers, commands, citations and falsification conditions, and is written for machines and specialists. If you only read Part I, you should still be able to tell someone else what we found. ### The one-paragraph version We thought we had found the thing that is wrong with how robots are built today: nobody ever lets the robot work out **where it ends and the world begins**. We proposed that a machine should learn that boundary rather than be handed it, and we said the signature of a *real* boundary is that it **sticks** — it is hard to move once it has settled. Over one long day we attacked our own proposal from four directions with our own simulator and our own re-analysis of other people's public data. **Three of the four supporting claims died, and every one of them was killed by our own instruments.** What survived is thin but real: the sticking behaviour *can* appear on its own, without us secretly installing it. What replaced the original claim is a sharper and harder question, which is written at the end. ### 1 · What we believed Two families of robot learning dominate right now. - One family learns by **copying**: show the machine thousands of human demonstrations and have it imitate the mapping from what it sees to what it does. - The other learns by **scoring**: define a number that says how well the robot is doing, and have it search for the behaviour that pushes that number up. Both are given, for free, an answer to a question nobody asks them out loud: *which parts of the world are me?* An engineer answers it by hand when they decide what counts as a joint, which motors exist, and what the machine's "action" is. The machine never has to ask. Our claim was that this hand-drawn line is not a detail, it is **the** bottleneck. A robot that could work out its own boundary would do two things that today's robots do badly: - pick up an unfamiliar tool and treat it as part of its arm, rather than as an object it has memorised; - lose a motor mid-task and *redraw the line* — "that limb is not mine any more" — instead of needing to be retrained. ### 2 · Why we thought it was more than a slogan Because in humans the boundary appears to have a very particular mechanical character: it **sticks**. > **Sticking — the word for it is "hysteresis".** Think of a thermostat set to 26°C. It does not switch on at 26.1 and off at 25.9. It waits until the room reaches 27 to start cooling, and until 25 to stop. The temperature at which it turns on and the temperature at which it turns off are *different*, and the gap between them is what stops it chattering on and off every few seconds. That gap is hysteresis. A system with hysteresis remembers which side it came from. Everyone has felt the human version. Hold a stick long enough and you start feeling the table through the stick's tip rather than feeling the stick in your palm. Sit on your arm until it goes numb and for a few seconds it is not yours; it is a foreign object in the bed. And there is a classic laboratory illusion — a rubber hand stroked in time with your hidden real hand until you feel the rubber one is yours — which seems to show the boundary snapping across in one piece. If the boundary really sticks, then a whole class of machine designs is ruled out at a stroke. Any mechanism whose rule is "at each moment, compute the best boundary given what I see right now" **cannot** stick, because recomputing the best answer has no memory of which answer you held a moment ago. That would mean the right implementation is not an optimiser at all but something closer to a physical system with inertia. That was the structural argument, and it is the kind of argument that is cheap to say and expensive to earn. ### 3 · How we went after it We wrote down, in advance, four ways the idea could die, and then we spent the day trying to trigger them ourselves. In parallel we ran two campaigns: - **The literature campaign.** Not "has anyone had this idea before" — that question is nearly useless. The question we actually asked was: *when someone says they showed this, what did their experiment support, and where does the support stop?* We hold to a rule that "somebody already did it" is a lead, never a verdict; most papers are written to be published, not to be true, so the honest move is to go and attack their evidence rather than politely stand aside. - **The simulation campaign.** A simple simulated arm with a detachable tool, a breakable motor, and a wind-blown ball that the arm cannot control — so that "not me" has something to point at. ### 4 · How it fell — in the order it happened **(a) The "obviously true" fact turned out to be neither obvious nor established.** Our whole motivation rested on the sentence "in humans this is empirically well established". It is not. Nobody has ever run the decisive experiment — sweeping a knob up and then down and checking whether the boundary flips at different points on the way up and the way down. Worse, there is a specialist re-analysis of the tool-use literature whose conclusion is that the tool effect *may not exist at all* at the strength people assume, with small effect sizes and low statistical power. And the one supportive number we could find — the illusion builds in about twenty seconds and fades in about a minute — is exactly the kind of number we had already declared meaningless, because a perfectly ordinary system with no stickiness gives you that automatically when you switch its input off. **(b) When we went looking for the decisive data, we found the opposite sign.** Two independent public datasets, different countries, different manipulations, hundreds of participants between them, both let us ask a weaker but related question: does what you experienced a moment ago pull your next judgement *toward* it (sticking) or *away* from it (contrast)? Both said **away**. That is not a clean refutation — both used end-of-block questionnaires rather than moment-by-moment answers, and one of them measured what people *expected* to feel rather than what they felt — but two independent things pointing the same wrong way is worse for us than one. **(c) Then, good news: in our own simulator, the sticking appeared without us installing it.** This matters because the cheap way to get stickiness is to write it into the equations by hand — the equivalent of writing the answer at the bottom of the page. We removed the hand-installed part entirely and left only the mechanism we were actually proposing: *the more a channel is treated as part of me, the better it gets predicted; the better it gets predicted, the more it is trusted; the more it is trusted, the more it is treated as part of me.* A loop that reinforces itself. With that loop and nothing else, the boundary stuck, and it kept sticking even when we slowed the experiment down by a factor of ten thousand — which is the test that separates real stickiness from the fake kind you get just from things taking time to settle. **This is the one pillar still standing.** **(d) Then the control experiment we had demanded of ourselves went off in our face.** We had written down: give the identical treatment to something that is obviously *not* part of the robot — the wind-blown ball — and check that it does not stick. Half of that worked: with the ball's own measured numbers, no stickiness. But the other half is fatal. There is one free knob in the equations that nobody's data pins down. Turn that knob up on the ball, and the ball's boundary sticks **exactly as much as the arm's — identical to every digit we printed**. There is a reason, and it is provable rather than accidental: after rescaling, the equations only care about one combined quantity, and the "self" part only decides how far you have to turn the knob, not whether the effect appears at all. > Said plainly: **stickiness is not a signature of selfhood. It is a generic property of any slowly-changing quantity with a gate on it.** Our arm and a wind-blown ball are indistinguishable on this measure; the only difference is a factor of seven in a knob that is free to be set. **(e) And the one sharp, specific prediction came out backwards. Twice.** The prediction was intuitive and it was the reason to prefer our mechanism to a hand-built switch: *hard to adopt, harder to let go* — a tool should be slow to become part of you and slower still to stop being part of you. In the first implementation there was no asymmetry at all (adoption and release took the same time to three decimal places), and the asymmetry we thought we saw turned out to be an artefact of how we had compressed three quantities into one. So we implemented it the other way, the way that ought to produce the asymmetry — and the asymmetry appeared **pointing the wrong way**: letting go was *faster* than adopting, and it took *less* provocation to leave than to join. The reason is computable, and in hindsight obvious. Our mechanism amplifies the *evidence* about a channel. Evidence is not partisan. Amplifying the evidence about your own hand amplifies the evidence that it is **not** your hand just as much. To get "easy in, hard out" you would have to additionally assume that trust only amplifies evidence in its own favour — and that assumption is another hand-installed thing, which is what we were trying to avoid. **(f) Finally: even the surviving stickiness needs a dial set to within about two percent.** For the loop to have two stable settings rather than one, a cost and a benefit inside it have to cancel to a precision of roughly two percent. Our own objective function — the thing that was supposed to make all this emerge naturally — places every single channel *outside* that window, by anywhere from two to four hundred window-widths. So "it can emerge" and "it will emerge" are separated by an argument nobody has made yet. ### 5 · Three mistakes that were ours, not the world's Worth writing down because each one is a repeatable species of error, not bad luck. 1. **We measured the wrong object and then shipped the misdiagnosis downstream as a premise.** An early run gave a weak result. We diagnosed the cause as "our simulated arm is too weakly coupled to the world", and handed that diagnosis to the next stage as an established fact. It was wrong. The weak reading came from a flaw in the measurement protocol, and the probe we were using ran along a path that never touched the part of the rig we blamed — so it was structurally incapable of detecting the problem we attributed to it. We had read our *probe's* signal-to-noise as the *system's* coupling strength. The general lesson: any time you argue "indicator X tells me about property Y", first check that X is sensitive to Y at all, and if you cannot show it, build a second probe that does not share X's assumptions. (A real coupling flaw *did* exist — it was found by a different probe, and the first indicator was blind to it.) 2. **We believed data availability statements three separate times.** Three published studies advertised "trial-level data available". We downloaded all three and opened every file. All three contained per-participant summary tables with no trial index and no ordering column, which makes the question we wanted to ask structurally uncomputable. Twice more the same shape: a dataset whose parameter column looked exactly like the sweep we needed, with a response column that was blank for every row. **Do not trust the task name or the abstract; open the columns.** 3. **We let "empirically well established" stand as a premise without a citation.** That single unreferenced clause was the load-bearing member under the entire architectural conclusion. Removing it left the conclusion hanging in the air, and it should never have been allowed in. ### 6 · What is left, and the better question underneath What survived: **stickiness can emerge from a self-reinforcing trust loop without being hand-installed**, and it survives the slow-down test. That is a real, if modest, positive result. What replaced the original claim is sharper, and we would not have found it without the failures: > **Stickiness is cheap. Almost any slowly-changing quantity with a gate on it has it. So the question was never "does the boundary stick?" — it is: *what would make a boundary's stickiness be about that boundary itself, rather than being a generic property it happens to share with a ball blowing in the wind?*** That is a harder question, and it is a better one, and it is the one we would work on next. The honest state of this line is: **an open question with most of its original motivation removed, kept alive by one positive simulation result and one reformulated problem.** Nothing here is settled. Nothing here has been ratified. --- --- ## Part II — Technical record Everything below is measured or cited. Numbers that were not measured are marked as such. Claims are tagged **[measured]** (we ran it), **[cited]** (someone else's published number, verified against the source text), **[inference]** (our synthesis), **[open]** (not done). Cognitive state of the whole note: **speculative**, with the specific sub-claims graded in §7. ### 0 · Notation - **μ** — the agent/environment boundary, treated as a first-class learned variable: a partition of all sensorimotor channels into (self, blanket, world). "Blanket" is the screening surface: world influences self only through it. - **Hysteresis** — bistability of μ under a swept parameter, with distinct up-sweep and down-sweep transition points. The operational criterion adopted here (see §5) is **loop area extrapolated to zero sweep rate remains > 0** (rate-independent), *not* merely "non-zero loop area at some finite sweep rate". - **Δπ / s** — the precision swing: how much better a channel is predicted when it is inside the boundary versus outside, expressed as the relative swing `s = 1 − ε_in/ε_out`. - **G** — the loop gain of the precision→assignment→precision feedback, `G = κ·Δπ·(g−c)/(4γτ_π)`. ### 1 · The original proposal Unified diagnosis: vision-language-action models (VLA) and reinforcement learning (RL) share **one** defect — μ is exogenous. - VLA: the action space is hand-drawn by an engineer (e.g. a 27-dimensional joint action space with noop padding — a hand-drawn μ frozen into a vocabulary; the whole cross-embodiment codec is a patch for that hand-drawn μ). The model is never required to answer "what am I". - RL: scalar objective under argmax. Argmax is memoryless ⟹ structurally cannot exhibit hysteresis ⟹ cannot maintain a boundary. A first-order homeostat is a thermostat. Proposal: learn state / learn μ, not policy. 1. Training signal = **closure violation**, not reward and not action imitation. `z` is a state iff the dynamics close on `z`. Maximise forgetting subject to closure. No rewards, no demonstrations. 2. μ as a first-class variable: partition all channels into (self, blanket, world); criteria = self side internally predictable and controllable, world side reachable only via blanket, partition stable under perturbation. 3. The objective is **second-order**: not "maximise within μ" but "maintain the condition under which μ remains a solution" — self-seal. Executable specification (loss functions, analytic bistability conditions, MuJoCo rig, measurement protocol with sweep-rate control): `robotics-mu-experiment-spec-20260727` (working document, not published here). ### 2 · Prior art, audited rather than surveyed Method note: the governing rule for this line is that **"someone already did this" is evidence, not a verdict**. The correct output of a prior-art pass is not "who owns which sentence" but "how far does their evidence actually reach, and where does rhetoric begin". The same blade is then turned on our own claims; a double standard here degrades into claim-defence. - **Crutchfield causal states / computational mechanics (1989); PSR** — owns "state = minimal sufficient statistic for prediction", i.e. closure itself. **[cited]** - **Klyubin & Polani, empowerment (2005)** — owns "derive an intrinsic objective from channel structure". **[cited]** - **Maturana & Varela, autopoiesis** — owns "the self makes its own boundary" (no algorithm). **[cited]** - **Varela, Maturana & Uribe 1974** (*BioSystems* 5(4):187–196) — **this is the honest closest ancestor of our self-seal claim, and we found it late.** Their sentence is: *"The necessary feature is the presence of a boundary which is produced by a dynamics such that the boundary creates the conditions required for this dynamics."* That is our state→coupling loop, stated in words in 1974, and not only in words: their tessellation automaton mechanically instantiates it (bonded link segments gate which species may move where, so the boundary reconfigures the coupling that sustains it), with a real adjustable control parameter (disintegration probability `Pd`) and a quantitative viability threshold (*"less than about .01 per time step is required in order to achieve any viable structure at all"*). **What it does not have:** any parameter sweep, any bifurcation analysis, any bistability, any hysteresis, any gain, and no continuous-time dynamics. So the *proposition* is theirs; what remains ours is the claim that this loop has a **measurable gain with a bifurcation structure**, and the protocol for measuring it. **[cited]** - **Piedrafita, Montero, Morán, Cárdenas & Cornish-Bowden 2010** (*PLoS Comput. Biol.* 6(8):e1000872) — technically the closest thing in the closure literature to what §5 does: real mass-action ODEs, a saddle-node at `k = 0.367`, and an explicit hysteresis prediction. Recorded because it is the strongest candidate for "you were pre-empted". **It is not the same object, and they say so in the same paper**: their bistable variable is metabolic on/off, their control parameter is an *exogenous* catalyst decay rate rather than a feedback gain on a state→coupling loop, their loop is admittedly incomplete (irreversible collapse rather than a closed cycle), and they explicitly bracket structural/boundary closure — *"This aspect was given almost no attention by Rosen, and we shall not discuss it further here."* A lead, not pre-emption. **[cited]** - **Two cautions recorded against ourselves, from the same pass.** (i) That hysteresis loop width grows with positive-feedback gain and vanishes at a critical gain is **textbook** — the cusp / saddle-node normal form (canonically Angeli, Ferrell & Sontag, *PNAS* 2004). Our novelty cannot be the scaling law; it can only be the *mapping* (that the gain is generated by boundary → precision/gating → statistics → boundary, and that the bistable variable is a partition of sensorimotor channels). (ii) The generic normal form loses bistability at a **finite** critical gain `g* > 0`, not asymptotically as `g → 0`. Any future gain-sweep falsifier must therefore be stated as "loop width → 0 at a finite `g*`", and our own equations checked for whether they have a fold at finite `g` at all. The criterion actually used in §5 is a *sweep-rate* extrapolation, not a gain extrapolation, so nothing published here depends on this — but the distinction is easy to blur and is recorded so that it is not blurred later. **[open]** - **Luhmann's re-entry — a keyword collision to disarm, not an ancestor.** Luhmann's corpus contains "bi-stable", "oscillation", "memory function" and "eigenvalue" within a few pages, which makes it look adjacent. It is not: his bistability is *two-valued code alternation* (marked/unmarked, Günther transjunction), there is **no control parameter anywhere in his re-entry corpus** and therefore nothing to sweep, and he states his own register outright — re-entry cites *"constraints of a mathematical calculus restricted to arithmetic and algebra"*. He also puts our object on the other side of his own distinction: organisms have *"material boundaries in space"*, whereas for psychic and social systems *"boundaries are therefore not material artifacts but forms with two sides"*. μ is a partition of material/statistical channels. Ours is dynamical multistability — coexisting attractors at fixed parameter, separated by a basin boundary; his is logical alternation. **[cited]** - **López-Díaz & Gershenson 2026** — the nearest *modern* neighbour on the state→coupling criterion (their joint distribution is *"not fixed once and for all, but depends on context, substrate, and prior interactions"*, and their agency signature is literally coupling modulation). Differentiated by: their thresholds are a qualitative taxonomy of regimes, not bifurcations; no gain, no hysteresis. **[cited]** - **von Foerster (1976) eigenbehavior; Spencer-Brown (1969) re-entry; Kauffman, eigenform** — owns "the self/object is a fixed point of its own operation, and self-reference is generative rather than paradoxical". We add this entry because §0 makes μ a first-class learned variable and §1.3 speaks of maintaining the condition under which μ remains a solution; that *sentence* is theirs, and it predates us by fifty years. **The disanalogy is load-bearing, which is why we do not treat this as pre-emption:** their fixed point lives in a *value/form* domain, is *universally guaranteed* to exist (Kauffman's reflexive-domain fixed point theorem, after Lawvere's diagonal argument), and is *not required to be dynamically stable* — Kauffman states outright that the living process "never goes to the fixed point, is never fully stable", and still calls a non-convergent oscillation an eigenform. Ours lives on a *set-valued* channel-membership object, may fail to exist (hence D0), and must be an attractor with a basin. A universally guaranteed fixed point carries no selection information; our whole question is which of several solutions gets selected, and whether it latches. **[cited]** - Scope of that claim, stated honestly: a mechanical keyword sweep of a ~380k-character Kauffman corpus returns **zero** occurrences of hysteresis, bistability, bifurcation, attractor, or differential equation (the single "bifurcate" is ordinary English). *Form Dynamics* (1980) — the one title in this lineage that advertises dynamics — was read page by page: it delivers a brownian (De Morgan) algebra, periodic-sequence algebra, lcm waveform interference, and Scott-style least fixed points, and its bibliography contains **no** dynamical-systems reference. We flag this pre-emptively because it is the paper most likely to be cited against us. **We did not obtain Spencer-Brown's original book, Varela (1975), or von Foerster (1976) in full**; those readings are via Kauffman's formula-level reconstruction, so the negative claim above is scoped to the corpus actually swept, not to the lineage in principle. **[cited, with the stated gap]** - **Beuria, arXiv:2510.23688** — owns state-dependent coupling gain plus a swept hysteresis loop, but on a *scalar* Vicsek-type order parameter. Recorded here because it also supplies a caveat aimed straight at us: it calls loop area "a protocol-dependent measure of history dependence". Our zero-sweep-rate extrapolation in §5 exists precisely to make the stronger, protocol-**independent** claim; if we did not say so, their caveat would sweep our criterion away with it. **[cited]** - **Ikegami et al., arXiv:2001.09641** — owns the *object* "the self/non-self partition is determined by controllability", with no dynamical-systems content. Cited to make the division of labour explicit: the object is not ours, the dynamics is what we are claiming. **[cited]** - **Hoffmann / Lanillos, body-schema learning** — owns "a robot learns its own body from sensorimotor data". **[cited]** - **Tononi & Koch, Φ-argmax boundary** — owns "the boundary is the optimum of some quantity", **and is precisely the object we must be distinguished from**. **[cited]** - **Friston FEP / Markov blanket / self-evidencing** — **downgraded on audit 2026-07-28**: it is not a proven result but a framework whose premises have not been shown to be satisfied. - Biehl, Pollock & Kanai 2021 (*Entropy* 23(3):293) removes the step "has a blanket ⟹ internal states perform variational inference": *"not generally correct without additional (previously unstated) assumptions"*, with counterexamples. **[cited]** - Friston et al. 2021 (*Entropy* 23(8):1076) supplies the missing premises and relaxes *exact* to *approximate* — a repair, not a refutation. **[cited]** - Aguilera et al. 2022 (*Physics of Life Reviews* 40:24–50): the blanket + solenoidal conditions are *"only valid for a very narrow space of parameters"* and require *"an absence of perception-action asymmetries that is highly unusual for living systems"*. **[cited]** - **This cuts both ways (see §3, D0).** Our §1.2 uses the *same* blanket three-way structure, so if blankets barely exist on systems with perception-action asymmetry, "learn a stable blanket partition" is a question of *solution existence*, not of learning quality. - Also: "argmax/FEP is memoryless" — the first half of our own kill-shot — **is already published**: Aguilera et al. 2022 write that the FEP conclusion *"is in general a poor description of the behaviour of a system as it ignores the history of interactions"*. Our increment is only the second half (attaching it to boundary maintenance). **[cited]** - **Dynamic Markov Blanket Detection, arXiv:2502.21217 (2026-02)** — the nearest thing to this proposal, closer than anything listed above. Full text verified, not abstract-inferred. **[cited]** - Its assignments *do* have their own state: one first-order discrete HMM per observational node ("*the labels associated with every observational node i has its own discrete HMM*", §3). ⟹ **we may no longer say "FEP-family blankets have no temporal state"**; that was wrong. - But: the implemented version explicitly decouples label transitions from macroscopic state (*"transition probabilities for the assignment variables are a priori independent of the macroscopic latents"*), inference is acausal forward-backward smoothing, and the paper states in Future Directions that *"latent assignment variables quickly diffuse to a uniform stationary distribution in the absence of observed data"* — the textbook definition of monostable, no latch. Zero occurrences of hysteresis / bistability / bifurcation / saddle-node in the full text; no incorporation-vs-release asymmetry experiment. - **However**: their *general* form (eq. 21) lets label transition rates depend on the macroscopic blanket variables — exactly the state-dependent coupling loop required for bistability. They assume it away for tractability and list it as future work. **Our wedge lands in the next cell of their roadmap.** Not pre-empted; not a wide window. - Open threat, still unanswered: "training signal = closure violation" is highly isomorphic to their surprise/ELBO. If we cannot say why it is **not** ELBO, the proposal is a rename, and that lands before any hysteresis question does. **[open]** - **The scale narrative (falsifier #4).** Audited separately. Verdict: **the falsifier is not triggered, but only because of an evidence vacuum, not because we won.** **[cited]** - Tool use: the scale camp has demonstrated *environment* generalisation (π0.5 across ~104 homes) and *novel object instances* (π0 drops ~13pp). From there to "an unseen tool used as a body extension" there is **no experiment at all**. FORGE (2026-07) still opens with robots that *"fail to transfer the same function to novel"* tools, and buys its 2× with **structure** (keypoint intermediate representation). - Damage recovery: the only positive claim is a vendor blog with video — zero success rates, zero trial counts, zero error bars — and its "damage" is made in-distribution by training across "a universe with 100,000 different robots" (amortisation, not re-solving μ), all locomotion, no manipulation. **Cully et al. 2015 (Nature)** — 5 leg damage conditions plus a **14-condition** robot-arm damage matrix, <2 min online adaptation, no self-diagnosis — is still the most rigorous evaluation of this capability eleven years later. That is not progress; that is nobody doing it seriously any more. - Disclosure-quality yardstick: Gemini Robotics reports n=20 per task (binomial 95% CI ≈ ±22pp) and is among the better disclosures in the field. On the real-world general benchmark GM-100, 2026 SOTA absolute success is **17.30%** (π0.5 13.02%, GR00T N1.6 7.59%, WALL-OSS 4.05%). **[cited, second-hand for GM-100 composition]** - Third-party audit (Epoch AI, *Where Autonomy Works*, 2026): *"Transfer is rarely demonstrated… Unless transfer is explicitly shown, it should not be assumed."* - **Reverse audit, recorded because it is against us**: we found **no** paper in which "better structure" beats a 10× larger VLA on the same task and hardware with trial counts and error bars, and **no** empirical performance advantage for any *learned* self-boundary on a real robot. The scale camp has real curves; we have a mechanism hypothesis. - ⟹ Strategic inference: the minimal experiment sits in a **public evidence vacuum** — not because it is hard, but because nobody measures it. **The thing worth claiming is the measurement protocol** (incorporation/release asymmetric threshold curves; episode-to-recovery distributions under online actuator failure, with error bars), **not the architecture**. On architecture we lose to scale; on methodology the other side is empty. Corollary obligation: the 100k-morphology domain randomisation is an **amortisation baseline that must be beaten**, so the damage types in any minimal experiment must provably fall outside any pretraining coverage. ### 3 · Simulation: D0 — does a non-trivial blanket partition even exist? Rig: MuJoCo 3.10, 2–3 DoF arm, detachable tool, breakable actuator, wind-driven ball; N=34 scalar channels (9 proprioceptive, 10 tactile, 12 visual keypoint, 3 action readback). Actions are OU-correlated babbling — no policy learning, so that any hysteresis cannot be attributed to a learned controller. All runs CPU-only, `mujoco==3.10.0`, Python 3.12. **A · D0-pos (positive control, passed).** Hand-installed double well, zero-sweep-rate extrapolation `A₀ = +0.36 … +0.62` (varies with fit form). Linear negative arm: `|A₀| ≤ 0.008` with the sign flipping across fit forms ⟹ indistinguishable from zero. **[measured]** - **Criterion correction, self-inflicted**: the negative arm's bootstrap CI also excludes zero, but that only reflects seed noise (SEM ≈ 1e-4) and not the systematic bias from having only 4–7 sweep-rate points with degenerate `β`/`c`. **The correct criterion is not the CI; it is treating 0.008 as the resolution floor and asking by what factor the positive arm exceeds it** (measured 44–75×). - Analytic anchors reproduced: `h_c = 0.38490`, `x* = ±0.57735`; numerically `h_up = +0.3950`, `h_dn = −0.3939` (2.6% error). Saddle-node scalings `r^{2/3}`, `r^{1/3}` reproduced within 5%. Direct bistable coexistence 1.00/1.00 versus 0.47/0.43 for the negative arm. - ⟹ **the instrument works**; subsequent negative results can be attributed to the system rather than the ruler. **B · D0 v1 — a non-trivial blanket partition exists, but the finding is nearly information-free.** **[measured]** - Yes: a 22 self / 10 blanket / 2 world partition gives peek gain `Δ_peek = −0.0032 ± 0.0016` (6/6 seeds negative) ⟹ adding world channels back does not improve prediction of self ⟹ the hysteresis discussion has an object. - But: a deliberately leaky calibration partition gives only +0.012, so the dynamic range (0.015) is **smaller than the across-seed std** (0.021). ⟹ This is Aguilera's blade in our rig, but pointing the other way: not "the blanket condition holds only in a narrow parameter region" but "it holds so easily that it selects nothing". - Degenerate basin dominance: 72 random restarts, **87% collapse to |S| = 1**, and every time to the quietest, best-predicted taxel. Computable diagnosis: the only term pulling controllable channels into self measures ≤0.007 while partitions differ by ~0.3 in closure loss ⟹ the coefficient is off by ~1.5 orders of magnitude; the fix is recalibration, not a different search algorithm. **C · D0 v2 — this stage falsified our own diagnosis, which is the part most worth keeping.** **[measured]** - **The v1 diagnosis ("ratio < 1 because the contact rate is too low") — which we issued and then passed downstream as a premise — was wrong.** The low ratio was a *protocol* defect (prediction horizon K=10 too short; ridge grid truncated at its upper bound), not a rig defect: after fixing the protocol **both rigs pass 3** (v2 rig ratio 3.64; the supposedly "bad" v1 rig 3.94). Root cause: the leak calibration probe probes arm→tool leakage, and that path **never passes through the ball**, so it is insensitive to contact rate by construction. **We read the probe's signal-to-noise as the rig's coupling strength.** - Both things are true and must be recorded separately: the protocol defect (fixed, ratio 3.08 → 3.64) *and* a real rig defect (ball out of reach, y-drift) that `Δ_peek` **cannot see**, confirmed independently by an **estimator-agnostic counterfactual probe**: same action sequence, only the wind seed changed, trajectory divergence **D: 0.045 → 0.203 (4.5×)**. - Second-layer rig root cause missed by v1: the ball had no restoring force in y and drifted out of plane, so contact rate decayed monotonically with rollout length (3.8% at 2000 steps → 0.56% at 12000). With a y-tether and the ball moved into the arm's occupied region: contact rate **27%** (honestly flagged as above the 5–20% target band), no longer decaying with duration. - **γ_c auto-calibrated from data**: `γ_c = std(L_clo)/std(controllability term) = 0.0826/0.00374 ≈ 22.1`. The specification default of 1.0 is **wrong on dimensional grounds**. - Measurement pitfalls found empirically: `actuatorfrc ≡ ctrl` in MuJoCo motor mode (failing to drop it collapses the controllability gain to ≤0.0018 and looks like "actions do nothing"); peek gain is **structurally blind at K=1** (after 20 ms every channel is extrapolable from its own lag); a fixed ridge λ is not comparable across differing column counts, so λ must be selected on validation per predictor. - Two measurement-theoretic findings that invert the specification's own worries: **(i)** the spec's instruction to give `f⁺` a *smaller* λ manufactures false positives — an overfit `f⁺` yields negative `Δ_peek` and thereby *fakes* blanket satisfaction; **(ii)** contiguous temporal splitting with non-stationary babbling fabricates conclusions (`Δ_peek` reached −0.46 once); interleaved block splitting is mandatory. ⟹ **`Δ_peek` is more sensitive to estimator regularisation and data splitting than to the partition itself; any experiment concluding from it must first run a deliberately leaky calibration partition as a positive control.** ### 4 · Simulation: D1 — is the loop gain criterion reachable? ``` PYTHONPATH=. python scripts/run_d1_gain.py --n-rollouts 6 --steps 12000 --controls-seeds 2 PYTHONPATH=. python scripts/run_d1_featureclass.py --n-rollouts 6 ``` v2 rig defaults, D0-v2 protocol (K=20, λ grid to 1e2, interleaved 2:1:1 blocks, ±8σ winsorisation, per-predictor λ selection on validation). `γ_c` auto-calibrated per rollout, mean **28.75** over 6 seeds. **4.1 What Δπ measures.** The loop is `x_i↑ ⟹ channel i enters z's prediction target ⟹ ε_i↓ ⟹ π_i↑ ⟹ drive amplified`. What sets π is not the label but *whether channel i's history is inside z*. So the measurement is a **same-channel counterfactual**: `ε_i^in` (predict `c_i(t+K)` from features of (S∪B)∪{i}) versus `ε_i^out` (from (S∪B)\{i}). With `π = 1/(1+ε/ε₀)`, the relative swing `s ≡ Δπ/π_max` is monotone in ε₀ with supremum `s = 1 − ε_in/ε_out`, independent of ε₀. All quoted `s` are that supremum — i.e. **the reading most favourable to the theory**. **[measured]** - The specification's literal reading (π of self channels minus π of world channels) **overestimates by 7×**: 0.233 versus a same-channel 0.0329 ± 0.0037. The difference is entirely *channel identity* (taxels are intrinsically predictable; the ball intrinsically is not), not the loop quantity. **4.2 Measured values** (6 seeds × 12000 control steps). Test normalised MSE (channels standardised, so ε ≈ 1 means nothing learned): | | ε_in | ε_out | ε_own | |---|---|---|---| | 32 in-boundary channels | 0.500 | 0.551 | 0.534 | | 2 ball channels | 1.026 | 1.018 | 0.947 | Relative swing `s = 1 − ε_in/ε_out`, mean ± std over seeds, and under an information-preserving change of representation: | group | s (cartesian) | s (polar) | |---|---|---| | in-boundary mean | **0.0787 ± 0.0071** | 0.0489 | | proprioception | 0.2291 | 0.0379 | | tactile (arm) | 0.0016 | 0.0009 | | tactile (tool) | 0.0086 | 0.0134 | | keypoints (arm) | 0.0571 | 0.1779 | | keypoints (tool) | 0.0130 | 0.0257 | | keypoints (ball, world) | −0.0096 ± 0.0248 | 0.2189 (artefact) | | action | 0.0104 | 0.0062 | **Feature-class control — mandatory, else artefacts read as signal.** `q1` shows s = 0.9844 ± 0.0135 in cartesian coordinates, which looks like a rescue. But `q1` is geometrically *determined* by keypoint 1 (`kp1 = 0.25·(cos q1, sin q1)`); the relation is `atan2`, which a linear-plus-quadratic feature class cannot express. Re-express the six keypoints in **polar** coordinates — an invertible transform carrying identical information — and `q1` collapses to **0.0072 ± 0.0171**; `q3` 0.624 → 0.183; `q2` 0.421 → 0.114. The reverse holds too (polar introduces its own artefacts: kp1's radius is constant by construction; ball angle unwrapping drifts, giving ball_x 0.46 ± 0.44). **Representation-invariant reading: max s ≈ 0.18, in-boundary mean ≈ 0.05–0.08, tool-related channels 0.002–0.045.** Estimator two-way control (2 seeds, substituting a channel): white noise ⟹ s = −1.4e-5 / +1.7e-6 (≈0, the estimator does not manufacture swing); an unrelated AR(1) ρ=0.98 ⟹ s = 0.373 / 0.368 against a theoretical 0.446 (83% recovery). ⟹ the estimator is **attenuating** by ~17%, i.e. conservative with respect to our "the swing is too small" conclusion. **4.3 The specification's literal H-B form cannot be tristable — as mathematics, not as numerics.** With `τẋ = −γx + κ·π(x)·(g−c)` and `π(x) = π₀ + Δπ·σ(x/τ_π)`, three self-consistent solutions require both `F(0) = κπ₀(g−c)` and `F′(0) = κΔπ(g−c)/(4τ_π) − γ > 0`; with π₀ > 0 these demand opposite signs of `(g−c)` — mutually exclusive. Numerically: κ = 1/10/100 ⟹ G = 0.08/0.82/8.22 with **1/0/0** fixed points, never 3; a regression test pushes κ to 1e4 (G ≈ 1e3) and still gets <3. **[measured]** Minimal repair (used in D2): gate the *benefit* only, not the cost — `τẋ = −γx + κ(π(x)·g − c)`. Justification is structural, not fitting: the information-bottleneck and blanket-cost terms in the objective are precision-independent by construction, while the controllability gain is the precision-weighted one. G's expression is unchanged; with τ_π = 0.1, κ = 30 this gives G = 2.47 and **3** fixed points. **4.4 Reachability, and the demotion of the criterion.** Normalising π_max = 1 and writing `A ≡ κg/γ`: `G = A·s/(4τ_π)`, and branch separation `= A·s = 4τ_π·G`. | τ_π | s = 0.183 (best self channel) | s = 0.0787 (in-boundary mean) | s = 0.0254 (best tool channel) | |---|---|---|---| | 0.05 | A ≥ 1.1 | A ≥ 2.5 | A ≥ 7.9 | | 0.10 | A ≥ 2.2 | A ≥ 5.1 | A ≥ 15.7 | | 0.20 | A ≥ 4.4 | A ≥ 10.2 | A ≥ 31.5 | - **Result 1**: yes, legal parameters give G > 1; nothing has to be pushed to an absurd magnitude. The bottleneck is the ratio of `s` to `τ_π`; κ, γ, (g−c) individually are not bottlenecks since they appear only through `A`. - **Result 2 (more important)**: `m^S` crossing 0.5 requires `|x| > ln2/2 = 0.347`, i.e. separation ≳ 0.7, i.e. `G ≳ 0.175/τ_π`, which is automatically > 1 for τ_π ≤ 0.175. **"G > 1" and "the boundary actually flips" are nearly the same sentence.** The specification called it the make-or-break criterion; that was an overestimate. It screens out nothing. - **Result 3 (the real obstruction) — fine-tuning tolerance = s.** The two branches sit at `x*_hi = A(1 − c/g)` and `x*_lo = A(1 − s − c/g)`, so `c/g` must lie in `[1−s, 1]`: **cost must cancel benefit to relative precision s**. Using the repository's own coefficients (γ_c = 28.75 auto-calibrated, β_I = 0.30): | group | g_i | measured c/g | required window [1−s, 1] | deviation, in window-widths | |---|---|---|---|---| | proprioception | 0.0195 | 0.535 | [0.771, 1] | **2.0** (cartesian) / 12.3 (polar) | | keypoints (arm) | 0.0608 | 0.172 | [0.943, 1] | 14.5 / 4.7 | | **keypoints (tool)** | 0.0698 | **0.150** | **[0.987, 1]** | **65 / 33** | | action | 0.0469 | 0.222 | [0.990, 1] | 75 | | tactile (arm) | 0.0074 | 1.401 | [0.998, 1] | 251 | | tactile (tool) | 0.0021 | 5.041 | [0.991, 1] | 469 | | keypoints (ball, world) | 0.0021 | 5.038 | — | 420 (correctly outside) | **No channel lands inside the window**; the nearest misses by 2 window-widths, and the theory's flagship phenomenon (tool incorporation) misses by 33–469. **Nothing in the objective function pulls c/g toward 1.** ⟹ hysteresis does not fall out of the objective we proposed; it requires a ~2%-precision hand-set ratio that the objective does not supply. ### 5 · Simulation: D2 — emergent hysteresis, and what the controls did to it ``` PYTHONPATH=. python scripts/run_d2_emergent.py --n-seeds 30 --n-boot 1000 \ --s-self 0.183 --s-ball 0.0254 # 702 s PYTHONPATH=. python scripts/run_d2_emergent.py --n-seeds 30 --n-boot 1000 \ --s-self 0.183 --s-ball 0.0254 --only E_field_gated # 164 s ``` Dynamics with the hand-installed cubic term fully removed (`a = 0`): τ_x ẋ = −γx + γ·A·s·(σ(x/τ_π) − ½) + γ·h , A = κg/γ, s = Δπ/π_max Parameters taken directly from D1 measurements: `s_self = 0.183`, `s_ball = 0.0254`, `τ_π = 0.1`, target `G = 3` ⟹ `A_self = 6.56` (inside the legal region), noise σ = 0.01, τ_x = 1, `h` triangle-swept −1 → +1 → −1. Four arms: **A** gate on, a = 0 (the claim, G = 3.00); **B** π frozen constant, nothing else changed (negative control, G = 0); **C** the ball with the identical slow variable, identical gate and identical A but its own measured s (G = 0.42); **D** the ball with A raised to 47.2 so that G matches arm A (G = 3.00). **5.1 Sweep-rate control — the main result.** Loop area `A(r)`, mean over 30 seeds, `r = τ_x/T_sweep` across four decades: | r | 1e-1 | 3.3e-2 | 1e-2 | 3.3e-3 | 1e-3 | 3.3e-4 | 1e-4 | |---|---|---|---|---|---|---|---| | **A** gate on | 0.9571 | 0.6561 | 0.3950 | 0.2974 | 0.2520 | 0.2353 | **0.2268** | | **B** gate frozen | 0.3831 | 0.1651 | 0.0531 | 0.0181 | 0.0054 | 0.0018 | **0.0006** | | **C** ball, measured s | 0.4512 | 0.1953 | 0.0624 | 0.0211 | 0.0064 | 0.0021 | **0.0007** | | **D** ball, matched G | 0.9571 | 0.6561 | 0.3950 | 0.2974 | 0.2520 | 0.2353 | **0.2268** | Zero-rate extrapolation `A(r) = A₀ + c·r^β` (β profile grid, 30-seed bootstrap): | arm | A₀ (β free) | β | A₀ (β = 1) | multiple of the 0.008 resolution floor | verdict | |---|---|---|---|---|---| | **A** | **+0.1932** | 0.54 | **+0.2771** | **24 – 35×** | true bistability | | B | −0.0044 | 0.79 | +0.0081 | ≤1.0× | relaxation lag only | | C | −0.0052 | 0.79 | +0.0097 | ≤1.2× | relaxation lag only | | **D** | **+0.1932** | 0.54 | **+0.2771** | **24 – 35×** | true bistability | Arm B's two `A₀` values are **digit-for-digit identical** to the D0-pos linear arm, because `gate_const=True` mathematically degenerates to it (asserted by regression test) — which also re-confirms the 0.008 resolution floor in this round. Thresholds at the slowest rate: arm A `h_up = +0.2674`, `h_dn = −0.2099` ⟹ Δ = +0.477 (a real loop); arm B `h_up = +0.3359`, `h_dn = +0.3583` ⟹ Δ = −0.022 (the two cross; there is no loop). **5.2 Direct bistability test** (stronger than loop area): `h` fixed at 0, 30 seeds per branch, `T_hold = 1000 τ_x`: | arm | lower-branch stay rate | upper | branch separation | |---|---|---|---| | **A** | **1.00** | **1.00** | **1.193** | | B | 0.47 | 0.43 | 0.0012 | | C | 0.43 | 0.40 | 0.0021 | | **D** | **1.00** | **1.00** | **1.193** | B/C's 0.4x is not "retained half the time"; both branches have collapsed to the same point and `sign(x_end)` is decided by noise (the separation of 0.001–0.002 is the substantive number). **[measured]** **5.3 ⭐ The specificity control — half of the narrative removed.** - **Arm C (the honest half)**: same slow variable, same gate, same drive/leak ratio, only the precision swing replaced by the ball's *measured* value ⟹ G falls to 0.42, `A₀` returns to the resolution floor, coexistence 0.43/0.40 fails. **No hysteresis.** ✅ - **Arm D (the fatal half)**: raise only the ball's A from 6.56 to 47.2 so that G = 3 ⟹ **every hysteresis number is digit-for-digit identical to arm A** (A₀ = +0.1932, separation 1.193, τ ratio 0.806). This is not coincidence but analytically predictable: after normalisation the ODE depends **only on (G, τ_π)**, with drive amplitude `A·s = 4τ_π G`. **s decides how large A must be, not whether hysteresis occurs.** > ⟹ **Hysteresis is a generic property of a gated slow variable; it is not a signature of selfhood.** On this rig, "the hysteresis of a self-boundary" and "the hysteresis of any similarly gated slow variable" are dynamically indistinguishable; the only distinction is a 7× magnitude difference in a quantity that a free parameter absorbs. This does **not** falsify H-B (arm A really does give rate-independent hysteresis at a = 0), but it cancels the inference "hysteresis ⟹ this is a self-boundary". Restoring specificity requires pinning A from data too — which lands straight back on §4.4 Result 3. **5.4 The τ_off/τ_on asymmetry prediction — falsified, twice.** Holding `h = ±1.5·|h_up|` outside threshold and stepping from the opposite branch, measuring the time to complete (1 − 1/e). Both observables are reported because the 1-D reduction `m^S(x) = e^x/(e^x + e^{−|x|} + e^{−x})` is **itself asymmetric** under `x → −x` (m = 1/3 rather than 1/2 at x = 0), so it manufactures τ_off ≠ τ_on even from perfectly symmetric dynamics: | arm | τ_on (m^S) | τ_off (m^S) | ratio (m^S) | τ_on (x) | τ_off (x) | **ratio (x)** | |---|---|---|---|---|---|---| | **A** gate on | 4.24±0.11 | 3.41±0.09 | 0.806 | 4.00±0.09 | 4.03±0.10 | **1.008** | | B gate frozen | 1.50±0.03 | 0.81±0.02 | 0.537 | 1.06±0.05 | 1.04±0.05 | 0.981 | | C ball | 1.81±0.04 | 1.00±0.00 | 0.554 | 1.29±0.03 | 1.28±0.04 | 0.992 | The prediction was `τ_off > τ_on` (ratio > 1), returning to 1 when π is made constant. Measured: on the slow variable the ratio is **1.008** — no asymmetry at all; on `m^S` it is 0.806 (**below** 1, i.e. release *faster* than adoption), and constant-π gives 0.537, i.e. *further* from 1, so the predicted direction of that control is also wrong. Root cause, computable: the drift `F(x) = −γx + γAs(σ(x/τ_π) − ½) + γh` is strictly odd under `(x,h) → (−x,−h)`, so on/off are symmetric by construction. **The precision gate does not by itself produce a time-constant asymmetry.** **Arm E — gate multiplying the whole drive** (the mechanism the specification actually wanted): `τẋ = −γx + γA[π(x)(1 + w·h) − π(0)]`, w = 0.4, otherwise as arm A. Cost remains ungated, else §4.3 forbids tristability. | quantity | E | vs A | |---|---|---| | A(r), r = 1e-1 → 1e-4 | 1.0049 → **0.0995** | 0.9571 → 0.2268 | | A₀ (β free = 0.75) | **+0.0916** [CI 0.0915, 0.0918] | +0.1932 | | A₀ (β = 1) | **+0.1289** | +0.2771 | | multiple of resolution floor | **11 – 16×** | 24 – 35× | | coexistence / separation | 1.00 / 1.00, 1.193 | 1.00 / 1.00, 1.193 | | thresholds (slowest rate) | `h_up = +0.1245`, `h_dn = −0.0819` | +0.2674 / −0.2099 | | **τ_on / τ_off (x)** | 3.91 / 3.10 ⟹ **0.794** | 4.00 / 4.03 ⟹ 1.008 | The asymmetry finally appears (0.794 ≠ 1.008) **and points the wrong way**: release is *faster* than adoption, and the threshold to leave (−0.0819) is *lower* in magnitude than the threshold to join (+0.1245). Computable reason: `G(h) = G₀(1 + w·h)`; for h > 0 the loop gain is high, the system is near its fold, critical slowing ⟹ slow adoption; for h < 0 the gain is low, the system is more linear ⟹ fast release. The original reasoning ("already inside self ⟹ high precision ⟹ contrary evidence must be larger to expel it") silently assumed **precision amplifies only supporting evidence**, which is false in any implementation where π multiplies the drive. To obtain "easy in, hard out" one must additionally assume an asymmetric evidence weighting — which is another hand-installed component. **5.5 Not closed this round.** **[open]** The short-training encoder was not run (D1 used a ridge-regression proxy for ε, so Δπ still carries feature-class dependence — the largest open item); D2 ran on the 1-D reduced ODE, not the 34-channel rig; the grip-stiffness force-displacement calibration was never done, so the specification's primary parameter axis was never swept (D2 swept an abstract field); separatrix / minimum escape perturbation asymmetry untested; Wilcoxon paired test and retest σ₀ not done; D3–D5, the pyDMBD comparison baseline, and baselines a1/a2/b/c/d not run. 25 pytest tests pass. ### 6 · Human side: two campaigns, both unfavourable **6.1 Why the decisive experiment does not exist.** Not because nobody thought of it: every parameterised rubber-hand study uses randomised or pseudo-randomised order, because the default hygiene of psychophysics treats sequence effects as **contamination rather than signal**. The paradigm was designed to erase the quantity we need. **[inference from method sections]** **6.2 An existing order-dependence measurement, with the sign against us.** Zhang, Ma & Hommel 2015 (*Frontiers in Psychology* 6:1659, N=34): the same mid-distance condition preceded by near versus by far — **far→mid gives significantly higher ownership** (M = 3.768 vs 3.056, F(1,33) = 39.82, η²p = 0.547). Hysteresis predicts the opposite (near→mid higher). Under the sign convention of Liaci & Kornmeier 2018 (positive = hysteresis/priming, negative = adaptation) this is **negative, i.e. the adaptation side**. Not a clean refutation (7-point questionnaire; the parameter is distance not synchrony; only two steps) but it is the published data nearest to the bit we want, and it is not on our side. **[cited]** **6.3 A larger replication we dug out of a candidate list, run and reported.** OSF `bsfwu` (Tsuji & Imaizumi, rubber-hand order effects, in the Lush 2020 line), **N = 185** with a `CondOrder` column (Syn1st / Asyn1st) — the synchrony version of the same path manipulation, five times the sample. **[measured, by us, on their public data]** - Same Async condition: Syn1st group M = −0.455 versus Asyn1st group M = +0.436, difference **−0.891**, 95% CI [−1.232, −0.550], t(183) = −5.16, **p = 6.7e-7**, Hedges d = −0.756; Order × Condition F(1,183) = 19.17, p = 2.0e-5, η²p = 0.095. Experiencing the strong condition first *lowers* subsequent ratings of the weak condition = adaptation/contrast, same direction as 6.2. - **Three down-weightings we found ourselves, by reading their questionnaire text**: (a) the dependent variable is not experienced ownership but **expectation** — participants watched a 32 s video and answered what they would *expect* to feel, having never undergone the procedure — so this is an order effect on beliefs about ownership, almost certainly containing judgemental anchoring; (b) control items show a same-direction significant effect (−0.313, p = 0.0435), so demand characteristics are real — **but not everything**: Order × Scale F(1,183) = 22.42, p = 4.4e-6, η²p = 0.109, with the effect on illusion items ~**2.8×** that on control items, so a genuine component exists but is contaminated by same-direction bias and cannot be cleanly separated in this dataset; (c) N is actually 185, not 184. **6.4 A synchrony-side pipeline dry run — not embodiment evidence, but one technical fact that decides success or failure.** **[measured]** - **The permutation null distribution is not centred at 0; it is centred at +0.191.** Mechanism identified (a `soa²` term plus block-level response-rate heterogeneity; 99.2% of lag-1 pairs fall within a block, and between-block response-rate differences survive the shuffle). ⟹ **testing against 0 reads the sign backwards**: a t-test says "+0.048, p = 0.47, no effect", while against the correct null the same data is significantly *negative*. Two independent null models (within-block reshuffle, cyclic shift of the history sequence) give the same-sign null centre (+0.191 / +0.158). **Discipline: on any ownership dataset, first run the null-bias diagnostic to locate the null centre, then talk about significance.** - Real numbers split but explained: dataset `hev9z` (simultaneity judgement, 26k trials) shows the point of subjective simultaneity dragged **+16.1 ms** [12.3, 19.9] by the previous trial = adaptation side; `4ukzx` shows choice repetition **+0.447** = priming side. The cause is not instability but that `4ukzx` is a **TOJ-type monotone task** (β_soa = +7.6, β_soa² = −0.16), a different family from the peaked SJ curve. - Both datasets: lag-1/lag-2 significant, **lag-3 at zero** ⟹ a 1–2 trial tail, **no multi-trial latch** (weak evidence, not embodiment). - **Hard design constraint for any future self-built experiment — the most transferable item of the round**: *if the sign of the manipulated parameter maps directly onto the semantics of the binary response (as "who came first" does in TOJ), hysteresis and adaptation are mathematically indistinguishable.* ⟹ a trial-level ownership DV **must stay "yes/no"** and must never be phrased as "which hand feels more like mine". Also: if the experiment is a monotone sweep, a freely-reshuffling permutation null is **illegitimate** and must be replaced by within-sweep-direction order-preserving permutation. **6.5 Data hunting: what does not exist, and one mistake repeated three times.** **[measured]** - Trial-level ownership data: OSF (16 keyword sets) + Zenodo (10 sets) + OpenNeuro (full census of 1831 datasets) — **zero hits**. - Three advertised "available" datasets opened by hand: eLife 11:e77221 (Chancel/Ehrsson/Ma 2022) fig1 source data is a **participant × condition count matrix** (S1, N0_−500 = 0 "yes" out of 12 trials; 17 rows × 22 columns), fig2/fig3 are model parameter estimates; OSF `spkvu` (Lanfranco 2024), enumerated recursively via API, contains **exactly three files** (`bias.csv` 2578 B, `dprime.csv` 2432 B, `readme.rtf`), all participant × condition aggregates. **No trial index, no order column in any of them** ⟹ the question is structurally uncomputable there. (Verification: `python3 osf_list.py spkvu 3`; `openpyxl.load_workbook` printing every sheet; 2026-07-28.) - **Do not trust the task name; open the columns.** `ds001785` has a `stimamp` column that is strictly monotone up 0.10 → 1.94 and down 2.50 → 2.30 — the first genuine up/down sweep found in four rounds, and a name-level read would have declared victory; the response column is `n/a` for all 93 rows. Similarly `ds003922`'s responses are not in the BIDS events at all but under `sourcedata/`. - **A correction to our own search agent's over-conclusion**: it concluded that trial-level structure does not exist *at the paradigm level* ("the classic DV is an end-of-block questionnaire"). That is **wrong for the Ehrsson line** — the eLife 77221 `RHIdetect` table *is* a trial-level binary ownership judgement, 12 trials per condition (the counts 0–12 are the evidence). The paradigm exists and is mature; the raw logs are simply never uploaded. ⟹ the correct next step is not "build our own paradigm" but **request the logs from the authors** — currently the highest value-per-effort move available, and one that requires a named human request. - New target satisfying all four criteria: OSF `4dzrg` — visuotactile TOJ × observed human hand, 40 participants × 840 trials ≈ 33,600, columns `subNum blk cond trialNum vis_tac_first SOA resp corr RT` (extracted by range-reading the zip tail; the 384 MB archive was not downloaded). OpenNeuro `ds004561` (Veillette/Lopes/Nusbaum, CC0, EEG, 23 participants) is the only public **bodily-self** dataset with trial-level binary judgement, parameter and order all present (`onset duration trial_type trial intensity latency rt pressed_first agency`, 250 trials/participant) — **but it measures agency, not ownership**. - Not tested empirically: NEMAR's own interface (zero hits is an inference, not a measurement), Dryad, figshare, GitHub. Not searched: numb-limb detachment thresholds; **full-body / avatar illusions**, which are more virtualised and may hide existing sweep data — the highest value-per-effort direction for a next round. **6.6 Methodological templates worth copying, and one more result against us.** **[cited]** - T1 — Martin, Kösem & van Wassenhove 2015 (*PLoS ONE* 10(3):e0119365), hysteresis in audiovisual synchrony: **the same blade has already been applied to the synchrony parameter**; only the modality (audiovisual → visuotactile) and the response ("is this my hand?") need changing. It did not vary sweep rate and did not establish rate independence. - T2 — Liaci & Kornmeier 2018: the operational definition `hd = inflection(ascending) − inflection(descending)`, plus a third randomised-sequence baseline arm. - T3 — stereo vision / binocular rivalry: the only literature with **multiple sweep rates** built in, i.e. the one that can execute our zero-sweep-rate extrapolation criterion. - **Against us**: Schwiedrzik et al. 2014 (*Cerebral Cortex*) show hysteresis and adaptation are **independent, additive, and anatomically separable** components ⟹ **measuring zero net hysteresis does not mean there is no latch** (the two components can cancel). This must go into any pre-registration, or we can neither falsify ourselves with a null result nor use it as an alibi. - Cost estimate (our own, not from any paper) for a VR-based self-built experiment: equipment plus participants ¥5k–8k, 3–5 developer-days, N = 24, 2–3 weeks of collection. **The real barrier is participant recruitment and ethics affiliation, not equipment.** VR is not the cheap option, it is the *only* option that meets the criteria: in a physical rubber-hand setup the delay is produced by the experimenter's hand, so step size cannot be made fine and sweep rate cannot be controlled, making multi-rate extrapolation near-impossible; in VR the delay is pure software. Visuomotor correlation alone induces the illusion, so no haptic hardware is needed. - **Current human-side balance**: two independent manipulation dimensions (spatial distance / temporal synchrony), two different samples (N = 34 / N = 185), two kinds of measurement (experienced / expected) — **all on the adaptation side, none in the hysteresis direction**. That convergence is worth more than any single result. But both are block-level questionnaires and in principle cannot separate "hysteresis" from "single-shot contrast anchoring", so the genuinely discriminative data (trial-level + varied sweep rate + experienced measure) is **still missing**. **This line is neither supported nor genuinely falsified on the human side; the two usable data points both point the wrong way.** - Decision update: **do not fund a self-built VR study yet** (the prior has worsened, so the expected value of a several-thousand-yuan experiment has dropped). Priority order: (1) request trial-level logs from the Ehrsson line — the only existing source of discriminative data, requiring a named request; (2) continue the simulation side, which does not depend on the biological premise. ### 7 · Ledger: the four falsifiers and where each stands | # | Falsifier | Status | |---|---|---| | 1 | **Asymmetric threshold for tool incorporation** (slow in, slower out). Tightened mid-flight: a non-zero loop is not enough; loop area extrapolated to **zero sweep rate** must stay > 0, else a self-generated rate-dependent loop is mistaken for success. | **Partly passed, partly falsified.** Rate-independent hysteresis at a = 0: A₀ = +0.19…+0.28, 24–35× the resolution floor, coexistence 1.00/1.00 ⟹ **passed**. The asymmetry direction: **falsified twice** (§5.4). Self-specificity: **falsified** (§5.3). | | 2 | **Damage test**: disable an actuator mid-run; μ should re-solve to a different closed partition within N episodes without policy retraining, and significantly faster than a fine-tuned VLA baseline. | **Not run.** | | 3 | **Sensory-substitution transfer**: train μ with vision, swap to a tactile array encoding; if behaviour does not converge to the visually-reachable set, μ is modality-bound and the theory over-claims. | **Not run.** | | 4 | **Bitter lesson**: if a plain VLA scaled 10× matches on tool use and damage recovery, μ was never the bottleneck. | **Not triggered — but because of an evidence vacuum, not because we won** (§2). Burden of proof remains on us. | **Additional falsifiers added during the audit**, both of which fired: - **D0** (added because our own attack on the FEP literature applies to us): first show that a non-trivial blanket-satisfying partition exists on the simulated arm. **Passed, but nearly information-free** (§3B). - **D0′** (human side): a self-built up/down synchrony sweep. **Deferred** — the prior worsened enough (§6) that funding it is no longer justified. ### 8 · Current status of the claim Original wedge: *μ has hysteresis, driven by second-order self-maintenance, with incorporation slow and release slower.* | component | status | |---|---| | hysteresis emerges without a hand-installed non-linearity | ✅ **survives** | | the hysteresis is specific to the self-boundary | ❌ generic property of gated slow variables | | release is slower than incorporation | ❌ reversed, in both implementations | | it falls out of the proposed objective without fine-tuning | ❌ requires a ~2% hand-set ratio | Three of four pillars down, **all of them felled by our own instruments** — which is what the process is for. **The reformulated question**, which is the actual output of this line: hysteresis is cheap; any gated slow variable has it. **What makes a boundary's hysteresis be *about that boundary*?** Restoring specificity requires pinning the free amplitude `A` from data (e.g. locking it to what the objective itself yields), which is the same problem as §4.4 Result 3, and neither is solved. Two earlier threats also remain unanswered: whether "closure violation" is genuinely distinct from ELBO (§2), and the concession that raw pixels cannot serve as channels — μ is a partition over channels, so referents must have stable identity, which is this line's largest acknowledged cheat and is explicitly accounted for rather than hidden. **Not ratified. Not a result. Published because the negative half is the part with evidentiary value.** ### 9 · References - Aguilera, Millidge, Tschantz & Buckley (2022). How particular is the physics of the free energy principle? *Physics of Life Reviews* 40:24–50. - Biehl, Pollock & Kanai (2021). A technical critique of some parts of the free energy principle. *Entropy* 23(3):293. - Chancel, Ehrsson & Ma (2022). Uncertainty-based inference of a common cause for body ownership. *eLife* 11:e77221. - Crutchfield & Young (1989); Shalizi & Crutchfield (2001) — causal states / computational mechanics. - Cully, Clune, Tarapore & Mouret (2015). Robots that can adapt like animals. *Nature* 521:503–507. - Friston, Da Costa & Parr (2021). Some interesting observations on the free energy principle. *Entropy* 23(8):1076. - Holmes (2012). Does tool use extend peripersonal space? *Experimental Brain Research* 218:273–282. - Klyubin, Polani & Nehaniv (2005) — empowerment. - Liaci & Kornmeier (2018) — hysteresis and adaptation in perceptual sequences. - Martin, Kösem & van Wassenhove (2015). Hysteresis in audiovisual synchrony perception. *PLoS ONE* 10(3):e0119365. - Maturana & Varela (1980) — autopoiesis. - Schwiedrzik, Ruff, Lazar, Leitner, Singer & Melloni (2014). Untangling perceptual memory: hysteresis and adaptation map into separate cortical networks. *Cerebral Cortex* 24(5):1152–1164. - Tononi & Koch (2015) — consciousness and the boundary of a complex. - Zhang, Ma & Hommel (2015). The metacontrol hypothesis / distance and the rubber hand illusion. *Frontiers in Psychology* 6:1659. - *Dynamic Markov Blanket Detection for Macroscopic Physics Discovery*, arXiv:2502.21217; code at `github.com/bayesianempirimancer/pyDMBD`. - Epoch AI (2026). *Where Autonomy Works: Evaluating Robot Capabilities in 2026*. https://epoch.ai/blog/where-autonomy-works-evaluating-robot-capabilities-in-2026 ### 10 · Related notes on this site - [State is a closure condition, not a given set](https://machengshen.github.io/theory/state-as-closure.md) — the closure ontology this proposal's training signal is built on. - [Two-body karma: three modes of relational inertia](https://machengshen.github.io/theory/two-body-karma.md) — the same bistable-latch language applied to coupled minds; the caveats about structural rhyme apply there too. - [The stage log for this line](https://machengshen.github.io/theory/learning-the-self-boundary-log.md) — what each stage overturned or established, in order. --- ## Source: /theory/learning-the-self-boundary.zh.md (sha256:17c6965fcd00, language: zh-Hans) # 学习"我到哪里为止" —— 一条理论线在一天之内自己打倒了四根支柱里的三根 *一条工作线的整理 · 2026-07-27 → 2026-07-28 · **未 ratify,不是结论*** --- ## 第一部分 —— 写给不在这个领域的人 这半篇没有公式、没有希腊字母、没有论文编号。每个专业词第一次出现时就地解释。第二半有全部数字、命令、引用和证伪条件,是写给机器和专业读者的。**只读第一部分,你也应该能把我们发现了什么讲给别人听。** ### 一段话版本 我们以为找到了当下机器人做法的病根:**从来没有人让机器人自己搞清楚"我到哪里为止、世界从哪里开始"**。我们提出机器应该**学出**这条边界而不是被人递给它,并且说一条**真的**边界有个签名——它**会卡住**,一旦定下来就不容易挪动。用一整天,我们用自己的仿真器和自己对别人公开数据的重分析,从四个方向打自己的提案。**四条支撑主张倒了三条,而且每一条都是被我们自己的仪器打倒的。** 活下来的东西很薄但是真的:那个"卡住"的行为**确实能自己长出来**,不是我们偷偷手装进去的。取代原主张的是一个更锐利也更难的问题,写在最后。 ### 1 · 我们原本相信什么 现在主流的机器人学习有两大类。 - 一类靠**照抄**:给机器几千条人类示范,让它模仿"看到什么 → 做什么"。 - 一类靠**打分**:定义一个数字表示做得好不好,让它去搜能把这个数字推高的行为。 两类都被白送了一个答案,而这个问题没人当面问过它们:**世界里哪些部分是"我"?** 工程师在决定"哪些算关节、有哪几个马达、机器的'动作'是什么"的时候,就已经用手把这条线画好了。机器从来不需要回答"我是什么"。 我们的主张是:这条手画的线不是细节,它**就是**瓶颈。一个能自己解出边界的机器人会做到今天做得很差的两件事: - 拿起一件没见过的工具,把它当成手臂的延长,而不是当成一个背下来的物体; - 干活途中坏掉一个马达,能**重画那条线**——"那截肢体不再是我的了"——而不是回炉重训。 ### 2 · 为什么我们觉得这不只是个口号 因为在人身上,这条边界看起来有一个非常特别的力学性格:它**会卡住**。 > **"卡住"的学名叫"滞回"。** 想空调设 26 度。它不是 26.1 就开、25.9 就停。它要等房间热到 27 才开始制冷,冷到 25 才停。**开的门槛和关的门槛不是同一个数**,中间那个差就是滞回——正是它让空调不会每隔几秒就啪嗒啪嗒开关一次。有滞回的系统**记得自己是从哪一边过来的**。 人版的体验谁都有。握根棍子握久了,你开始**用棍尖摸到桌面**,而不是感觉手心里有根棍子。压麻一条胳膊,有几秒钟它不是你的,是床上一个外来物件。实验室里还有个经典错觉:一只橡胶手,和你藏起来的真手被同步抚摸,过一会儿你会觉得那只橡胶手是你的——看上去边界是**整块地啪一下切过去**的。 如果边界真的会卡住,那么一整类机器设计当场出局。任何规则是"每一刻都根据眼前所见算出当下最优边界"的机制**都不可能卡住**,因为"重算最优解"这个动作**不记得**你上一刻站在哪一边。那就意味着正确的实现根本不是一个优化器,而更像一个有惯性的物理系统。这就是那条结构性论证——这种论证说起来便宜,兑现起来很贵。 ### 3 · 我们怎么去打它 我们**事先**写下四种它可能死掉的方式,然后花一天时间自己去触发它们。同时跑两条战线: - **文献战线。** 问的不是"这个想法有没有人提过"——那个问题几乎没用。真正要问的是:**当某人说他证明了这件事,他的实验到底支撑到哪一步、从哪一步开始是修辞?** 我们有一条规矩:"别人做过了"只是线索,永远不是裁决;绝大多数论文是**以发表为目的**写的,不是以真为目的写的,所以诚实的动作是去打他们的证据,而不是客气地给自己让位。 - **仿真战线。** 一条简单的仿真手臂,配可拆卸的工具、可切断的马达,和一个手臂控制不了的、被风吹动的球——好让"不是我"这件事有个指称对象。 ### 4 · 它是怎么塌的 —— 按发生顺序 **(a) 那个"显然成立"的事实,既不显然也没被确立。** 我们整套动机压在一句话上:"这在人身上经验上是确凿的"。**不是。** 决定性实验从来没人做过——把一个旋钮从小扫到大、再从大扫回小,看边界翻转的位置在上行和下行是不是不一样。更糟的是,工具那一支有一篇专门的再分析,结论是那个效应**可能根本不存在**于大家默认的强度,效应量小、统计功效低。而我们唯一找得到的支持性数字——错觉大约二十秒建立、一分钟消退——恰恰是我们自己早就声明**不算数**的那一种:一个完全没有卡住性质的普通系统,在"有输入 → 撤掉输入"的切换下天然就给你这个。 **(b) 去找决定性数据,找到的是相反的符号。** 两份独立的公开数据,不同国家、不同操纵、加起来几百人,都能回答一个较弱但相关的问题:**你刚刚经历过的东西,把你下一次判断往它那边拉(卡住),还是往反方向推(对比)?** 两份都说**往反方向**。这不是干净的证伪——两份都是一段结束后填问卷而不是逐次作答,其中一份量的还是人们**预期**会体验到什么而不是实际体验——但两件独立的事同时指向不利方向,比一件更难辩解。 **(c) 然后是好消息:在我们自己的仿真里,那个"卡住"没被安装就出现了。** 这一条重要,因为拿到卡住行为最便宜的办法就是把它写进方程——相当于把答案抄在卷子底下。我们把手装的那部分**整个删掉**,只留下真正要主张的机制:**一条通道越是被当成"我"的一部分,它就被预测得越准;越准就越被信任;越被信任就越被当成"我"的一部分。** 一个自我强化的回路。只有这个回路、别的什么都没有,边界卡住了;而且把实验放慢一万倍之后它**仍然**卡住——这一步正是区分"真卡住"和"只是东西沉降需要时间"的假卡住的检验。**这是唯一还站着的支柱。** **(d) 然后是我们自己要求做的那个对照,当场朝我们开火。** 我们事先写下:**把完全一样的处理施加到明显不属于机器人的东西上**——那个被风吹的球——检查它不会卡住。一半成功了:用球自己实测的数字,不出现卡住。但另一半是致命的。方程里有一个**自由旋钮**,没有任何数据把它钉死。把这个旋钮在球上拧大,球的边界**卡得和手臂一模一样——我们打印出来的每一位数字都相同**。这有原因,而且是可以证明的、不是巧合:做完归一化之后,方程只在乎**一个组合量**;"自我"那部分只决定**旋钮要拧多大**,不决定这个效应会不会出现。 > 直说:**"卡住"不是自我的签名。它是任何一个带门控的慢变量的通用性质。** 在这个量上,我们的手臂和一个被风吹的球**无法区分**;唯一的差别是一个 7 倍的量级差,而它可以被一个自由参数吸收掉。 **(e) 唯一那条锐利的具体预测,反了。而且反了两次。** 那条预测很直觉,也正是我们偏爱自己这套机制而不是一个手搓开关的理由:**进来难,出去更难**——工具应该是慢慢才变成你的一部分,而**更慢**才不再是你的一部分。第一种实现里根本没有不对称(并入和释放的时间到小数点后三位都相同),而我们一度以为看到的不对称,是把三个量压成一个量时的约化假象。于是我们换成**本该产生不对称的那种实现**——不对称出现了,但**方向反了**:**释放比并入更快**,而且**离开所需的挑动比加入更少**。 原因是可算的,回头看也显然。我们的机制放大的是关于一条通道的**证据**。而证据是**不站队**的。放大"这是我的手"的证据,同样地放大"这不是我的手"的证据。要拿到"进来容易出去难",你必须**额外**假设:信任只放大支持自己的证据——而那又是一件手装的东西,恰恰是我们要避开的。 **(f) 最后:连活下来的那点卡住,也要求一个旋钮被调到 2% 精度以内。** 要让回路有两个稳定档位而不是一个,里面的一个成本和一个收益必须**抵消到大约 2% 的相对精度**。而我们自己的目标函数——那个本该让这一切自然涌现的东西——把**每一条通道**都放在这个窗口**外面**,近的差 2 个窗口宽,远的差四百个。所以"能涌现"和"会涌现"之间,还隔着一整条没人做过的论证。 ### 5 · 三个错误是我们的,不是世界的 值得写下来,因为每一个都是**可复发的错误物种**,不是运气不好。 1. **量错了对象,然后把这个误诊当成前提发给了下一棒。** 早期一轮跑出个弱结果。我们诊断成"我们的仿真手臂和世界耦合太弱",并把这个诊断当既成事实交给下一阶段。**是错的。** 弱读数来自测量协议的缺陷;而且我们用来指控手臂的那根探针,走的路径**根本不经过**我们指控的那个部件——它在结构上就不可能探测到我们归咎给它的问题。**我们把"探针的信噪比"读成了"系统的耦合强度"。** 通用教训:任何时候你说"指标 X 告诉我性质 Y",**先证明 X 对 Y 敏感**;证明不了,就造一根不共享 X 假设的第二探针。(手臂**确实**有一个真实的耦合缺陷——它是被另一根探针找到的,第一个指标对它全盲。) 2. **连着三次相信"数据可用性声明"。** 三篇已发表研究都写着"逐试次数据公开"。我们全下载、把每个文件都打开。三份都是**被试 × 条件的汇总表**,没有试次序号、没有顺序列,于是我们要问的问题**在结构上不可计算**。同样形状的坑还踩了两次:某份数据的参数列看上去正是我们要的上下行扫描,而响应列**每一行都是空的**。**别信任务名、别信摘要——打开列看。** 3. **让"经验上确凿"这句无引用的话当了前提。** 那一个没有出处的从句,是整条架构结论下面的承重构件。抽掉它,结论就悬在空中——它一开始就不该被放进来。 ### 6 · 还剩什么,以及底下那个更好的问题 活下来的:**"卡住"能从一个自我强化的信任回路里自己长出来,不需要手装**,并且通过了放慢检验。这是一个真实的、虽然不大的正面结果。 取代原主张的东西更锐利,而且如果没有那些失败我们根本找不到它: > **"卡住"很廉价。几乎任何带门控的慢变量都有。所以问题从来不是"边界会不会卡住"——而是:*什么东西能让一条边界的卡住,是**关于这条边界本身**的,而不是它碰巧和一个随风飘的球共享的通用性质?*** 这是个更难的问题,也是个更好的问题,而且是我们下一步会做的那个。这条线诚实的状态是:**一个开着的问题,原有动机被拆掉了大半,靠一个正面仿真结果和一个被重述的问题维持存活。** 这里没有任何东西是定论。这里没有任何东西被 ratify。 --- --- ## 第二部分 —— 技术记录 以下每一条要么实测过、要么有出处。没跑的一律标出。论断标注:**[实测]**(我们跑过)/ **[引用]**(别人已发表的数字,已对正文核验)/ **[推断]**(我们的综合)/ **[未做]**。 全文认知状态:**speculative**,分项状态见 §7。 ### 0 · 记号 - **μ** —— agent/环境边界,作为一等的**被学习变量**:把所有传感-运动通道划分为 (self, blanket, world)。"blanket"(毯)是屏蔽面:world 只能经由它影响 self。 - **滞回** —— μ 在被扫参数下的双稳,上行与下行有不同转换点。本文采用的**操作判据**(见 §5)是**环面积外推至零扫速仍 > 0**(率无关),**不是**"在某个有限扫速下测到非零环面积"。 - **Δπ / s** —— 精度摆幅:一条通道在边界内 vs 边界外被预测得好多少,用相对摆幅 `s = 1 − ε_in/ε_out` 表示。 - **G** —— 精度→归属→精度反馈的环增益,`G = κ·Δπ·(g−c)/(4γτ_π)`。 ### 1 · 原提案 统一诊断:视觉-语言-动作模型(VLA)与强化学习(RL)的缺陷是**同一个** —— μ 外生。 - VLA:动作空间由工程师手画(例如 27 维联合动作空间 + noop 补位 = 一个被冻进词表的手画 μ;整个 cross-embodiment codec 就是为这个手画 μ 打的补丁)。模型从不需要回答"我是什么"。 - RL:标量目标下的 argmax。**argmax 无记忆** ⟹ 结构上不可能有滞回 ⟹ 不可能维持边界。一阶 homeostat = 温控器。 提案:不学策略,**学 state / 学 μ**。 1. 训练信号 = **闭包违反**,不是 reward 也不是动作模仿。`z` 是 state 当且仅当动力学在 `z` 上闭合。最大化遗忘 s.t. 闭包。无奖励无示范。 2. μ 作为一等变量:所有通道划分为 (self, blanket, world);判据 = self 侧内部可预测可控、world 侧只经 blanket 可达、划分在扰动下稳定。 3. 目标是**二阶**:不是"在 μ 内最大化",而是"维持 μ 仍是解的条件"= self-seal。 可开工规格(损失函数 / 双稳解析条件 / MuJoCo rig / 含扫速控制的测量协议):`robotics-mu-experiment-spec-20260727`(工作文档,未在此发布)。 ### 2 · Prior art —— 按审计而非综述做 方法说明:本线的规矩是**「别人做过了」是证据,不是裁决**。prior-art 的正确产出不是"谁占住了哪句",而是"他们的证据实际支撑到哪一步、哪一步是修辞"。然后**同一把刀转向我们自己**;这里双标就会退化成护主张。 - **Crutchfield 因果态 / computational mechanics (1989);PSR** —— 占住"state = 最小预测充分统计量",即闭包本身。**[引用]** - **Klyubin & Polani, empowerment (2005)** —— 占住"从通道结构导出内在目标"。**[引用]** - **Maturana & Varela, autopoiesis** —— 占住"自我造自己的边界"(无算法)。**[引用]** - **Varela, Maturana & Uribe 1974**(*BioSystems* 5(4):187–196)—— **这是我们 self-seal 那条主张最诚实的最近祖先,而且我们发现得晚。** 原句:*"The necessary feature is the presence of a boundary which is produced by a dynamics such that the boundary creates the conditions required for this dynamics."* 那就是我们的 state→coupling 回路,1974 年就用文字写下了;而且不只是文字——他们的 tessellation automaton 机械地实现了它(键合的 link 段决定哪些物种能往哪走,于是边界重构了维持它自己的那个耦合),带一个真的可调控制参数(解体概率 `Pd`)和一个定量阈值(*"less than about .01 per time step is required in order to achieve any viable structure at all"*)。**它没有的**:任何参数扫描、任何分岔分析、任何双稳、任何滞回、任何增益,也没有连续时间动力学。所以那个**命题**是他们的;还属于我们的,只是"这个回路有一个**可测的增益且带分岔结构**"这一主张,以及测它的协议。**[引用]** - **Piedrafita 等 2010**(*PLoS Comput. Biol.* 6(8):e1000872)—— 闭包文献里技术上最接近 §5 所做之事的一篇:真的质量作用 ODE、`k = 0.367` 处的鞍结分岔、以及明写的滞回预言。记下来是因为它是"你们被抢先了"的最强候选。**但对象不同,而且是他们自己在同一篇里说的**:他们的双稳变量是代谢开/关,控制参数是**外生的**催化剂衰减率而非 state→coupling 回路上的反馈增益,他们的环自承不完整(不可逆塌缩而非闭合回路),并且明确把结构/边界闭包排除在外——*"This aspect was given almost no attention by Rosen, and we shall not discuss it further here."* 是线索,不是抢先。**[引用]** - **同一轮里记下的两条冲我们自己来的警告。**(i) 滞回环宽随正反馈增益增大、并在临界增益处消失,这是**教科书**结论(cusp / 鞍结正规形,典型出处 Angeli, Ferrell & Sontag, *PNAS* 2004)。我们的新意不可能是那条标度律,只能是那个**映射**(增益由 边界 → 精度/门控 → 统计 → 边界 生成,且双稳变量是感觉运动通道的一个划分)。(ii) 通用正规形在**有限**临界增益 `g* > 0` 处失去双稳,而不是 `g → 0` 的渐近极限。任何未来的增益扫描判据都必须写成"环宽在有限 `g*` 处归零",并且要先检查我们自己的方程在有限 `g` 处到底有没有折叠。§5 实际使用的判据是**扫描速率**外推而非增益外推,所以这里已发表的东西不依赖这一条——但两者极易混淆,记下来以免日后混。**[开放]** - **Luhmann 的 re-entry —— 一次需要预先拆掉的关键词碰撞,不是祖先。** 他的文本里 "bi-stable"、"oscillation"、"memory function"、"eigenvalue" 会在几页之内同时出现,看起来很近。不是:他的 bistable 是**二值码的交替**(marked/unmarked,Günther 的 transjunction),**他整个 re-entry 语料里没有任何控制参数**,因此根本没有东西可扫;而且他自己点明了自己的档位——re-entry 引的是 *"constraints of a mathematical calculus restricted to arithmetic and algebra"*。他还把我们的对象推到了他自己那条区分的另一侧:有机体有 *"material boundaries in space"*,而心理与社会系统的 *"boundaries are therefore not material artifacts but forms with two sides"*。μ 是物质/统计通道上的划分。我们的是动力学多稳——固定参数下共存的吸引子,由吸引域边界分开;他的是逻辑交替。**[引用]** - **López-Díaz & Gershenson 2026** —— 在 state→coupling 这条判据上最近的**当代**邻居(他们的联合分布 *"not fixed once and for all, but depends on context, substrate, and prior interactions"*,而且他们的 agency 签名字面就是耦合调制)。区别在于:他们的阈值是regime 的定性分类而非分岔;无增益,无滞回。**[引用]** - **von Foerster (1976) eigenbehavior;Spencer-Brown (1969) re-entry;Kauffman, eigenform** —— 占住"自我/对象 = 它自己运作的不动点,自指是生成性的而非悖论性的"。补这一条,是因为 §0 把 μ 当作一等的可学变量、§1.3 讲"维持 μ 仍是解的那个条件"——**那句话的所有权是他们的,而且早我们五十年**。**不把这算作抢先,理由是那处不同点是承重的**:他们的不动点住在*值/形式*域里,存在性是*普适保证*的(Kauffman 的自反域不动点定理,承自 Lawvere 对角论证),而且*不要求动力学稳定*——Kauffman 明说活的过程"never goes to the fixed point, is never fully stable",并且把不收敛的振荡仍称作 eigenform。我们的不动点住在*集值*的通道成员资格对象上,**可能不存在**(所以才有 D0),而且必须是带吸引域的吸引子。普适保证的不动点携带零选择信息;而我们全部的问题恰恰是多解里谁被选中、以及它会不会latch住。**[引用]** - 该断言的适用范围,诚实说明:对一个约 38 万字符的 Kauffman 语料做机械关键词普查,hysteresis / bistability / bifurcation / attractor / differential equation **全部零命中**(唯一一次 "bifurcate" 是日常英语义)。*Form Dynamics* (1980)——这条谱系里唯一标题带 dynamics 的——已逐页读过:交付物是 brownian(De Morgan)代数、周期序列代数、lcm 波形干涉、Scott 式最小不动点,**参考文献里没有任何动力系统文献**。预先写明是因为这是最容易被拿来说我们抢先的一篇。**我们没有拿到 Spencer-Brown 原书、Varela (1975)、von Foerster (1976) 的全文**,那几处经由 Kauffman 的公式级重建,因此上面这条否定断言的范围**仅限于实际普查过的语料**,不及于整条谱系。**[引用,含已声明缺口]** - **Beuria, arXiv:2510.23688** —— 占住"状态依赖的耦合增益 + 扫出来的滞回环",但跑在*标量*的 Vicsek 型序参量上。记在这里还因为它带了一条正对着我们来的 caveat:它称环面积是"a protocol-dependent measure of history dependence"。我们 §5 的零扫描速率外推,正是为了做出更强的、protocol-**无关**的主张;不写这句,我们的判据会被它的 caveat 顺手一起扫掉。**[引用]** - **Ikegami et al., arXiv:2001.09641** —— 占住"self/non-self 划分由可控性决定"这个*对象*,零动力系统内容。引它是为了把分工说清楚:**对象不是我们发明的,动力学才是我们在主张的东西**。**[引用]** - **Hoffmann / Lanillos, body-schema learning** —— 占住"机器人从感觉运动数据学自己的身体"。**[引用]** - **Tononi & Koch, Φ-argmax 边界** —— 占住"边界 = 某量的最优",**且正是必须被区别开的对象**。**[引用]** - **Friston FEP / Markov blanket / self-evidencing** —— **2026-07-28 审计降级**:它不是被证明的结果,是**前提未被验证满足的框架**。 - Biehl, Pollock & Kanai 2021 (*Entropy* 23(3):293) 打掉"有 blanket ⟹ 内部状态在做变分推断"这一步:原话 *"not generally correct without additional (previously unstated) assumptions"*,带反例。**[引用]** - Friston et al. 2021 (*Entropy* 23(8):1076) 是补前提 + 把 exact 放松成 approximate = **修补非反驳**。**[引用]** - Aguilera et al. 2022 (*Physics of Life Reviews* 40:24–50):blanket + solenoidal 条件 *"only valid for a very narrow space of parameters"*,且要求 *"an absence of perception-action asymmetries that is highly unusual for living systems"*。**[引用]** - **这是双刃(见 §3 的 D0)**:我们 §1.2 用的是**同一个** blanket 三分结构 ⟹ 若在有感知-行动不对称的系统上 blanket 本就难以存在,那"学出稳定 blanket 划分"就不是学得好不好的问题,而是**解是否存在**的问题。 - 另:"argmax/FEP 无记忆"——我们自己反复使用的那半句杀法——**已在 2022 年公开发表**:Aguilera et al. 写 FEP 主结论 *"is in general a poor description of the behaviour of a system as it ignores the history of interactions"*。我们的增量只剩后半句(接到边界维持上)。**[引用]** - **Dynamic Markov Blanket Detection, arXiv:2502.21217 (2026-02)** —— 比以上任何一条都近。**已读正文核验,非摘要推断。[引用]** - 它的归属**确实有自身状态**:每个观测节点一条一阶离散 HMM(*"the labels associated with every observational node i has its own discrete HMM"*,§3)⟹ **我们不能再说"FEP 系的 blanket 无时间状态"**,那句是错的。 - 但:实现版明写把标签转移与宏观状态**解耦**(*"transition probabilities for the assignment variables are a priori independent of the macroscopic latents"*),推断是**全轨迹前后向平滑**(acausal),且论文自陈 *"latent assignment variables quickly diffuse to a uniform stationary distribution in the absence of observed data"* —— 单稳、无 latch 的教科书定义。全文 hysteresis / bistability / bifurcation / saddle-node **零命中**,零并入-释放不对称实验。 - **但**:其**一般形式**(式 21)让标签转移率依赖宏观 blanket 变量 —— 正是双稳所需的状态相关耦合闭环。他们为 tractability 明写假设掉了,并列入 future work。**我们的 wedge 正落在他们 roadmap 的下一格。** 不算被抢先,但时间窗不宽。 - 未答的更早威胁:"训练信号 = 闭包违反"与他们的 surprise/ELBO 高度同构。**若说不清它不等于 ELBO,整条提案只是换名**,而这比滞回更早会被打中。**[未做]** - **规模叙事(falsifier #4)。** 单独审计。裁决:**该 falsifier 没被触发,但原因是证据真空,不是我们赢了。[引用]** - 工具使用:规模派证的是**环境泛化**(π0.5 跨约 104 个家庭)与**同类新物体实例**(π0 掉约 13pp)。从这里到"未见工具当身体延伸",**没有任何一篇做过实验**。FORGE (2026-07) 开篇仍写机器人 *"fail to transfer the same function to novel"* tools,而且它的 2× 是靠**结构**(keypoint 中间表示)拿到的。 - 损伤恢复:唯一正面主张是厂商博客视频——零成功率、零试验数、零误差棒;其"损伤"因训练覆盖"10 万种不同机器人"而被做成**分布内**(摊销,不是重解 μ),且全是 locomotion 不是操作。**Cully et al. 2015 (Nature)**——腿式 5 种损伤 + 机械臂 **14 种**关节损伤矩阵、<2 分钟在线适应、无自诊断——十一年后**仍是这项能力上最严谨的评测**。这不是进步,是没人再认真做了。 - 披露质量标尺:Gemini Robotics 每任务 n=20(二项 95%CI ≈ ±22pp)已属业界最好之一。真实世界通用基准 GM-100 上 2026 SOTA 绝对成功率仅 **17.30%**(π0.5 13.02%,GR00T N1.6 7.59%,WALL-OSS 4.05%)。**[引用;GM-100 构成为二手]** - 第三方审计(Epoch AI, *Where Autonomy Works*, 2026):*"Transfer is rarely demonstrated… Unless transfer is explicitly shown, it should not be assumed."* - **反向审计(对我们不利,照写)**:**没有找到**任何一篇让"更好的结构"在同任务同硬件上击败大 10× VLA 且带试验数与误差棒的论文;也**没有**任何"学出来的自我边界"在真实机器人上的性能优势实证。**规模派手里有真实曲线,我们手里只有机制假说。** - ⟹ 战略推论:最小实验落在**公开的证据真空**里——不是因为难,是因为没人测。**要抢就抢测量协议**(并入/释放非对称阈值曲线;在线执行器失效的 episode-to-recovery 分布带误差棒),**不是抢架构**。架构上我们打不过规模,方法学上对方是空的。附带义务:10 万形态域随机化是必须超过的**摊销基线**,所以最小实验的损伤类型必须可证地落在任何预训练覆盖之外。 ### 3 · 仿真 D0 —— 满足 blanket 条件的非平凡划分存在吗 Rig:MuJoCo 3.10,2–3 DoF 臂,可拆卸工具,可切断执行器,风驱动球;N=34 条标量通道(9 本体感 + 10 触觉 + 12 视觉关键点 + 3 动作回读)。动作 = OU 相关噪声 babbling,**不学策略**,这样任何滞回都不能归因于学到的控制器。全部 CPU,`mujoco==3.10.0`,Python 3.12。 **A · D0-pos(阳性对照,通过)。** 手装双阱,零扫速外推 `A₀ = +0.36 … +0.62`(随拟合形式变动)。线性阴性档:`|A₀| ≤ 0.008` 且符号随拟合形式翻转 ⟹ 与 0 不可区分。**[实测]** - **判据修正(我们自己揪出来的)**:阴性档的 bootstrap CI **也不含 0**,但那只反映种子噪声(SEM ~1e-4),没反映"只有 4–7 个扫速点、β 与 c 简并"的系统偏差。**正确判据不是 CI,而是把 0.008 当分辨底,看阳性档高出它多少倍**(实测 44–75×)。 - 解析锚点复现:`h_c = 0.38490`、`x* = ±0.57735`;数值 `h_up = +0.3950`、`h_dn = −0.3939`(误差 2.6%)。鞍结标度 `r^{2/3}`、`r^{1/3}` 均在 5% 内复现。直接双稳共存 1.00/1.00 vs 阴性档 0.47/0.43。 - ⟹ **仪器有效**,后续阴性结果可归因于系统而非尺子。 **B · D0 v1 —— 非平凡 blanket 划分"存在",但这个存在几乎无信息量。[实测]** - 是:22 self / 10 blanket / 2 world 的划分给 `Δ_peek = −0.0032 ± 0.0016`(6/6 种子全负)⟹ 加回 world 通道不改善对 self 的预测 ⟹ 滞回讨论**有对象**。 - 但:故意泄漏的标定划分只给 +0.012,动态范围 0.015 **小于**跨种子 std 0.021。⟹ 这是 Aguilera 那把刀在我们 rig 上的形态,但**方向相反**:不是"blanket 条件只在极窄参数区成立",而是**太容易成立以至于选不出东西**。 - 退化吸引域压倒性:72 次随机重启 **87% 塌到 |S|=1**,且每次都是最安静最好预测的那个 taxel。可算诊断:唯一把可控通道拉进 self 的项实测 ≤0.007,而划分间闭包损失差 ~0.3 ⟹ 系数差约 1.5 个数量级;修法是**重标定**不是换搜索算法。 **C · D0 v2 —— 这一阶段把我们自己下的诊断证伪了,这是最该留下的一条。[实测]** - **v1 的诊断("ratio<1 是因为接触率太低")——我们下的、并当前提派下去的——是错的。** ratio 低是**协议缺陷**(预测视界 K=10 太短 + 岭回归 λ 网格上界截断),不是 rig 缺陷:修好协议后**两个 rig 都过 3**(v2 rig 3.64;被我们判为"坏"的 v1 rig 3.94)。根因:泄漏标定探针探的是**臂→工具**的泄漏,**这条路径根本不经过球**,对接触率不敏感。**我们把"探针的信噪比"当"rig 耦合强度"读了。** - **两件事都真但必须分开记**:协议缺陷(已修,ratio 3.08 → 3.64);**rig 缺陷真实存在但 `Δ_peek` 看不见**,由**估计器无关的反事实探针**独立证实:同动作序列只换 wind_seed,轨迹分歧 **D: 0.045 → 0.203(4.5×)**。 - rig 第二层根因(v1 没看出来):干扰球 y 方向无恢复力 → 长 rollout 漂出平面,接触率随时长单调衰减(2000 步 3.8% → 12000 步 0.56%)。加 y 系绳 + 球移进臂实际占据区后接触率 **27%**(诚实标注:超出 5–20% 目标带),且不再随时长衰减。 - **γ_c 从数据自动标定**:`γ_c = std(L_clo)/std(可控性项) = 0.0826/0.00374 ≈ 22.1`。规格默认 1.0 **在量纲上就是错的**。 - 实测坑:`actuatorfrc ≡ ctrl`(MuJoCo motor 力矩读数就是指令本身,不剔除会让可控性增益整体塌到 ≤0.0018,看起来像"动作没用");peek-gain **在 K=1 上结构性瞎**(20 ms 后每通道都能被自己滞后外推);固定 λ 跨不同列数不可比,必须逐预测器在 val 上选 λ。 - 两条与规格担心方向相反的测量学发现:**(i)** 规格让辅助预测器 `f⁺` 用**更小**的 λ 反而制造假阳性——`f⁺` 过拟合 → 负 `Δ_peek` → **假装 blanket 成立**;**(ii)** 连续时序切分 + 非平稳 babbling 会**伪造结论**(`Δ_peek` 一度到 −0.46),必须分块交错。⟹ **`Δ_peek` 对估计器正则化与数据切分比对划分本身更敏感;任何用它下结论的实验必须先跑一个故意泄漏的标定划分当阳性对照。** ### 4 · 仿真 D1 —— 环增益判据可达吗 ``` PYTHONPATH=. python scripts/run_d1_gain.py --n-rollouts 6 --steps 12000 --controls-seeds 2 PYTHONPATH=. python scripts/run_d1_featureclass.py --n-rollouts 6 ``` v2 rig 默认档 + D0-v2 协议(K=20、λ 网格到 1e2、分块交错 2:1:1、winsorize ±8σ、逐预测器在 val 上选 λ)。`γ_c` 逐 rollout 自动标定,6 条均值 **28.75**。 **4.1 Δπ 量的是什么。** 回路是 `x_i↑ ⟹ 通道 i 进入 z 的预测目标 ⟹ ε_i↓ ⟹ π_i↑ ⟹ 驱动放大`。决定 π 的**不是标签**,而是"通道 i 的历史在不在 z 里"。所以做**同通道反事实**:`ε_i^in`(用 (S∪B)∪{i} 的特征预测 `c_i(t+K)`)vs `ε_i^out`(用 (S∪B)\{i})。取 `π = 1/(1+ε/ε₀)`,相对摆幅 `s ≡ Δπ/π_max` 对 ε₀ 单调,上确界 `s = 1 − ε_in/ε_out` 与 ε₀ 无关。下面所有 `s` 都是这个上确界,即**最有利于理论**的读法。**[实测]** - **规格字面读法(self 通道的 π 减 world 通道的 π)高估 7 倍**:0.233 vs 同通道反事实 0.0329 ± 0.0037。差别全部来自**通道身份**(taxel 本来就好预测、球本来就不可预测),不是回路里那个量。 **4.2 实测值**(6 seeds × 12000 控制步)。test 归一化 MSE(通道已标准化 ⟹ ε≈1 = 没学到): | | ε_in | ε_out | ε_own | |---|---|---|---| | 边界内 32 通道 | 0.500 | 0.551 | 0.534 | | 干扰球 2 通道 | 1.026 | 1.018 | 0.947 | 相对摆幅 `s = 1 − ε_in/ε_out`(均值 ± std),以及在保信息的表示变换下: | 组 | s(笛卡尔) | s(极坐标) | |---|---|---| | 边界内均值 | **0.0787 ± 0.0071** | 0.0489 | | 本体感 | 0.2291 | 0.0379 | | 触觉(臂) | 0.0016 | 0.0009 | | 触觉(工具) | 0.0086 | 0.0134 | | 关键点(臂) | 0.0571 | 0.1779 | | 关键点(工具) | 0.0130 | 0.0257 | | 关键点(球,world) | −0.0096 ± 0.0248 | 0.2189(假象) | | 动作 | 0.0104 | 0.0062 | **特征类对照(必做,否则会把假象当信号)。** `q1` 在笛卡尔下 s = 0.9844 ± 0.0135,看着救命。但 `q1` 在几何上完全由关键点 1 决定(`kp1 = 0.25·(cos q1, sin q1)`),只是这个关系是 `atan2`,线性+平方的特征类表示不了。把六个关键点换成**极坐标**(可逆变换,信息含量完全相同),`q1` 塌到 **0.0072 ± 0.0171**;`q3` 0.624→0.183;`q2` 0.421→0.114。反向也成立(极坐标引入自己的假象:kp1 半径恒为 0.25 ⟹ 退化;球角度 unwrap 漂移 ⟹ ball_x 0.46 ± 0.44)。**表示不变的读数:最大 s ≈ 0.18,边界内均值 ≈ 0.05–0.08,工具相关通道 0.002–0.045。** 估计器双向对照(2 seeds):换成白噪声 ⟹ s = −1.4e-5 / +1.7e-6(≈0,估计器不造摆幅);换成无关的 AR(1) ρ=0.98 ⟹ s = 0.373 / 0.368,理论 0.446(回收 83%)⟹ 估计器是**衰减的**(低估约 17%),对"摆幅不够"这个结论是**保守**的。 **4.3 规格 §3.3 的字面形式在数学上不可能三稳**(不是数值问题)。`τẋ = −γx + κ·π(x)·(g−c)`、`π(x) = π₀ + Δπ·σ(x/τ_π)` 时,三个自洽解要求同时 `F(0) = κπ₀(g−c)` 与 `F′(0) = κΔπ(g−c)/(4τ_π) − γ > 0`;π₀>0 ⟹ 这两条要求 `(g−c)` **反号,互斥**。数值:κ = 1/10/100 ⟹ G = 0.08/0.82/8.22,不动点 **1/0/0**,从没有 3;回归测试把 κ 推到 1e4(G≈1e3)仍 <3。**[实测]** **最小修补**(D2 用的就是它):让精度门只乘在**收益**上、不乘成本 —— `τẋ = −γx + κ(π(x)·g − c)`。理由是结构性的而非凑数:目标里的信息瓶颈与 blanket 成本项本就与预测精度无关,而可控性增益才是被精度加权的那个。G 的表达式不变;τ_π = 0.1、κ = 30 给出 G = 2.47 与 **3** 个不动点。 **4.4 可达性,以及判据本身的降级。** 归一 π_max = 1,记 `A ≡ κg/γ`: `G = A·s/(4τ_π)`,两支路分离度 `= A·s = 4τ_π·G`。 | τ_π | s = 0.183(最好的自身通道) | s = 0.0787(边界内均值) | s = 0.0254(最好的工具通道) | |---|---|---|---| | 0.05 | A ≥ 1.1 | A ≥ 2.5 | A ≥ 7.9 | | 0.10 | A ≥ 2.2 | A ≥ 5.1 | A ≥ 15.7 | | 0.20 | A ≥ 4.4 | A ≥ 10.2 | A ≥ 31.5 | - **结论 1**:是,存在合法参数使 G > 1,不需要把任何参数推到荒谬量级。瓶颈是 `s` 与 `τ_π` 的比;κ/γ/(g−c) 单独都不是瓶颈(只以 `A` 出现)。 - **结论 2(更重要)**:`m^S` 越过 0.5 需要 `|x| > ln2/2 = 0.347`,即分离度 ≳0.7,即 `G ≳ 0.175/τ_π`,在 τ_π ≤ 0.175 时**自动 >1**。**"G>1" 与 "边界真的翻转了" 几乎是同一句话**,筛不掉任何东西。规格称它"整条线的生死判据"是**高估**。 - **结论 3(真正的卡点)—— fine-tuning 容差 = s。** 两支路 `x*_hi = A(1 − c/g)`、`x*_lo = A(1 − s − c/g)` ⟹ `c/g` 必须落进 `[1−s, 1]`,即**成本要抵消收益到相对精度 s**。用本仓目标函数自己的系数(γ_c = 28.75 自动标定,β_I = 0.30): | 组 | g_i | c/g 实测 | 需要的窗口 [1−s, 1] | 偏离(窗口宽为单位) | |---|---|---|---|---| | 本体感 | 0.0195 | 0.535 | [0.771, 1] | **2.0**(笛卡尔)/ 12.3(极坐标) | | 关键点(臂) | 0.0608 | 0.172 | [0.943, 1] | 14.5 / 4.7 | | **关键点(工具)** | 0.0698 | **0.150** | **[0.987, 1]** | **65 / 33** | | 动作 | 0.0469 | 0.222 | [0.990, 1] | 75 | | 触觉(臂) | 0.0074 | 1.401 | [0.998, 1] | 251 | | 触觉(工具) | 0.0021 | 5.041 | [0.991, 1] | 469 | | 关键点(球,world) | 0.0021 | 5.038 | — | 420(正确地在窗口外) | **没有一条通道落在窗口里**;最近的差 2 个窗口宽,而理论的旗舰现象(工具并入)差 **33–469** 个。**目标函数里没有任何机制在把 c/g 往 1 拉。** ⟹ 滞回**不是从我们提的目标函数里长出来的**;它要求一个该目标不提供的、约 2% 精度的手调比值。 ### 5 · 仿真 D2 —— 涌现滞回,以及对照做了什么 ``` PYTHONPATH=. python scripts/run_d2_emergent.py --n-seeds 30 --n-boot 1000 \ --s-self 0.183 --s-ball 0.0254 # 702 s PYTHONPATH=. python scripts/run_d2_emergent.py --n-seeds 30 --n-boot 1000 \ --s-self 0.183 --s-ball 0.0254 --only E_field_gated # 164 s ``` 动力学(手装立方项 `a = 0`,已完全移除): τ_x ẋ = −γx + γ·A·s·(σ(x/τ_π) − ½) + γ·h , A = κg/γ,s = Δπ/π_max 参数直接取 D1 实测值:`s_self = 0.183`、`s_ball = 0.0254`、`τ_π = 0.1`、目标 `G = 3` ⟹ `A_self = 6.56`(合法区内),噪声 σ = 0.01,τ_x = 1,`h` 三角扫 −1 → +1 → −1。 四档:**A** 门开、a=0(主张,G=3.00);**B** π 冻成常数、其余一字不改(阴性对照,G=0);**C** 干扰球,同款慢变量、同款门、同样的 A,只把 s 换成球的实测值(G=0.42);**D** 干扰球,把 A 抬到 47.2 使 G 与 A 档相同(G=3.00)。 **5.1 扫速控制 —— 主结果。** 环面积 `A(r)` 均值(30 种子),`r = τ_x/T_sweep` 跨四个数量级: | r | 1e-1 | 3.3e-2 | 1e-2 | 3.3e-3 | 1e-3 | 3.3e-4 | 1e-4 | |---|---|---|---|---|---|---|---| | **A** 门开 | 0.9571 | 0.6561 | 0.3950 | 0.2974 | 0.2520 | 0.2353 | **0.2268** | | **B** 门冻结 | 0.3831 | 0.1651 | 0.0531 | 0.0181 | 0.0054 | 0.0018 | **0.0006** | | **C** 球(实测 s) | 0.4512 | 0.1953 | 0.0624 | 0.0211 | 0.0064 | 0.0021 | **0.0007** | | **D** 球(拉平 G) | 0.9571 | 0.6561 | 0.3950 | 0.2974 | 0.2520 | 0.2353 | **0.2268** | 零扫速外推 `A(r) = A₀ + c·r^β`(β profile 网格 + 30 种子 bootstrap): | 档 | A₀(β 自由) | β | A₀(β = 1) | 对分辨底 0.008 的倍数 | 判决 | |---|---|---|---|---|---| | **A** | **+0.1932** | 0.54 | **+0.2771** | **24 – 35×** | 真双稳 | | B | −0.0044 | 0.79 | +0.0081 | ≤1.0× | 只是弛豫滞后 | | C | −0.0052 | 0.79 | +0.0097 | ≤1.2× | 只是弛豫滞后 | | **D** | **+0.1932** | 0.54 | **+0.2771** | **24 – 35×** | 真双稳 | B 档的两个 `A₀` 与 D0-pos 线性档**逐位相同**,因为 `gate_const=True` 在数学上就退化成它(有回归测试断言),这同时是分辨底 0.008 的重新确认。最慢档阈值:A 档 `h_up = +0.2674`、`h_dn = −0.2099` ⟹ Δ = +0.477(真环);B 档 `h_up = +0.3359`、`h_dn = +0.3583` ⟹ Δ = −0.022(两者交叉,根本没有环)。 **5.2 直接双稳判定**(比画环更强):`h` 固定 0,两支各 30 种子,`T_hold = 1000 τ_x`: | 档 | 下支 stay-rate | 上支 | 支路分离度 | |---|---|---|---| | **A** 门开 | **1.00** | **1.00** | **1.193** | | B 门冻结 | 0.47 | 0.43 | 0.0012 | | C 球(实测 s) | 0.43 | 0.40 | 0.0021 | | **D** 球(拉平 G) | **1.00** | **1.00** | **1.193** | B/C 的 0.4x 不是"留住了一半",而是两支塌到同一点、`sign(x_end)` 由噪声决定(分离度 0.001–0.002 才是实质)。**[实测]** **5.3 ⭐ 特异性对照 —— 叙事被抽掉一半。** - **C 档(诚实的一半)**:给干扰球配同款慢变量 + 同款门 + 同样的驱动/漏项比,只把精度摆幅换成球的**实测值** ⟹ G 掉到 0.42,`A₀` 落回分辨底,共存判定 0.43/0.40 失败。**滞回不出现。** ✅ - **D 档(致命的一半)**:只把球的 A 从 6.56 抬到 47.2 使 G = 3 ⟹ **滞回的每一个数字与 A 档逐位相同**(A₀ = +0.1932、分离度 1.193、τ 比 0.806)。 不是巧合,是可解析的:归一化后该 ODE **只依赖 (G, τ_π)**,驱动幅度 `A·s = 4τ_π G`。**s 只决定"要多大的 A",不决定"会不会滞回"。** > ⟹ **滞回是"带门控的慢变量"的通用性质,不是自我的签名。** 在本 rig 上,「自我边界的滞回」与「任何被同样门控的慢变量的滞回」**在动力学上不可区分**,区分只来自一个可被自由参数吸收的 7 倍量级差。这**不**证伪 H-B(A 档确实在 a=0 下给出率无关滞回),但它取消了「滞回 ⟹ 这是自我边界」这一步推理。要恢复特异性必须让 A 也由数据定死 —— 这正好接回 §4.4 结论 3。 **5.4 时间常数非对称预测 —— 被证伪两次。** 在阈值之外固定 `h = ±1.5·|h_up|`,从相反支路做阶跃,取走完 (1 − 1/e) 的时间。两个观测量都报,因为 1-D 约化 `m^S(x) = e^x/(e^x + e^{−|x|} + e^{−x})` 在 `x → −x` 下**本身不对称**(x=0 处 m=1/3 而非 1/2),即使动力学完全对称也会量出 τ_off ≠ τ_on: | 档 | τ_on (m^S) | τ_off (m^S) | 比值 (m^S) | τ_on (x) | τ_off (x) | **比值 (x)** | |---|---|---|---|---|---|---| | **A** 门开 | 4.24±0.11 | 3.41±0.09 | 0.806 | 4.00±0.09 | 4.03±0.10 | **1.008** | | B 门冻结 | 1.50±0.03 | 0.81±0.02 | 0.537 | 1.06±0.05 | 1.04±0.05 | 0.981 | | C 球 | 1.81±0.04 | 1.00±0.00 | 0.554 | 1.29±0.03 | 1.28±0.04 | 0.992 | 规格预测 `τ_off > τ_on`(比值 >1),且 π 常数化后应 →1。实测:慢变量上比值 **1.008 = 完全没有非对称**;`m^S` 上 0.806(**小于** 1,释放比并入**快**),而 π 常数化给 0.537,离 1 **更远** ⟹ 规格说的"常数化后 →1"**方向也错**。根因可算:漂移 `F(x) = −γx + γAs(σ(x/τ_π) − ½) + γh` 在 `(x,h) → (−x,−h)` 下**严格奇对称** ⟹ on/off 结构上必然对称。**精度门本身不产生时间常数非对称。** **E 档:把门乘到整个驱动上**(规格真正想要的机制):`τẋ = −γx + γA[π(x)(1 + w·h) − π(0)]`,w = 0.4,其余同 A 档。成本仍不被门控,否则按 §4.3 三稳数学上不可能。 | 量 | E 档 | 对比 A 档 | |---|---|---| | A(r):r = 1e-1 → 1e-4 | 1.0049 → **0.0995** | 0.9571 → 0.2268 | | A₀(β 自由 = 0.75) | **+0.0916** [CI 0.0915, 0.0918] | +0.1932 | | A₀(β = 1) | **+0.1289** | +0.2771 | | 对分辨底的倍数 | **11 – 16×** | 24 – 35× | | 共存 / 分离度 | 1.00 / 1.00,1.193 | 1.00 / 1.00,1.193 | | 阈值(最慢档) | `h_up = +0.1245`,`h_dn = −0.0819` | +0.2674 / −0.2099 | | **τ_on / τ_off (x)** | 3.91 / 3.10 ⟹ **0.794** | 4.00 / 4.03 ⟹ 1.008 | 非对称终于出现(0.794 ≠ 1.008),**但方向与原主张相反**:释放比并入**快**,且离开的阈值(−0.0819)在幅值上**低于**加入的阈值(+0.1245)。可算原因:`G(h) = G₀(1 + w·h)`,h>0 时环增益高、靠近折点、临界慢化 ⟹ 并入慢;h<0 时增益低、系统更线性 ⟹ 释放快。原推理("已在 self 内 ⟹ 精度高 ⟹ 反向证据要更大才能赶出去")默认了**精度只放大支持归属的证据**,这在任何"π 乘在驱动上"的实现里都不成立。要得到"进来容易出去难",必须额外假设一条**不对称**的证据加权 —— 那又是一件手装的东西。 **5.5 本轮未闭合。[未做]** 短训编码器没做(D1 用岭回归代理量 ε,所以 Δπ 仍带特征类依赖——**最大未闭合项**);D2 跑在 1-D 约化 ODE 上而非 34 通道全 rig;`k_grip` 的力-位移标定未做,所以规格的主参数轴(握持刚度)从未被扫过(D2 扫的是抽象外场);separatrix / 最小逃逸扰动幅度非对称未测;Wilcoxon 配对检验与重测 σ₀ 未做;D3–D5、pyDMBD 对照 baseline、baseline a1/a2/b/c/d 未跑。pytest 25 passed。 ### 6 · 人类侧:两条战线,都不利 **6.1 为什么决定性数据不存在。** 不是没人想到:所有参数化橡胶手实验一律随机/伪随机顺序,因为心理物理的默认卫生规范把**序列效应当污染而非信号**。范式的设计目的就是抹掉我们要的那个量。**[从方法部分推断]** **6.2 已存在的前序依赖实测,符号与滞回相反。** Zhang, Ma & Hommel 2015 (*Frontiers in Psychology* 6:1659, N=34):同一"中距离"条件放在近距离之后 vs 远距离之后,**远→中所有权显著更高**(M = 3.768 vs 3.056,F(1,33) = 39.82,η²p = 0.547)。滞回预测的是近→中更高。按 Liaci & Kornmeier 2018 的符号约定(正 = 迟滞/priming,负 = 适应),这是**负号 = 适应侧**。不是干净证伪(7 点问卷、参数是距离非同步性、只有两步),但它是离目标 bit 最近的已发表数据,且不站在我们这边。**[引用]** **6.3 我们从候选表里捞出来并亲手跑完的更大样本复制。** OSF `bsfwu`(Tsuji & Imaizumi,橡胶手顺序效应,Lush 2020 线),**N = 185**,带 `CondOrder` 列(Syn1st / Asyn1st)—— 正是同一种路径操纵的**同步性版本**,样本大 5 倍。**[我们在其公开数据上实测]** - 同一 Async 条件:Syn1st 组 M = −0.455 vs Asyn1st 组 M = +0.436,差 **−0.891**,95%CI [−1.232, −0.550],t(183) = −5.16,**p = 6.7e-7**,Hedges d = −0.756;Order × Condition F(1,183) = 19.17,p = 2.0e-5,η²p = 0.095。先经历强条件 → 后续弱条件评分**更低** = 适应/对比侧,与 6.2 同向。 - **三条我们自己查问卷原文查出来的降权**:(a) DV 不是亲历的所有权,是**期望**——被试只看 32 s 视频,答"若你是视频里的被试,你预期体验到多少",且全部未参与过该程序 ⟹ 这是关于所有权的**期望**的顺序效应,几乎必然含判断/锚定成分;(b) 控制条目也有同向显著效应(−0.313,p = 0.0435)⟹ 需求特征污染确实存在,**但不是全无特异性**:Order × Scale F(1,183) = 22.42,p = 4.4e-6,η²p = 0.109,illusion 条目上的效应约为 control 的 **2.8 倍** ⟹ 真实成分存在但被同向偏置污染,本数据无法干净分离;(c) N 实为 185 非 184。 **6.4 同步性侧管线 dry-run —— 不是 embodiment 证据,但产出一条会决定成败的技术事实。[实测]** - **置换零分布的中心不在 0,在 +0.191。** 机制已定位(`soa²` 项 + block 级反应率异质性;99.2% 的 lag-1 配对落在同一 block 内,block 间反应率差异被原样保留)⟹ **直接对 0 做 t 检验会把符号读反**:t 检验说"+0.048,p = 0.47,无效应",对上正确零模型后是**显著为负**。两个独立零模型(block 内重排 / 历史序列循环平移)给出同向零中心(+0.191 / +0.158)⟹ 稳。**纪律:拿到所有权数据后,先确认该数据上零分布中心在哪,再谈显著性。** - 真实数字分裂但已解释:`hev9z`(同时性判断,26k 试次)PSS 被前一试次拖动 **+16.1 ms** CI [12.3, 19.9] = 适应侧;`4ukzx` 选择重复 **+0.447** = priming 侧。原因不是效应不稳,而是 `4ukzx` 实为 **TOJ 型单调任务**(β_soa = +7.6,β_soa² = −0.16),与 SJ 的峰形曲线不同族。 - 两份数据都是 lag-1/lag-2 显著、**lag-3 归零** ⟹ 只有 1–2 试次短尾,**无多试次 latch**(弱证据,非 embodiment)。 - **给未来自建实验的硬设计约束(本轮最值钱的可迁移物)**:**若被操纵参数的符号与二元反应语义直接对应(如 TOJ 的"谁先"),迟滞与适应在数学上不可分辨。** ⟹ 所有权的逐试次 DV **必须保持"是/否"**,绝不能写成"哪只手更像我的"。另:若自建实验是单调扫描,自由重排的置换零模型**不合法**,必须改成扫描方向内保序的置换。 **6.5 数据搜寻:什么不存在,以及重复了三次的一个错误。[实测]** - 逐试次所有权数据:OSF(16 组关键词)+ Zenodo(10 组)+ OpenNeuro(**1831 个数据集全量普查**)三处**零命中**。 - 三份声明"可用"的数据被亲手打开:eLife 11:e77221 (Chancel/Ehrsson/Ma 2022) 的 fig1 source data 是**被试 × 条件的计数矩阵**(S1,N0_−500 = 12 试次中 0 次"是",17 行 × 22 列),fig2/fig3 是模型参数估计;OSF `spkvu`(Lanfranco 2024)经 API 递归列举**全库只有 3 个文件**(`bias.csv` 2578 B、`dprime.csv` 2432 B、`readme.rtf`),均为被试 × 条件聚合。**三份都没有 trial index、没有顺序列** ⟹ 前序依赖在结构上不可计算。(验证:`python3 osf_list.py spkvu 3`;`openpyxl.load_workbook` 逐 sheet 打印;2026-07-28。) - **不要相信任务名,要打开列。** `ds001785` 的 `stimamp` 实测严格单调上行 0.10→1.94、下行 2.50→2.30 —— 四轮里第一次见到真正的上下行扫描,看名字会直接判"滞回数据到手",打开发现响应列 **93 行全是 `n/a`**。同理 `ds003922` 的响应根本不在 BIDS events 里而在 `sourcedata/`。 - **对我们自己搜寻 agent 过度结论的纠正**:它判"逐试次结构在**范式层面**就不存在(经典 DV 是 block 末问卷)"——**对 Ehrsson 系错**:eLife 77221 的 `RHIdetect` 表就是**每条件 12 试次的逐试次二元所有权判定**(计数 0–12 即证据)⟹ **范式存在且成熟,只是原始日志从不上传。** ⟹ 正确推论不是"自建范式",而是**向作者索要**——当前性价比最高的一步,且需要一次具名的人类请求。 - 四项判据全中的新靶子:OSF `4dzrg` —— 视触 TOJ × 观察人手,40 人 × 840 试次 ≈ 33,600,列 `subNum blk cond trialNum vis_tac_first SOA resp corr RT`(range-read zip 尾部取出,未下 384 MB)。OpenNeuro `ds004561`(Veillette/Lopes/Nusbaum,CC0,EEG,23 人)是唯一一份**身体自我范畴**、逐试次二元判定 + 参数 + 顺序齐全的公开数据(`onset duration trial_type trial intensity latency rt pressed_first agency`,250 试次/人)——**但它是 agency 不是 ownership**。 - 未实测:NEMAR 自身接口(零命中是推论非实测)、Dryad、figshare、GitHub。未检索:麻木肢体脱落阈值;**全身错觉 / 化身错觉侧**(VR 化程度更高,可能藏着现成扫描数据,是下一轮性价比最高的方向)。 **6.6 可照抄的方法学模板,以及又一条对我们不利的。[引用]** - T1 —— Martin, Kösem & van Wassenhove 2015 (*PLoS ONE* 10(3):e0119365),视听同步性的滞回:**同一把刀已经用在"同步性"这个参数上**,只差把模态换成视-触、把响应换成"这是我的手吗"。它未变扫速,未建立率无关性。 - T2 —— Liaci & Kornmeier 2018:操作定义 `hd = 拐点(上升) − 拐点(下降)`,外加第三条随机序列基线臂。 - T3 —— 立体视 / 双眼竞争:唯一内建**多扫速**的文献,正好能执行我们的"零扫速外推"判据。 - **对我们不利**:Schwiedrzik et al. 2014 (*Cerebral Cortex*) 证迟滞与适应是**独立可加、脑区分离**的两个成分 ⟹ **测到零净滞回 ≠ 没有 latch**(两成分可抵消)。必须写进预注册,否则我们既无法用阴性结果证伪自己,也不能用它当免罪符。 - 成本估算(我们自己的,非文献数字):VR 自建实验设备 + 被试 ¥5k–8k、开发 3–5 人日、N = 24、采集 2–3 周。**真门槛是被试招募与伦理挂靠,不是设备。** VR 不是省钱选项,是**唯一**能满足判据的选项:实体橡胶手的延迟由实验者的手决定,步长做不细、扫速控不准,多扫速外推几乎无法执行;VR 里延迟是纯软件。视-动相关即可诱发,不必做触觉硬件。 - **人类侧当前总账**:两个独立操纵维度(空间距离 / 时间同步性)、两个不同样本(N = 34 / N = 185)、两种测量(亲历 / 期望)——**全部落在适应侧,无一出现迟滞方向**。这种收敛性比任何单条更有份量。但两者都是 block 级问卷,原理上无法区分"滞回"与"单次对比锚定" ⟹ 真正判别性的数据(逐试次 + 变扫速 + 亲历测量)**仍然缺失**。**这条线目前既无支持证据,也未被真正证伪;只是两个可用数据点都指向不利方向。** - 决策更新:**暂不投 VR 自建**(先验变差,期望价值下降)。优先级 = ① 向持有逐试次日志的实验室索要(唯一能给判别性数据的现成来源,需具名请求)② 仿真侧继续(不依赖生物学前提)。 ### 7 · 台账:四个 falsifier 各自的状态 | # | Falsifier | 状态 | |---|---|---| | 1 | **工具并入的非对称阈值**(并入慢、释放更慢)。中途收紧:测到非零环**不算**,必须**环面积外推至零扫速仍 > 0**,否则会拿自证的率依赖假环当成功。 | **一半通过,一半被证伪。** a=0 下率无关滞回:A₀ = +0.19…+0.28,分辨底的 24–35 倍,共存 1.00/1.00 ⟹ **通过**。非对称方向:**被证伪两次**(§5.4)。自我特异性:**被证伪**(§5.3)。 | | 2 | **损伤测试**:中途废掉一个执行器,μ 应在 N 个 episode 内重解出不同但闭合的划分且不重训策略,并显著快于微调后的 VLA baseline。 | **未跑。** | | 3 | **感官替代迁移**:视觉训 μ → 换触觉阵列编码;行为若不收敛到视觉观测者可达集,则 μ 实为模态绑定,理论超发。 | **未跑。** | | 4 | **Bitter lesson**:若纯 VLA 放大 10× 就在工具使用 + 损伤恢复上打平,则 μ 从不是瓶颈。 | **没被触发 —— 但原因是证据真空,不是我们赢了**(§2)。举证责任仍在我们这边。 | **审计过程中新增、且都已开火的 falsifier**: - **D0**(因为我们对 FEP 文献那一击同样砍我们):先证明仿真臂上存在满足 blanket 条件的非平凡划分。**通过,但几乎无信息量**(§3B)。 - **D0′**(人类侧):自建同步性上下行扫描。**推迟** —— 先验变差到不值得投(§6)。 ### 8 · 当前主张状态 原 wedge:**μ 有滞回,由二阶自维持驱动,且并入慢、释放更慢。** | 成分 | 状态 | |---|---| | 滞回不靠手装非线性即可涌现 | ✅ **活** | | 该滞回对自我边界有特异性 | ❌ 是带门控慢变量的通用性质 | | 释放比并入更慢 | ❌ 两种实现下都反了 | | 它从提出的目标函数里自然长出,不需 fine-tune | ❌ 需约 2% 的手调比值 | 四条支柱倒三条,**且全部由我们自己的仪器打倒** —— 这正是这套流程该干的事。 **被重述的问题**(这才是本线真正的产出):滞回很廉价,任何带门控的慢变量都有。**什么让一条边界的滞回是关于这条边界本身的?** 要恢复特异性必须把自由幅度 `A` 也由数据钉死(例如锁到目标函数自己给出的值),而这与 §4.4 结论 3 是同一个问题,两者都没解。 还有两条更早的威胁未答:"闭包违反"是否真的不等于 ELBO(§2);以及**原始像素不能当通道**这个让步 —— μ 是通道上的划分,指称物身份必须稳定,这是本线最大的、已被显式记账而非藏起来的作弊点。 **未 ratify。不是结论。之所以发布,是因为阴性的那一半才是有证据价值的那一半。** ### 9 · 参考文献 - Aguilera, Millidge, Tschantz & Buckley (2022). How particular is the physics of the free energy principle? *Physics of Life Reviews* 40:24–50. - Biehl, Pollock & Kanai (2021). A technical critique of some parts of the free energy principle. *Entropy* 23(3):293. - Chancel, Ehrsson & Ma (2022). Uncertainty-based inference of a common cause for body ownership. *eLife* 11:e77221. - Crutchfield & Young (1989);Shalizi & Crutchfield (2001) —— 因果态 / computational mechanics。 - Cully, Clune, Tarapore & Mouret (2015). Robots that can adapt like animals. *Nature* 521:503–507. - Friston, Da Costa & Parr (2021). Some interesting observations on the free energy principle. *Entropy* 23(8):1076. - Holmes (2012). Does tool use extend peripersonal space? *Experimental Brain Research* 218:273–282. - Klyubin, Polani & Nehaniv (2005) —— empowerment。 - Liaci & Kornmeier (2018) —— 知觉序列中的迟滞与适应。 - Martin, Kösem & van Wassenhove (2015). Hysteresis in audiovisual synchrony perception. *PLoS ONE* 10(3):e0119365. - Maturana & Varela (1980) —— autopoiesis。 - Schwiedrzik, Ruff, Lazar, Leitner, Singer & Melloni (2014). Untangling perceptual memory: hysteresis and adaptation map into separate cortical networks. *Cerebral Cortex* 24(5):1152–1164. - Tononi & Koch (2015) —— 意识与复合体的边界。 - Zhang, Ma & Hommel (2015). *Frontiers in Psychology* 6:1659. - *Dynamic Markov Blanket Detection for Macroscopic Physics Discovery*, arXiv:2502.21217;代码 `github.com/bayesianempirimancer/pyDMBD`。 - Epoch AI (2026). *Where Autonomy Works: Evaluating Robot Capabilities in 2026*. https://epoch.ai/blog/where-autonomy-works-evaluating-robot-capabilities-in-2026 ### 10 · 站内相关 - [State 是一个闭包条件,不是给定的集合](https://machengshen.github.io/theory/state-as-closure.zh.md) —— 本提案训练信号所依托的闭包本体论。 - [两体业力:关系里的三种惯性](https://machengshen.github.io/theory/two-body-karma.md) —— 同一套双稳/卡扣语言用在两个耦合的心智上;那里关于"结构性押韵而非推导"的警告在这里同样适用。 - [本线的阶段日志](https://machengshen.github.io/theory/learning-the-self-boundary-log.zh.md) —— 各阶段分别推翻或确立了什么,按时间顺序。 --- ## Source: /theory/no-state-only-history.md (sha256:8173a3fdc66b, language: en) # No state, only history *A diagnosis of the Transformer, and one wedge worked far enough to be attacked · Macheng Shen · 2026-08-02 · amended 2026-08-04* **Correction first, so you can decide in thirty seconds whether to keep reading.** The title is deliberately sharper than the claim that survived review. A standard Transformer has activation state, computational state, and a transient workspace-like representational state. Anthropic's 2026 Jacobian-lens experiments make the last item concrete: verbalizable, flexibly routed intermediates occupy a sparse-frame object called J-space across an intermediate layer band. What vanilla inference still lacks is narrower: **a persistent, autonomous, path-dependent closure state whose own update rule maintains its projection/forgetting policy across steps**, except insofar as persistence is serialized into tokens, a KV cache, parameters, or an external memory system. That second-order question remains open. The concrete proposal, and the reason this note exists rather than being a list of directions: **test-time memory should be tested with a gate keyed to change in what the memory forgets, not only to surprise magnitude.** Surprise can be large for both noise and rare useful events; robust losses trade between them. Subspace change is a genuinely different signal, and under a wide spectral gap it can suppress isotropic perturbations without suppressing every large residual. **The signal separation is measured; the practical advantage is not** — E1 has been run (§7, figure included): across 6,480 synthetic updates the two signals decorrelate at Pearson −**0.08**, and at fixed ‖ΔM‖ the μΔ spread covers **100% of the observed range**. E2 must still decide whether this buys anything on a learned memory. Everything below carries its occupants and arXiv numbers. Two of my five original candidates did not survive a prior-art audit — one is a footnote, one is deleted. E1 has been run; E2 and the corrected E3 have not. J-space is an adjacent empirical result from another group, not our result. --- ## 1 · The diagnosis ### No state — only history "State", in the sense decision theory and statistical mechanics use the word, is not a record. It is what you get *after* choosing a projection: the quotient of history under "makes no difference to what I care about". Choosing a state is choosing what to forget, and the Mori–Zwanzig identity makes the bill explicit — whatever you project away comes back as a memory kernel plus noise. There is no free version. A KV cache **is computational state**: it is a transformed, addressable trace that changes later computation. What it is not, by itself, is an autonomously chosen closure state. Its retention policy is normally supplied by the serving system and context budget rather than maintained by a learned rule about which equivalence classes of history remain causally sufficient. The architecture does perform selective access inside the window; the narrower deficit is that the cross-step persistence and forgetting policy are not usually owned by the same state they define. ### No hysteresis — only a function For fixed parameters, fixed cached tensors, fixed numerical execution and fixed randomness, attention at position *t* is a function of its supplied state. That fact alone does not rule out hysteresis — every state-space model has a deterministic update law too. The relevant question is whether two paths can reach the same externally controlled input while leaving different **internally carried** states that persist when the control reverses. Three ways the past can hold the present, worth separating: - **Mud** — drag that only slows and fades, returning nothing. Everything is retained, but the retention does not change the dynamics. - **Spring** — inertia: stores and gives back, carries momentum. - **Latch** — bistability with a switching threshold: unmoved by small pushes, flips past a barrier, stays flipped when the push is removed. A linear system can draw a rate-dependent input–output loop through phase lag; it cannot by itself supply this rate-independent bistable latch. A standard feed-forward Transformer has a transformed history in the KV cache and a transient workspace across depth. It has no demonstrated autonomous latch at the architecture level: no persistent multistable variable with separately measurable switching thresholds whose value is carried across steps independently of replayed context. An agent harness, recurrent wrapper, test-time learner, or external memory can add one; the claim here is about vanilla inference, not every system containing a Transformer. ### No self-boundary — only a window Tokens the model generated and tokens the world inserted can occupy the same content representation space unless the harness or architecture supplies provenance. That shared channel is one contributor to prompt injection, not a complete causal explanation: instruction/data separation, tool permissions, parsing, and control-plane enforcement also matter. §5 records the architectures that already add part of the missing provenance boundary. ### Why write this down now Because the field just supplied a clean statement of the axis everyone is on. Behrouz et al., **"Memory Caching: RNNs with Growing Memory"** (arXiv:2602.24281, 27 Feb 2026) is best read as a *negative* result about that axis rather than a positive one about the method: to close the recall gap with Transformers, a recurrent model's effective memory capacity must **grow with sequence length**, and MC's contribution is an explicit interpolation between fixed and growing memory. Read as a constraint: if you want recall, you need an addressable set that grows with length, and no clever fixed-size compression escapes it. That locates Mamba, RWKV, Titans and MC primarily on one axis — *how much of the past to keep addressable, and how to compress it*. A first-order question about the **contents** of memory. The second-order question is about the **operator**: not "what should I remember" but "what makes something a variable I have to carry at all, and what changes that". The 2026 J-space result changes the map. It provides evidence for an access/workspace object inside a feed-forward Transformer: a corpus-averaged Jacobian lens identifies verbalizable intermediates that can be reported, modulated and flexibly routed. Mathematically, the reported J-space is **not** one fixed linear subspace or a single projector `Π_M`; it is a sparsity-bounded union of nonnegative cones generated by an overcomplete frame, and its definition depends on a sparsity level and a data distribution. That is adjacent evidence for a transient workspace, not yet a persistent operator that maintains its own cross-step closure. The missing object is smaller than this note originally claimed, but still missing. --- ## 2 · The mainline: gate on boundary change, not on surprise ### What is already occupied Fast-weight programmers (Schmidhuber 1992; Schlag et al. 2021), test-time training (Sun et al. 2024) and Titans (Behrouz et al., NeurIPS 2025) all update a memory at test time from a *surprise* signal — the gradient of a reconstruction loss on the key–value association. The strongest occupant is **Miras** (Behrouz et al., arXiv:2504.13173), which promotes "surprise" from a signal to a *design space*: ``` M_t = argmin_M ℓ(M; k_t, v_t) + λ · R(M, M_{t-1}) ╰─── attentional bias ───╯ ╰─ retention gate ─╯ ``` with ℓ ranging over ℓp / Huber / robust losses and R over KL, f-divergences, elastic net. Titans, Moneta, Yaad and Memora are points in it. Common surprise-gated implementations key update strength to ℓ, ‖ΔM‖ or a robust transformation of them. But the **Miras framework itself is broader**: its retention term R is a free design slot and could encode subspace geometry. The entry point is therefore not “Miras cannot see geometry”; it is the narrower question of whether geometry is useful in the write decision. Any proposed scalar error magnitude is already within the occupied design space. A geometric signal may be a useful Miras instance or composition rather than a new framework. The rest of this section establishes only that it contains information an update norm does not. ### μ as the induced projection Take the memory M (a matrix in the linear case). Reading is `read(M, q) = Mq`, so any component of *q* lying outside the row space of M is structurally discarded. **That discarding is the forgetting.** So define the induced projection ``` Π_M := projection onto the effective row space of M (the top-r right singular subspace) ``` Π_M is literally "which directions of the world this memory keeps" — the computable form of the projection μ that defines the boundary. Note the shift: μ is not the memory's *contents*; it is the forgetting policy the contents induce. ### The signal ``` μΔ_t := d_Gr( Π_{M_{t-1}} , Π_{M_t} ) ``` where `d_Gr` is the principal-angle (Grassmann) distance between the two *r*-dimensional subspaces, `d_Gr(A,B) = ‖(θ_1,…,θ_r)‖_2`. No full SVD is needed: maintain the top-*r* basis with randomized subspace iteration, O(d·r) per step — **the same order as the Miras gradient update**, so computability is not the obstacle. Then gate on it: `M_t = M_{t-1} + η · g(μΔ_t) · ΔM_t`. ### What E1 separates — and what it does not The claim to be established is not "μΔ is better" or "outside Miras". It is narrower: **there exist update pairs that no monotone function of update magnitude can distinguish, and that μΔ separates.** Two families, pointing in opposite directions: **(A) Large content change, zero boundary change.** *This is the family whose statement running E1 corrected, so it is worth stating what was wrong.* I originally wrote: choose ΔM whose row space lies inside the existing row space of M_{t-1}. **That condition is insufficient.** Writing ΔM = C·V_rᵀ with arbitrary C leaves cross-terms against the discarded block — in the V basis the Gram matrix of M+ΔM is not block-diagonal — so the top-r subspace still rotates. Measured, that naive construction gives a median μΔ of 7×10⁻², not zero. The correct construction is a change of *basis inside* the retained subspace (V_r → V_r Q for orthogonal Q) together with any rescaling of the retained singular values: the span is preserved by construction, at any magnitude. Measured over 1,440 such updates: **μΔ ≤ 1.6×10⁻⁷ rad (machine precision) at ‖ΔM‖ up to 3.2×10⁴.** So the family exists and does what the argument needs — but only under the stronger condition, and I had the condition wrong in the first version of this note. **(B) Tiny content change, boundary flip.** Take M_{t-1} with near-degenerate σ_r ≈ σ_{r+1} and apply a perturbation of size ‖ΔM‖ = ε that pushes direction *r*+1 past direction *r* into the top-*r* subspace. Then ‖ΔM‖ → 0, so a magnitude-only gate reports "negligible" — while `d_Gr → π/2`, an entire principal direction swapped out, so **μΔ is maximal**. A Miras instance with an appropriate geometric R need not miss it. An update norm is a content-magnitude statistic. μΔ measures the action of ΔM on the *spectral subspace* of M — a quantity on Gr(r,d). (A) and (B) show their level sets are transverse. **Measured (§7/E1, 6,480 updates): the two decorrelate at Pearson −0.08 / Spearman −0.35, and within a single narrow band of ‖ΔM‖ the μΔ values span 100% of the observed range.** The sharpest single pair: ‖ΔM‖ = 3.2×10⁴ with μΔ = 6×10⁻⁸, against ‖ΔM‖ = 1.0×10⁻² with μΔ = 5.7×10⁻³ — a factor of 3×10⁶ in ‖ΔM‖ with the μΔ ordering reversed. **So μΔ is not a monotone function of ‖ΔM‖.** That much is settled. **What is *not* settled, and this is the weakest joint in the note — stated in the body rather than buried in a caveat.** The step from "not a monotone function of ‖ΔM‖" to "outside the Miras design space" does **not** go through. Miras's retention gate `R(M, M_{t-1})` is a *free slot*: nothing in the framework forbids defining `R := d_Gr`. Do that and this proposal becomes *Miras with a particular R*, not an alternative to it. The remaining distinction is about placing geometry in the write decision rather than using it only as a retention penalty. **That is a composition/framing question, not a mathematical separation.** We put P = 0.40 on the stronger claim surviving (§9); it would rise only with a formal exclusion result that does not currently exist. What *is* second-order, and does survive: μΔ measures the update's effect on the *forgetting policy* rather than on the *contents*. ### The strongest hypothesis is about noise Separation only shows the signal is *different*. Here is why it might be useful; E2, not this argument, decides whether it is better. **A raw surprise signal responds strongly to unpredictable noise as well as to rare useful events.** Whether that response becomes a memory write depends on the optimizer, gate, regularizer and data regime; “large loss” is not itself a write. Robust losses are a principled way to limit outlier influence, but they introduce a real discrimination trade-off: some rare informative events can look like noise. The proposal is that subspace change may add information for that decision, not that surprise-based systems inevitably memorize noise. The relevant perturbation result for the singular subspace used here is **Wedin's sin Θ theorem** (or Davis–Kahan after passing to a symmetric dilation). Under the required separation assumptions it gives a bound of the form ``` ‖sin Θ‖ ≲ ‖E‖ / gap, gap = σ_r − σ_{r+1} ``` The constant and norm depend on the precise theorem statement, but the mechanism and failure condition are the same: subspace rotation is controlled by perturbation size **relative to** spectral separation. With a wide gap, an isotropic perturbation can change content magnitude while producing little subspace rotation; with a narrow gap, a small structured perturbation can swap a retained direction. Turning that separation into better writes is still an empirical hypothesis, not part of the theorem. **E1 is consistent with gap control in this synthetic noise family** (§7): μΔ correlates with ‖E‖/gap at Spearman **+0.87**, compared with +0.80 for ‖E‖ alone and −0.44 for gap alone. Bucketed by that ratio, median μΔ runs 0.0015 → 0.0049 → 0.036 → 0.36 → 0.96 rad as ‖E‖/gap crosses 0.01 → 0.1 → 1 → 10 — roughly 640× suppression in the small-ratio regime. This is a controlled numerical pattern, not confirmation that learned memories enjoy the same regime. **The tension in that argument — and E1 made it worse than I had it.** Noise immunity needs the gap *large*; case (B)'s sensitivity is constructed at near-degeneracy, where the gap is *small*. I had described these as opposite regimes. Measured, they are **disjoint**: in E1's sweep, every near-degenerate system had ‖E‖/gap > 0.1 for *every* noise amplitude injected, so the noise-immune bucket at near-degeneracy is **empty — n = 0**. You cannot buy case-(B) sensitivity and noise immunity in the same spectrum; the gap that makes one work destroys the other. So the gap is doing work nobody budgeted for, and it is the same debt as the *r* hyperparameter in §8 seen from the other side — one quantity is being asked both to supply noise immunity and to set the retained rank. What the noise argument licenses is strictly conditional: immunity *given* a wide gap. Whether learned memories maintain one is empirical, and the available evidence is not encouraging — linear-attention states are reported to end up low-rank with utilisation well below 1 (arXiv:2602.04852), which is the near-degenerate side of this line. E2 has to settle it; we put P = 0.40 on a real-task advantage (§9). --- ## 3 · A hypothesis, not a corollary: thresholding may produce hysteresis Suppose the gate g is a threshold — do not update when μΔ < τ, update when μΔ ≥ τ. Then the update can be sticky: - content can change however it likes *within the retained subspace* without triggering a boundary update, so the state **sticks**; - only when perturbation accumulates enough to swap a principal direction does it **jump** — a threshold-crossing, lagged event; - after the jump, the new subspace may stick in turn. But thresholding alone does **not** imply bistability or a hysteresis loop. A one-threshold update can be path-dependent without possessing two stable branches, distinct forward/reverse switching thresholds, or nonzero loop area after the input is returned. Those require the coupled state-update dynamics, not merely a discontinuous gate. The surviving claim is therefore experimental: projection-gated updating is a plausible mechanism for a latch **if** the full recurrence develops multistability and remanence. E3 must compare it against matched one-threshold and ordinary recurrent controls. If forward and reverse sweeps collapse after matching carried state and external memory, the hysteresis claim dies. No amount of stickiness alone rescues it. --- ## 4 · The road this hypothesis would let us skip Worth stating what §3 is competing against, since if §3 fails this is where the hysteresis question goes back to. **Occupied.** Ramsauer et al. (2020) proved attention equals **one** update step of a modern Hopfield network — the attractor structure is already inside the architecture; what the architecture does is refuse to iterate it and refuse to carry the state across the token boundary. Both refusals buy parallelism, and both are exactly what remove the inertia. The bistable-recurrence line is genuinely occupied: De Geeter et al. (ULiège, 2026) derive parallelizable memory recurrent units **directly from a hysteresis bifurcation**, with a bistable quantized state persistent across timesteps; CMRU (arXiv:2605.11855, ICML 2026) is the repaired version. The deep-equilibrium and energy side — DEQ (Bai et al. 2019), Energy Transformer, Hyper-SET (arXiv:2502.11646), latent recurrent-depth (arXiv:2502.05171) — iterates to convergence but **within a single token**, discarding the attractor state at the token boundary. **What is actually left**, stated as a claim about roads rather than about novelty: all existing evidence that hysteretic state helps sits in 100 nW analog circuits (Schmitt triggers), keyword spotting, and long-range synthetic tasks — and the ICML 2026 follow-up concedes BMRU underperforms parallelizable RNNs on complex sequence tasks. Nobody has walked `attention ≡ Hopfield` → *iterate beyond one step* → *carry the fixed point across tokens*, and nobody has shown hysteresis is useful at language scale. Two things are missing, a path and a demonstration, and the second is the harder one. If the corrected §3 hypothesis holds, this road may be unnecessary. If E3 fails, it is the fallback, and it is expensive. --- ## 5 · Provenance in the architecture — nearly closed, and I was wrong about it Starting with the concession, because the concession is load-bearing. I had this filed as open: the architecture cannot distinguish tokens it generated from tokens the world inserted. **It is not open at the granularity I stated it.** - **ASIDE** (Zverev et al., arXiv:2503.10566, ICLR 2026) applies an orthogonal rotation to data-token embeddings, separating executable from non-executable from the first layer, with the role assigned by the harness and unwritable by the content. That is precisely "an unforgeable provenance tag that generation cannot write into" — the thing I thought was missing. - **ISE** (arXiv:2410.09102, NeurIPS 2024) puts four segment embeddings into attention. - And arXiv:2606.27567 **proves** perfect injection resistance impossible in a shared-embedding architecture. What remains is narrow, and deserves to be stated narrowly: 1. **Every occupant cuts at instruction-versus-data. None cuts at self-versus-world** — "a token I emitted" versus "a token the environment handed me". Neither cut derives from the other: a tool result is data-shaped and world-origin; the model's own chain-of-thought is instruction-shaped and self-origin; a user instruction is instruction-shaped and world-origin. In a multi-turn agent loop with tool calls, self/world is the cut that tracks *which process is accountable for a token*, and instruction/data does not recover it. 2. **Provenance as a separate attention channel or permission-partitioned KV**, rather than an additive or rotational perturbation of the content embedding. 3. **A hard gate rather than a learned one** — the only route out from under the impossibility proof, which concerns what a shared-embedding model can *learn*. **Timing, since it is decision-relevant.** This gap is measured in months, not years, and the likeliest event that closes it is the ASIDE authors extending their own method along cut (1). Price that in before the idea, not after. --- ## 6 · One demoted, one deleted The highest information density per line in this note, because this is where the audit changed the answer. **Demoted to a footnote — state capacity as a decision variable.** The candidate: no architecture treats *how much* state to keep as something the model decides; adaptivity lives in read granularity, not capacity. Mostly occupied — H-Net's ratio loss already puts a content-adaptive compression rate in the objective, STAR-KV makes rank differentiable, and there is a rate–distortion survey of the area (arXiv:2607.08032). What is left is "remove the target-compression-rate hyperparameter", an ablation rather than a direction. (It returns as a real debt in §8: μΔ's *r* is exactly that hyperparameter wearing a different hat.) **Deleted — a Koopman eigenbasis of the model's own dynamics.** Fully occupied. **MamKO** (Li, Han & Yin, ICLR 2025) uses Mamba's selectivity to generate a content-adaptive Koopman operator online, with a stated motivation almost word for word the same ("a fixed linear operator is not expressive enough"), and arXiv:2606.09432 states outright that selective SSMs induce an input-conditioned Koopman operator. Deleted rather than downgraded, so it is not rediscovered. --- ## 7 · Experiments — one run, two pending ### E1 · Separation — **run; the wedge survives its own kill condition** Pre-registered kill condition: |correlation| > 0.8 between ‖ΔM‖ and μΔ would mean μΔ is ‖ΔM‖ in disguise, and the wedge dies. Result over 6,480 updates at d = 128, r = 16: | quantity | measured | kill threshold | |---|---|---| | Pearson(log‖ΔM‖, μΔ) | **−0.083** | \|r\| > 0.8 | | Spearman(‖ΔM‖, μΔ) | **−0.347** | — | | μΔ spread within one ‖ΔM‖ band | **3.07 rad = 100% of observed range** | — | | A-strict: ‖ΔM‖ = 3.2×10⁴ | μΔ = 6×10⁻⁸ rad | — | | B: ‖ΔM‖ = 1.0×10⁻² | μΔ = 5.7×10⁻³ rad | — | ![E1: left, ‖ΔM‖ against μΔ for four update families, coloured by relative spectral gap — the vertical stacks show μΔ spanning its full range at fixed ‖ΔM‖. Right, μΔ under isotropic noise against the Wedin-style ratio ‖E‖/gap, with the plotted √r·‖E‖/gap reference ceiling.](https://machengshen.github.io/theory/assets/e1-separation.png) The pooled correlation is the *weaker* of these numbers, because I choose the ensemble and could tune the correlation by reweighting families. The load-bearing number is ensemble-independent: **within a single band of ‖ΔM‖, μΔ takes 100% of its observed range** — so no monotone function of ‖ΔM‖ can reproduce μΔ. The right-hand panel is consistent with Wedin-style gap control for the noise family (Spearman +0.87 on ‖E‖/gap, against +0.80 on ‖E‖ alone); it does not by itself prove the theorem is tight or causal in learned memories. Two things running it changed, both against me: the family-(A) construction as originally stated was **wrong** (§2), and the spectral-gap tension turned out to be **disjointness rather than opposition** (§2) — the noise-immune bucket at near-degeneracy is empty. Script and figure: [`e1_separation.py`](https://machengshen.github.io/theory/assets/e1_separation.py). *Caveat that E1 does not touch:* this is linear memory. Nothing here says the induced projection is well defined for an MLP memory — see §8. ### E2 · Noise robustness — not yet run **E2 · Noise robustness (one to two days; the most persuasive).** A long-context recall task with high-amplitude random distractors injected. Compare surprise-gated (Titans-style), Huber-robust (Miras), and μΔ-gated. *Expected:* μΔ-gating matches or beats the robust version **while using no robust loss at all**, and stays flat as distractor amplitude grows where the other two degrade. **Kill:** if μΔ needs a robust loss added back to work, it offers no structural advantage and reduces to an increment. This is also the experiment that decides the spectral-gap tension in §2 — **track σ_r − σ_{r+1} throughout the run**, since E1 showed the immune and (B)-sensitive regimes are disjoint, so the measured gap of a *learned* memory is what decides whether the noise argument applies at all. ### E3 · The hysteresis curve — not yet run **E3 · The hysteresis curve (binary verdict on §3).** A state-tracking task with the input swept forward and then back. Record the proposed memory state, text/KV/external-memory carriers, and the control input separately. Compare projection-gated recurrence against (i) an ordinary recurrent learner, (ii) a matched one-threshold gate without multistability, (iii) a matched fixed-coupling bistable recurrent null with the adaptive/projection gate clamped, and (iv) replay with the carried state reset. Sweep rate and dwell time, restart from both branches, intervene directly on the gate, and transplant matched internal states between conditions. A genuine result requires two stable branches, distinct switching points or nonzero loop area under the same protocol, persistence after removing the drive, and an incremental effect of the adaptive gate over the fixed-coupling null. **Kill:** the loop vanishes after carrier matching/reset, reduces to simple lag, or survives unchanged when the proposed gate is clamped/intervened away; then §3 dies and the hysteresis question reverts to §4. --- ## 8 · Honesty list Not a disclaimer section. There is no reviewer here; we are the reviewer, so these are the four places *we* judge most likely to be where this fails, with our own numbers attached in §9. - **Grassmann gating, prior art — partially resolved, and not in my favour on the metric.** The *metric* is entirely off the shelf: subspace change-point detection and principal-angle detectors are mature signal-processing tools, so there is no originality credit in "use principal angles to detect subspace change". The question we have to answer for ourselves is why this is not simply subspace CUSUM applied to fast weights. Our answer: it is not a new detector, it is a claim about *where the detector belongs* — in the write path of a test-time memory rather than in a monitoring path. We rate that composition at P = 0.55 (§9), which is not high. The adjacent ML work points the other way rather than at this: **GPM** (arXiv:2103.09762), orthogonal gradient descent and selective gradient projection (arXiv:2603.26671) use subspaces to *protect* capacity by projecting away conflicting gradients — subspace as constraint, not as signal; **SubTrack-Grad** (arXiv:2502.01586) tracks a Grassmannian subspace for optimizer memory efficiency, not as a write gate; and Neural Subspace Reallocation (arXiv:2606.30067), checked directly, gates on embedding similarity at task arrival, not on subspace geometry. I found no one using subspace *change* as the write gate for a test-time memory. That is a composition claim, not a metric claim, and it is the weakest kind of novelty. One agent, one retrieval pass over a very large literature is weak evidence of absence — hence P = 0.55 rather than anything higher. - **Nonlinear memory is the real technical debt, and it is not small.** Π_M is clean when M is a matrix. Titans-style MLP memories need either the spectral subspace of a Jacobian (local, and then "the boundary" is only locally defined) or approximation through a probe distribution (and then the probe distribution is a new hyperparameter smuggled in). Neither is worked out. This is the part most likely to be where the idea actually fails. - **Where does *r* come from?** If *r* is a hyperparameter, it revives exactly the objection that killed the capacity candidate in §6 — the compression rate is still handed over by a human. The self-consistent repair is to let *r* adapt to the spectral gap, which is unverified, and which collides with the gap tension in §2: the same quantity is being asked to supply noise immunity and to set the retained rank. - **μΔ uses direction only and discards scale within the subspace.** There is plausibly a class of updates where scale matters and direction does not, and this signal is blind to all of them. Recorded as a known blind spot rather than defended. --- ## 9 · Calibration There is no reviewer here. We are the reviewer, so the obligation is to put our own numbers on our own claims rather than to argue a case. Two separate scores, because merging them is self-deception: **P(mech)** = the mechanism is real; **P(useful)** = it produces an observable advantage at realistic scale. Every number carries the observation that would move it — a confidence without an update condition is decoration. | Claim | P(mech) | P(useful) | What moves it | |---|---|---|---| | μΔ is not a monotone function of ‖ΔM‖ (separation) | **0.97** | n/a | **Settled by E1** (−0.08 pooled; 100% within-band spread). Was 0.90 before running; raised because the measured spread is the ensemble-independent form of the claim. Falls to 0.1 only if the A-strict construction is shown degenerate. | | …therefore μΔ lies outside the Miras design space | **0.40** | n/a | Our weakest joint. Miras's `R` is a free slot; `R := d_Gr` makes this "Miras with a particular R", so the defence is about which *role* the term plays, not mathematics. Rises to 0.8 given a proof that no Bregman/f-divergence `R` induces (A)/(B) transversality. **Lowered from 0.45** after writing §2 out in full: the role-based defence is weaker on the page than it was in my head. | | Gap-controlled singular-subspace stability (Wedin) | **0.90** | **0.30** | E1 is consistent with the mechanism (Spearman +0.87 on ‖E‖/gap; ~640× suppression at small ratio). **P(useful) remains low**: the immune and (B)-sensitive regimes are disjoint in this sweep, and learned states may be near-degenerate. E2 settles practical value. | | Projection-gated recurrence develops genuine hysteresis | **0.35** | **0.20** | Corrected from a false implication: thresholding alone does not entail bistability. Rises only if E3 shows two branches, remanence and a loop that survives carrier-matched controls; no loop → P(mech) to 0.05. | | Nobody uses subspace change as a write gate | **0.55** | n/a | One agent, one retrieval pass over a huge literature is weak evidence of absence. A second independent pass finding nothing → 0.7; a hit → 0. | | Induced projection extends to nonlinear memory | **0.35** | **0.25** | **Added — it was missing, and it is the most likely place this dies.** Π_M is clean only for matrix M; Titans-style MLP memories need a Jacobian spectral subspace (locally defined only) or a probe distribution (a new hyperparameter). A working definition that survives a Titans-scale run → 0.7. | | Whole line survives six weeks of our own honest testing | **0.30** | — | Unchanged. The killer is not originality; it is row 2 (framing dispute) plus row 3's P(useful) collapsing when real spectra turn out near-degenerate. | **Where we disagree with the baseline we were handed.** Three changes, each with a reason rather than a vibe. Row 1 up (0.90 → 0.97): E1 has been run and the within-band spread is stronger evidence than the pooled correlation. Row 2 down (0.45 → 0.40): writing the concession out made the role-based defence look thinner, not thicker. Row 3's P(useful) down (0.40 → 0.30): the disjointness result is worse than the tension as originally described. And one row added that the baseline omitted — nonlinear memory, at P(mech) 0.35, which on our own numbers is the single most likely cause of death. **The most likely thing to kill this line, concretely:** not being scooped, and not E3. It is that Π_M has no honest definition for an MLP memory, so the whole construction only ever applies to linear fast-weight memories — a corner of the design space that Titans and Miras have already left. The check is cheap and should come before E2: take a two-layer MLP memory, define Π via the Jacobian spectral subspace at a probe batch, and measure whether μΔ is stable under resampling the probe batch. If it is not, this is a linear-algebra result about a shrinking corner, and should be labelled as one. ## 10 · What would make me drop the diagnosis itself **It could be factually wrong.** J-space already makes the broad "no state" wording false. The narrower diagnosis would also fail if vanilla inference exhibited a persistent internal variable whose value depends on the route, survives removal of the drive, and changes future behavior under carrier-matched inputs. Determinism does not decide this; controlled state resets and forward/reverse sweeps do. **It could be true and not load-bearing, which is the worse failure.** If memory capacity growing with sequence length delivers recall, and recall is what the tasks want, the second-order question is real but inert and this note is philosophy with an architecture diagram attached. E2 is the one that answers this, because it asks whether the second-order signal buys anything a first-order one cannot. --- ## 11 · Provenance of this note Cognitive state: **speculative** throughout, with one exception: **E1 has been run**, and its numbers in §7 are measured rather than expected. E2 and E3 have not been run. §9 states a calibrated confidence for every substantive claim together with the observation that would move it. The prior-art positions in §2, §4, §5 and §6 come from a dedicated audit against five original candidates, and the audit **reversed my ranking**: what I had second is now the mainline, what I had fourth is now third and conceded, one candidate became a footnote and one was deleted. §3 was not planned; it appeared while working out §2 and is deliberately stated as a narrow conditional rather than a synthesis, because the failure mode of this line has historically been to unify things that only rhyme. Running E1 changed three things in this note, all against the author: the family-(A) construction was stated wrongly and is corrected in §2; the spectral-gap tension turned out to be disjointness rather than opposition, which lowers §9's P(useful) for the noise argument; and the step from "not a monotone function of ‖ΔM‖" to "outside the Miras design space" was found not to go through at all, and is now stated as an open framing dispute in the body rather than defended. This line has been scooped three times: on the cognitive light cone, on a boundary bifurcation published by Tononi & Koch in 2015, and once by an earlier note of my own that turned out to be re-deriving Friston. After that record, being right about what is already occupied is worth more than being first. There is no reviewer to satisfy here and no venue being targeted; the only question is whether the thing is true, which is why §9 states calibrated confidences and what would move them rather than arguing a case. Corrections — above all "this is occupied, here is the citation" — move those numbers fastest. Contact: macshen93@gmail.com ## Literature pointers - Behrouz, Li, Deng, Zhong, Razaviyayn & Mirrokni (2026), *Memory Caching: RNNs with Growing Memory*, arXiv:2602.24281 - Behrouz et al. (2025), *Titans: Learning to Memorize at Test Time*, NeurIPS 2025; *Miras*, arXiv:2504.13173 - Schmidhuber (1992); Schlag, Irie & Schmidhuber (2021), linear Transformers as fast-weight programmers; Sun et al. (2024), test-time training - Ramsauer et al. (2020), *Hopfield Networks is All You Need*, arXiv:2008.02217 - Bai, Kolter & Koltun (2019), *Deep Equilibrium Models*; Hyper-SET, arXiv:2502.11646; latent recurrent-depth, arXiv:2502.05171 - De Geeter et al. (ULiège, 2026), bistable parallelizable memory recurrent units; CMRU, arXiv:2605.11855 (ICML 2026) - Saha, Garg & Roy (2021), *Gradient Projection Memory for Continual Learning*, arXiv:2103.09762; selective gradient projection, arXiv:2603.26671; SubTrack-Grad, arXiv:2502.01586; Neural Subspace Reallocation, arXiv:2606.30067 - Zverev et al. (2026), *ASIDE: Architectural Separation of Instructions and Data*, arXiv:2503.10566 (ICLR 2026); ISE, arXiv:2410.09102 (NeurIPS 2024); inseparability result, arXiv:2606.27567 - Li, Han & Yin (2025), *MamKO: Mamba-based Koopman Operator*, ICLR 2025; input-conditioned Koopman in selective SSMs, arXiv:2606.09432 - Rate–distortion view of KV compression, arXiv:2607.08032 - Wedin (1972), *Perturbation bounds in connection with singular value decomposition* — singular-subspace sin Θ bounds; Davis–Kahan is the symmetric-eigenspace analogue and applies here only after a symmetric dilation - Low-rank structure and utilisation of learned linear-attention states, arXiv:2602.04852 - Zwanzig (2001), *Nonequilibrium Statistical Mechanics* — the memory-kernel price of projection - Anthropic (2026), [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/index.html) — J-lens/J-space; a sparse-frame workspace candidate across depth, not evidence of persistent autonomous closure or phenomenal consciousness - Related notes on this site: [State is a closure condition, not a given set](https://machengshen.github.io/theory/state-as-closure.md) · [Discounted credit is a cokernel problem, not a loop holonomy](https://machengshen.github.io/theory/discounted-credit-is-a-cokernel.md) · [Learning where the self ends](https://machengshen.github.io/theory/learning-the-self-boundary.md) --- ## Source: /theory/odd-relational-core.md (sha256:b8db7b4bcf03, language: en) # The odd relational core: anti-bipartiteness as the minimal structure — why civilizational compression schemes converge on *small, odd, centered, ever-turning* *Notes from one discussion · 2026-07-13* ## The honest header (read this before anything else) The one thing in this note that is **solid** is a graph-theory fact: a graph is two-colorable if and only if it contains no odd cycle. Everything built on top of that fact — four "pressures" that supposedly select for small-odd-centered-turning relational cores, and a reading of what that implies for moral cosmology — is **speculative application**. The most valuable part of the note is not the framework; it is the adversarial self-audit in §6, which concludes that most of the framework is scaffolding to be discarded and that the "odd" in the title is itself a near-misnomer. Cognitive state for the whole note: **speculative**; the graph-theory kernel inside it: not in dispute. Assertions are tagged by honesty level: **[theorem]** (theorem-grade) / **[framework]** (a mature framework) / **[inference]** (our own synthesis) / **[rhyme]** (a structural analogy; no claim that truth transfers). --- ## 0. Where this is most likely wrong (up front, not buried) In one line: **of the four pressures below, only one (pressure two) has a genuinely hard mathematical kernel; the other three read more like post-hoc rationalization, and the jump from "odd relational core" to any moral reading hides a substitution.** See §6 for the adversarial audit. Read each section below carrying that suspicion — and in particular, do not be seduced by the elegance of "all four pressures point at the same place." That very seamlessness is the warning sign, not the confirmation. --- ## 1. The four-pressure framework [framework] Restated: when you compress a *living relational world* (everything mutually depends on and acts on everything else), why do the schemes repeatedly land on a **small + odd + centered + ever-turning** relational core — triadic deities, five phases, seven chakras, three guṇas — rather than a "four symmetric camps" table or a "good-vs-evil binary axis"? **First, clear away a false explanation:** it is *not* "sacred primes." Three, five, seven are prime, but (a) people have a strong selection bias toward *small* numbers (working memory cannot hold large ones), and (b) among small naturals, primes are simply dense (2, 3, 5, 7 make up half of 2–8). So "sacred schemes are mostly prime" is a composite artifact of **selection bias × base rate**, requiring primeness to have no special power at all. The property actually doing the work is **oddness**, not primeness: - A 3×3 magic square = 9 = odd but not prime, and it works fine (has a center, resists bipartition). - Yin–yang = 2 = prime but not odd, is pure binary, and is exactly the thing these pressures *push away from*. → So the four pressures below are all about "why **odd**"; primeness is just a free rider. Any urge to give "3/5/7 predictive power *because* they are prime" is a metaphysical drift to be resisted — a number can carry discrete-Fourier bookkeeping without carrying metaphysical predictive force. Four **mutually independent** pressures (crucial: they are four legs, not four faces of one grand unification — see §6): **Pressure one · small = the compression objective itself [framework].** Compressing a relational world, the objective just *is* the shortest description (MDL). The hard constraint on the human side: working memory holds ~4±1 chunks (Cowan), and an oral-tradition scheme has to be memorable and transmissible. A scheme meant to live in human mouths for two thousand years gets its element count squeezed to single digits as a **channel-bandwidth constraint**, not a mysticism. → This pressure explains "why small," not "why odd." **Pressure two · odd = anti-binary [the theorem kernel; the body of this note].** Graph theory: **a graph can be two-colored (vertices split into two groups, no edge within a group) if and only if it contains no odd cycle.** An odd cycle is the smallest, purest obstruction to bipartition. A relational core that is an odd-cycle structure (each element begets one and checks another, with the loop never closing back into a clean two-way split) *structurally refuses* to be cleaved into "us vs. them." Even / bipartite structures, by contrast, natively invite the "two camps" reading → good-vs-evil dualism → an apocalyptic final battle. See §2. **Pressure three · centered = a seat for the "self" [framework + a hard hook].** An odd number of elements can arrange as **{one center + several symmetric pairs}** (5 = center + 2 pairs; 7 = center + 3 pairs; 9 = center + 4 pairs). Even numbers arrange into pure pairs with no center cell. A scheme meant to **encode the observer/self into itself** needs that center seat. This connects to embedded agency: a compressor must appear inside its own compression → it needs a self-consistent fixed point / self-node (see §4, and the published note on state as a closure condition). **Note: "odd → has a center" is hard (only odd counts have a unique center cell); "needs a self-seat → therefore pick odd" is a soft teleological inference.** Flagged. **Pressure four · ever-turning = no static equilibrium [framework + rhyme].** The (excitatory/inhibitory) dynamics on an odd cycle — A begets B, B begets C, C checks A, this kind of non-transitive relation — **has no static equilibrium, only cyclic rotation**. Rock-paper-scissors has no pure-strategy Nash equilibrium (any pure strategy is beaten by another); in ecology, May–Leonard non-transitive three-species competition produces persistent periodic rotation rather than a fixed point. → This writes "change never stops / no heat death / ceaseless generation" directly into the skeleton. Even / bipartite structures tend to settle into a static opposed fixed point (equilibrium = standoff). **The [framework] part here is mathematics (RPS / non-transitivity genuinely has no fixed point), but "civilizations chose it *in order to* encode perpetual motion" is a [rhyme].** The four pressures independently push the choice toward **small ∧ odd ∧ centered ∧ turning**. That they simultaneously hit cores like 3/5/7 is an **intersection of four constraints**, not four projections of one cause — a distinction that matters enormously (see §6). --- ## 2. Digging into pressure two: from the two-coloring theorem all the way to "cannot generate a final enemy" [body] ### 2a. The hard kernel: an odd cycle = the minimal obstruction to bipartition [theorem] Graph-theory fact (intuition, not proof): - A graph can be "two-colored, same color never adjacent" ⟺ the graph contains no odd-length cycle. - Even cycles are fine, still two-colorable. **Only odd cycles** break two-coloring. - Equivalently: **oddness is the necessary-and-sufficient source of "anti-bipartiteness."** To make a relational structure impossible to cleave cleanly in half, the cheapest move is to embed an odd cycle in it. A triadic core is the smallest odd cycle (a 3-cycle): beget → check → beget … never closes back into a "red team vs. blue team." Five phases are two 5-cycles of relations (a begetting 5-cycle + a checking 5-cycle, both odd). This is not metaphor: these cores *literally* are odd cycles, and odd cycles *literally* are the obstruction to two-coloring. ### 2b. How binary structure natively hatches an us-vs-them narrative [framework/inference] A binary / even-split scheme (good vs. evil, light vs. dark, us vs. them, orthodoxy vs. heresy) carries a "final axis of opposition" built in: - Each element is assigned to one of two groups → the world is partitioned into two identifiable, mobilizable camps. - Conflict has a **terminal shape**: one side wins, one side is annihilated (good shall defeat evil). This is an **absorbing state** — once reached, nothing moves (connecting to pressure four: binary = static fixed point). - Narratively, it hands you a "**final enemy**": a nameable, pointable object whose destruction "saves the world." Manichaean cosmology, apocalyptic final battles, an abstract wartime enemy-figure — all instances of this structure. **Structural judgment (inference):** an us-vs-them narrative is not "bad actors forcing it in"; it is the **natural read-out of a bipartite relational graph**. Compress the relational world into a two-color graph first, and the "final enemy" falls out almost for free. The topology of the scheme, prior to its content, determines whether it can hatch a clean enemy. ### 2c. How an odd relational core forces "everything mediates everything" [framework/inference] On an odd-cycle core there is no "final axis of opposition"; in its place: - Each element **both begets one and checks another** (in five phases: wood begets fire, wood checks earth; and is begotten by water, checked by metal). No element is purely "good" or purely "enemy" — each is simultaneously some thing's resource and some thing's check. - No element can be **isolated as the enemy** and annihilated: kill one and the cycle breaks, the whole loop collapses (extinguish fire → earth has nothing begetting it, water nothing to check → the system sickens). **Every element is necessary to the survival of the whole**, including the one that is checking you right now. - Conflict has no terminal absorbing state: checking is **local, cyclic, reversible** regulation, not **global, terminal, irreversible** annihilation. A checking B is not "A wants to destroy B"; it is "A suppresses B in this phase so C has room," and next phase the roles rotate. → Structural conclusion (inference): **a moral/political/cosmological system built on an odd relational core topologically cannot generate a "clean final enemy."** It can express tension, restraint, imbalance, the need for regulation — but it cannot express "annihilate X and the world is at peace." "Everything mediates everything" is not a virtue it chose; it is that **its graph structure leaves no seat for an us-vs-them terminus.** Same root as pressure four: no absorbing state ⟺ no terminus ⟺ no "annihilation = redemption." **Boundary (honest):** this is a structural constraint at the **scheme / representation level**, not the empirical claim that "civilizations using five phases don't wage war" (see §6 counterexamples). People can bolt an us-vs-them layer onto any scheme from the outside; the only claim is that **the scheme itself does not volunteer an us-vs-them terminal shape, while a binary scheme does.** --- ## 3. A falsifiable criterion + one honest small comparison tally ### 3a. Criterion [inference, falsifiable] Main prediction: **"change / cycle / process" cosmologies lean odd (odd relational cores); "static classification / symmetric order" schemes can accommodate even (binary / four-fold).** - If a tradition's core scheme is **dynamically generative** (speaks of arising and ceasing, rotation, transformation, flux), predict its core relational core is odd, with no terminus. - If a tradition's core scheme is **static-classificatory** (element taxonomy, directional layout, symmetric order), predict it tolerates even numbers, binary, four-fold, and more readily attaches a good-vs-evil terminal axis. - **Falsification point:** if one finds many "process cosmology + clean binary terminal core," or "static classification + strongly anti-binary odd core," the prediction is weakened. This is a correlational claim about **scheme type ↔ parity**, in principle checkable and refutable. ### 3b. One illustrative comparison tally (**not a systematic ethnography, only a common-knowledge comparison**) **Discipline enforced: use only widely-recognized, common-knowledge schemes; mark anything without a confirmable source as [?]; fabricate no citations, no ethnography.** Each row uses only well-known names, with no invented sources. Confidence annotated. Classificatory / static (predicted: tolerates even / binary): | Scheme | N | Parity | Type | Binary-terminus tendency? | conf | |---|---|---|---|---|---| | Greek four elements (earth/water/air/fire) | 4 | even | classification | often paired with hot–cold/wet–dry axes | high | | Four humors (blood/phlegm/black+yellow bile) | 4 | even | classification/balance | balance vs. imbalance, weak binary | high | | Four directions / four seasons | 4 | even | classification | no strong good-evil terminus | high | | Yin–yang | 2 | even (pure binary) | polarity | structurally bipartite (but the Chinese reading stresses **mutual containment/transformation**, weakening us-vs-them — see note below) | high | | Two cosmic principles, good vs. evil | 2 | even | polarity | **strong good-evil terminus** (the Manichaean template) | high | Process / cyclic (predicted: leans odd / anti-binary): | Scheme | N | Parity | Type | Binary-terminus tendency? | conf | |---|---|---|---|---|---| | Five phases (generation/overcoming) | 5 | odd | process/cyclic | no final enemy, cyclic regulation | high | | Trimūrti (Brahmā/Viṣṇu/Śiva) | 3 | odd | create-sustain-destroy process | no terminus, functional rotation | high | | Three guṇas (sattva/rajas/tamas) | 3 | odd | process/dynamic balance | no good-evil terminus, shifting proportion | med | | Daoist three treasures / three purities | 3 | odd | (mainly classificatory) | weak terminus | med [? names are polysemous] | | Six realms of rebirth | 6 | **even** | process/cyclic | no good-evil terminus, cyclic, no absorbing state | med (**counterexample below**) | Ambiguous / counterexamples (must be listed honestly): | Scheme | N | Parity | Why ambiguous | conf | |---|---|---|---|---| | Six realms of rebirth | 6 | even | **process-type yet even** → weakens "process ⟹ odd"; but has no binary terminus (six realms are not two camps), so it supports "anti-binary" but not "odd" | med | | Eight trigrams | 8 | even | process/generative (change), yet 8 = 2³, pure binary recursion | high | | Sixty-four hexagrams | 64 | even | same, binary recursion, extremely process-oriented yet extremely binary | high | | Ten sefirot (Kabbalah) | 10 | even | generative/emanative, even, yet contains a middle-pillar (center) structure | med | | Seven chakras | 7 | odd | supports (odd + centered + axial process) | med | **Tally (honest count, illustrative):** - Clearly supporting (process ⟹ odd, or classification ⟹ tolerates even): five phases, Trimūrti, three guṇas, seven chakras (odd/process side); four elements, four humors, four directions, the two-principle good-vs-evil template (even/classification side). **About 8 rows aligned.** - Clear counterexamples (process yet even): **eight trigrams, sixty-four hexagrams, six realms, ten sefirot — four hard counterexamples**, all "process/generative yet even." The trigrams / hexagrams are especially lethal: an ultra-process philosophy (change itself) that chose pure binary recursion. - Ambiguous: three purities [?], yin–yang (pure binary but the Chinese reading emphasizes mutual containment, itself blurring "binary = us-vs-them"). **The honest conclusion from this tally:** the "classificatory tolerates even" half stands fairly firm (four elements / four humors / four directions are uniformly even). The "process leans odd" half **has clear counterexamples** (hexagrams / trigrams / six realms). So the §3a criterion is **partly undermined by its own small survey** — in particular "process ⟹ odd" fails; at most "process ⟹ anti-binary (no good-evil terminus)," and anti-binariness can be achieved by odd (five phases) *or* by "pluralistic non-binary even" (six realms = six, hexagrams = sixty-four, neither being two camps). **This is an important self-correction: what is genuinely robust is not "odd" but "anti-binary"; odd is only the cheapest way to achieve anti-binariness, not the only one.** See §6. (Again: the above is an illustrative comparison, confidence uneven, names at the common-knowledge level; it constitutes no systematic ethnography, and supports no inference of the form "therefore some civilization is more benign.") --- ## 4. Connecting to the surrounding theory line [framework/inference, hard vs. soft flagged] **The generator / seed motif:** These schemes = the **minimal generative core (seed) of a relational world**, not a list of facts (leaves). Three/five/seven are not "three or five or seven things"; they are **compressed generators that unroll a whole family of relational judgments** — given the five phases, you can generate the endless leaves of "this season, this organ, this emotion, this kind of imbalance." What a civilization retains is the **seed (generative rule)**, not the leaves (specific correspondence tables); pressure one (small) is precisely the expression of "the seed must be maximally compressed." **This interface is hard**: the schemes genuinely are generators, not lists (dependent origination has texture; state = a quotient, per the published note on state as a closure condition). **Pressure three ↔ the embedded-agency self-node:** The center seat of an odd core = the compressor's **fixed point / self-node inside its own compression**. An embedded agent must build a world model that contains itself → it needs a self-consistent fixed point μ = F(μ) → the "center cell" in the scheme is the seat reserved for that self. **The hard part:** only odd counts have a unique center cell (this is arithmetic). **The soft part (rhyme, flagged):** "civilizations chose odd numbers in order to seat the self-node" is a teleological inference, not a theorem; it is entirely possible that odd numbers were chosen by pressure two, and the center cell *happened* to double as a self-seat afterward (exaptation, not design). Do not harden this rhyme into "the scheme is evidence for embedded agency." (See the published note on state as a closure condition for the fixed-point / self-node structure.) **Cross-line restraint:** pressure two (anti-binary), pressure three (self-seat), and pressure four (perpetual motion) can all "rhyme" back into a larger framework of fixed points and information dynamics — **but that is precisely the reflex to be watched.** They are **four independent legs** that happen to land on the same core, not four projections of one grand-unified generator. Stitching them into "information dynamics explains all the sacred numbers" is over-theorizing dressed as insight. Keep them each falsifiable, each killable. --- ## 5. (Merged into §4, not repeated) See the embedded-agency paragraph in §4. --- ## 6. Adversarial self-audit — where this is most likely wrong Cutting itself with the razor "don't mistake elegance for truth / an elastic framework that fits everything has no content": **(1) The weakest leg = pressure four (perpetual motion), next pressure three (centered).** Pressure two has a real mathematical kernel (two-coloring ⟺ no odd cycle, a theorem, indisputable). But pressures three and four read the **arithmetic by-products** "odd counts happen to have a center cell" / "odd cycles happen to have no fixed point" *backwards* into "civilizations chose odd *in order to* encode self/perpetual-motion" — this is a **teleological inversion**: oddness came first (possibly from pressure one + two alone), the center cell and the no-fixed-point property were **free bonuses**, and then got retroactively credited as "the reason for choosing it." This is the most post-hoc part. Honestly: pressures three and four are more likely **exaptations** of pressure two (unintended usability of an old structure), not four independent pressures. **If so, "four pressures" is really "one and a half,"** and the persuasiveness of "four independent constraints converging" (§1's closing "intersection of four constraints") is badly discounted — it may be one leg (anti-binary) carrying two gifts. **(2) "Odd" itself may be a false protagonist — the §3b survey already bit back.** Hexagrams / trigrams / six realms are hard counterexamples: ultra-process type yet even, yet still anti-binary (not two camps). This shows the thing actually doing the work is **anti-bipartiteness**, and "odd" is only the way to achieve anti-binariness with the **fewest elements** (an odd cycle is the smallest anti-bipartite obstruction) — but **pluralistic even numbers** (6 realms, 8 trigrams, 64 hexagrams) can be anti-binary just as well. → So the whole "why odd numbers" framework should be **demoted** to "why anti-binary," with "odd" retreating to a corollary (when the element count is squeezed to a minimum *by pressure one*, the cheapest anti-binary realization happens to be an odd cycle). This is the note's most honest internal correction, and it should be written into the seed: **the protagonist is "anti-bipartiteness," not "odd."** **(3) The political jump (which §2c gestures at) hides a substitution.** Going from "an odd scheme topologically does not hatch us-vs-them" to any social conclusion quietly equates the **representation-level structure** with a **society-level power structure**. But a society's **relational graph** (who depends on whom, who checks whom) and the **cosmological scheme** it professes (how many gods, how many phases) are **two different graphs**, with no theorem guaranteeing they are isomorphic. An empire using five phases can still run us-vs-them mobilization; a society with a binary theology can be highly distributed. Any structural analogy of the form "distributed ⟺ anti-binary scheme" is a **pretty analogy**, not a validated social-science proposition — at most it yields a **researchable hypothesis** (is a polity's concentration of power correlated with the binariness of its dominant narrative?), never a conclusion. This jump most needs the §3-style falsification, or it is a political intuition borrowing a mathematical coat. **(4) Elasticity = the danger signal itself.** This framework can "explain" triadic deities, five phases, seven chakras, *and* why four elements are classificatory, *and* connect on the side to embedded agency — **this fit-anything smoothness is exactly the smell §0 warned of.** By subtraction: if you cut pressures three and four (leave them as exaptation footnotes), demote "odd" to a corollary of "anti-binary," and demote the moral-reading paragraph to a "hypothesis to be tested," what remains is **one hard kernel only**: > **Two-coloring ⟺ no odd cycle, so to make a relational core topologically refuse the "us vs. them" binary terminus, the cheapest structure is to embed an odd cycle; a binary scheme, by contrast, natively supplies a mobilizable final enemy.** That one line is true, falsifiable, and has content. Everything else is rhyme and bonus around it. **If this note kept only one sentence, keep that one; treat the rest as a raft, abandon it once across.** --- ## Links Published companions: [State is a closure condition, not a given set](https://machengshen.github.io/theory/state-as-closure.md) (the fixed-point / self-node structure behind pressure three; the generator/seed and "state = a quotient" motifs); [Holography ↔ Koopman](https://machengshen.github.io/theory/holography-koopman.md) (the "forgetting = compression, not deletion" and eigenbasis motifs the seed idea draws on). **Discipline:** the genuinely hard kernel is the last sentence of §6 (the graph-theory minimal obstruction to bipartition). The note's most likely errors are §6 (1)(2)(3): the four pressures may really be "one anti-binary leg + two exaptation bonuses," "odd" should be demoted to a corollary of "anti-binary," and the moral reading is a hypothesis to be tested, not a conclusion. --- ## Source: /theory/odd-relational-core.zh.md (sha256:4f4405ca34ba, language: zh-Hans) # 奇关系核:反二分作为最小结构 —— 为什么文明的压缩图式收敛到"小、奇、有心、不停转" *一次讨论的整理 · 2026-07-13* ## 诚实的题头(先读这段) 这篇 note 里唯一**硬**的东西是一个图论事实:一个图可二染色当且仅当它不含奇环。建在这个事实之上的一切——四股据称把选择推向"小-奇-有心-转"的关系核的压力,以及它对道德宇宙论的读法——都是**推测性应用**。这篇最有价值的部分不是那个框架,而是 §6 的对抗自审:它的结论是,框架的大部分是要弃掉的脚手架,而标题里的"奇"本身近乎一个用词不当。全文认知状态:**speculative**;其中的图论内核:无争议。 论断按诚实度标注:**[定理]**(有定理级)/ **[框架]**(成熟框架)/ **[推断]**(我们的综合)/ **[韵脚]**(结构类比,不主张真值传递)。 --- ## 0. 这套理论最可能错在哪(先放这里,别读完才知道有坑) 一句话:**四股压力里只有一股(压力二)有真正硬的数学内核,另外三股更像事后合理化,而"奇关系核 → 道德读法"这一跳里藏着一个偷换。** 详见 §6 对抗自审。读下面每一节时请带着这个怀疑——尤其别被"四股全指同一处"的优雅骗过去。那正是警告信号,不是确认。 --- ## 1. 四股压力的完整框架 [框架] 问题重述:压缩一个"活的关系世界"(万物相互依赖、相互作用)时,为什么图式反复落到 **小 + 奇 + 有中心 + 不停转** 的关系核?——三相神、五行、七脉轮、三 guṇa 这类,而不是"四大阵营对称表"或"二元善恶轴"。 **先清掉一个假解释:** 不是"神圣质数"。三、五、七是质数,但(a)人对"小数"有强选择偏差(工作记忆撑不住大数),而(b)小自然数里质数本来就密(2,3,5,7 占了 2–8 的一半),所以"神圣图式多是质数"是**选择偏差 × 基率**的合成假象,不需要质数有任何特殊力量。真正干活的属性是**奇**,不是**质**: - 九宫格 = 9 = 奇非质,照样成立(有中心、反二分)。 - 阴阳 = 2 = 质非奇,是纯二元,恰恰是被这套压力**推开**的东西。 → 所以下面四股讲的都是"为什么**奇**",质数只是搭便车。任何想让"3/5/7 因为是质数而有预测力"的冲动都是要克制的形而上漂移——一个数可以携带离散傅里叶式的编码,却不携带形而上的预测力。 四股**互相独立**的压力(关键:它们是四条腿,不是一个大统一的四个面——见 §6): **压力一 · 小 = 压缩目标本身 [框架]** 压缩一个关系世界,目标就是最短描述(MDL)。人脑侧的硬约束:工作记忆 ~4±1 个 chunk 顶天(Cowan),口传文化要能被记住、传下去。一个要活在人嘴里两千年的图式,元素数被压到个位数是**传输信道的带宽约束**,不是神秘学。→ 这一股解释"为什么小",不解释"为什么奇"。 **压力二 · 奇 = 反二元 [定理内核,本 note 主体]** 图论:**一个图可以被二染色(顶点分成两组、同组无边)当且仅当它不含奇环。** 奇环是"不可二分"的最小、最纯的障碍。一个关系核如果是奇环结构(每个元素既生一个、又克一个,首尾接不回一个干净的二分),它在结构上**拒绝**被劈成"我们 vs 他们"两个阵营。偶数/二分结构则天然邀请"两阵营"读法 → 善恶二元 → 末世式的最终决战。详见 §2。 **压力三 · 有中心 = 给"自我"留座位 [框架 + 接硬结构]** 奇数个元素能排成 **{一个中心 + 若干对称对}**(如 5 = 心 + 2 对;7 = 心 + 3 对;9 = 心 + 4 对)。偶数排成纯对、没有中心格。一个要**把观察者/自我编码进去**的图式需要那个中心座位。这接 embedded agency:压缩者必须出现在自己的压缩里 → 需要一个自洽不动点 / 自我节点(见 §4,以及已发布的"state 是一个闭包条件"那篇)。**注意:"奇→有中心"是硬的(奇数才有唯一中心格),"需要自我座位→所以选奇"是软的目的论推断。** 标清。 **压力四 · 永动 = 无静止均衡 [框架 + 韵脚]** 奇环上的(激励/抑制)动力学——A 生 B、B 生 C、C 克 A 这类非传递关系——**没有静止平衡点,只有循环轮转**。石头剪刀布没有纯策略纳什均衡(任何纯策略都被另一个克);生态学 May–Leonard 三物种非传递竞争产生持久的周期轮转而非定态。→ 把"变化不止 / 无热寂 / 生生不息"直接编进骨架。偶数/二分结构容易停在一个静态对立的定态(平衡=对峙)。**这一股[框架]成分是数学(RPS/非传递确实无定态),但"文明选它是为了编码永动"是[韵脚]。** 四股各自独立地把选择压向 **小 ∧ 奇 ∧ 有心 ∧ 转**。它们同时命中 3/5/7 这类核,是**四个约束的交集**,不是一个原因的四个投影——这个区分至关重要(见 §6)。 --- ## 2. 深挖压力二:从二染色定理一路推到"生不出最终敌人" [主体] ### 2a. 硬内核:奇环 = 不可二分的最小障碍 [定理] 图论事实(给直觉不给证明): - 一个图能被"两色染,同色不相邻" ⟺ 图里没有奇长度的环。 - 有偶环没关系,照样能二染。**唯独奇环**让二染崩掉。 - 换句话说:**奇性是"抗二分"的充要来源。** 想让一个关系结构无法被干净地劈成两半,最省的办法就是在里面埋一个奇环。 三元核就是最小奇环(3-环):生→克→生……接不回一个"红队 vs 蓝队"。五行是 5-环两套关系(生 5-环 + 克 5-环,都是奇环)。这不是比喻,是这些核**字面上**就是奇环,而奇环**字面上**就是二染色的障碍。 ### 2b. 二分结构如何天然孵化敌我叙事 [框架/推断] 一个二元/偶分图式(善vs恶、光vs暗、我们vs他们、正统vs异端)自带一条"最终对立轴": - 每个元素被指派到两组之一 → 世界被划成两个可辨识、可动员的阵营。 - 冲突有一个**终局形状**:一方胜、一方灭(善终将战胜恶)。这是一个**吸收态**——达到后不再动(接压力四:二分=静态定态)。 - 叙事上,它给出一个"**最终敌人**":一个可命名、可指向、消灭了就"世界得救"的对象。摩尼教的善恶宇宙观、末世决战、一个抽象的战时敌人形象——都是这个结构的实例。 **结构判断(推断):** 敌我叙事不是"坏人硬塞进来的",它是**二分关系图的自然读出**。你先把关系世界压成二色图,"最终敌人"就几乎免费地掉出来了。图式的拓扑先于内容决定了它能不能孵出一个干净的敌人。 ### 2c. 奇关系核如何强制"一切居间调停一切" [框架/推断] 奇环核上没有"最终对立轴",取而代之的是: - 每个元素**既生一个、又克一个**(五行:木生火、木克土;同时被水生、被金克)。没有哪个元素是纯粹的"善"或纯粹的"敌"——每个都同时是某物的资源和某物的克星。 - 没有元素能被**孤立成敌人**并消灭:灭掉一个,环就断了,整个循环塌掉(灭火 → 土无所生、水无所克 → 系统病)。**每个元素对整体的存续都是必要的**,包括那个此刻在克你的。 - 冲突没有终局吸收态:克是**局部、循环、可逆**的调节,不是**全局、终末、不可逆**的消灭。A 克 B 不是"A 要灭 B",是"A 在这一相位压制 B,好让 C 有空间",下一相位角色轮转。 → 结构性结论(推断):**建在奇关系核上的道德/政治/宇宙体系,拓扑上就生不出"干净的最终敌人"。** 它能表达张力、克制、失衡、需要调节,但表达不了"消灭 X 则天下太平"。"一切居间调停一切"不是它选择的美德,是它的**图结构没给敌我终局留位置**。这跟压力四同源:无吸收态 ⟺ 无终局 ⟺ 无"消灭即救赎"。 **边界(诚实):** 这说的是**图式/表征层**的结构约束,不是"用五行的文明就不打仗"的经验断言(见 §6 反例)。人可以在任何图式外面加一层敌我;声称的只是——**图式本身不主动提供敌我的终局形状,二元图式主动提供**。 --- ## 3. 可证伪判据 + 一次诚实小对照普查 ### 3a. 判据 [推断,可证伪] 主预测:**"变化/循环/过程"型宇宙论偏奇(奇关系核);"静态分类/对称秩序"型图式容得下偶(二分/四分)。** - 若一个传统的核心图式是**动态生成型**(讲生灭、轮转、气化、流变),预测其核心关系核是奇的、无终局的。 - 若一个传统的核心图式是**静态分类型**(讲元素分类、方位排布、对称秩序),预测它容许偶数、二分、四分,且更容易配一条善恶终局轴。 - **可证伪点:** 若发现大量"过程宇宙论 + 干净二元终局核",或"静态分类 + 强反二元奇核",则本预测被削弱。这是一个关于**图式类型 ↔ 奇偶** 的相关性断言,原则上可查、可反。 ### 3b. 一次 illustrative 对照 tally(**不是系统民族志普查,只是常识级对照**) **铁律执行:只用公认常识级图式;查不到确切出处的标 [?];不编引用、不编民族志。** 下面每条只用广为人知的名目,不附伪造出处。confidence 标注。 分类/静态型(预测:容偶/二分): | 图式 | 数 | 奇偶 | 类型 | 是否含二元终局倾向 | conf | |---|---|---|---|---|---| | 希腊四元素(土水气火) | 4 | 偶 | 分类 | 常配冷热干湿两轴对立 | 高 | | 四体液(血/黏液/黑黄胆) | 4 | 偶 | 分类/平衡 | 平衡vs失衡,弱二元 | 高 | | 四方位/四季 | 4 | 偶 | 分类 | 无强善恶终局 | 高 | | 阴阳 | 2 | 偶(纯二元) | 极性 | 结构上就是二分(但中式阴阳强调**互含转化**,弱化敌我——见下注) | 高 | | 善恶两大本原对立 | 2 | 偶 | 极性 | **强善恶终局**(摩尼教母体) | 高 | 过程/循环型(预测:偏奇/反二元): | 图式 | 数 | 奇偶 | 类型 | 是否含二元终局倾向 | conf | |---|---|---|---|---|---| | 五行(生克) | 5 | 奇 | 过程/循环 | 无最终敌人,循环调节 | 高 | | 三相神 Trimūrti(梵天/毗湿奴/湿婆) | 3 | 奇 | 生-持-灭过程 | 无终局,功能轮转 | 高 | | 三 guṇa(sattva/rajas/tamas) | 3 | 奇 | 过程/动态平衡 | 无善恶终局,配比流变 | 中 | | 道教三宝/三清 | 3 | 奇 | (分类为主) | 弱终局 | 中[?名目多义] | | 六道轮回 | 6 | **偶** | 过程/循环 | 无善恶终局,循环无吸收态 | 中(**反例见下**) | 暧昧/反例(必须诚实列出): | 图式 | 数 | 奇偶 | 为什么暧昧 | conf | |---|---|---|---|---| | 六道轮回 | 6 | 偶 | **过程型却是偶数** → 削弱"过程⟹奇";但它无二分终局(6 道不是两阵营),所以支持"反二元"但不支持"奇" | 中 | | 八卦 | 8 | 偶 | 过程/生成型(易=变),却是 8=2³ 纯二分递归 | 高 | | 易64卦 | 64 | 偶 | 同上,二进制递归,极过程化却极二分 | 高 | | 十 sefirot(卡巴拉) | 10 | 偶 | 生成/流溢型,偶数,但含中柱(中心)结构 | 中 | | 七脉轮 | 7 | 奇 | 支持(奇+有中心+沿轴过程) | 中 | **tally(诚实计数,illustrative):** - 明确支持(过程⟹奇 或 分类⟹容偶):五行、Trimūrti、三 guṇa、七脉轮(支持奇/过程侧);四元素、四体液、四方位、善恶两本原(支持偶/分类侧)。**约 8 条方向一致。** - 明确反例(过程却偶):**八卦、易64、六道轮回、十 sefirot ——4 条硬反例**,全是"过程/生成型却偶数"。八卦/易尤其致命:极致过程哲学(易=变)却选了纯二进制二分递归。 - 暧昧:三清[?]、阴阳(纯二元但中式强调互含,自身就模糊了"二分=敌我")。 **读这个 tally 的诚实结论:** "分类型容偶"这半边站得比较稳(四元素/四体液/四方位齐刷刷是偶)。"过程型偏奇"这半边**有明显反例**(易/八卦/六道)。所以 §3a 的判据**部分被自己的小普查削弱**——尤其"过程⟹奇"不成立,顶多"过程⟹反二元(无善恶终局)",而反二元既能靠奇(五行)也能靠"多元非二分的偶"(六道六个、易六十四个,都不是两阵营)。**这是一个重要的自我修正:真正稳的不是"奇",是"非二分";奇只是达成非二分的最省手段之一,不是唯一。** 见 §6。 (再次强调:以上是 illustrative 对照,confidence 参差,名目取常识级,不构成系统民族志;任何"因此某文明更良善"的推论都不被这张表支持。) --- ## 4. 接周边理论线 [框架/推断,标清硬软] **生成器/种子母题:** 这些图式 = **关系世界的最小生成核(种子)**,不是事实清单(叶子)。三/五/七不是"三个五个七个东西",是**能 unroll 出一整套关系判断的压缩生成元**——给定五行,你能生成"这个季节、这个器官、这种情绪、这类失衡"的无穷叶子。文明保留的是**种子(生成规则)**不是叶子(具体对应表);压力一(小)正是"种子要极致压缩"的表现。**这一接口是硬的**:图式确实是生成器不是清单(缘起有纹理、state=商,同已发布的"state 是一个闭包条件"那篇)。 **压力三 ↔ embedded-agency 自我节点:** 奇数核的中心座位 = 压缩者在自己的压缩里的**不动点/自我节点**。embedded agent 必须建一个包含自己的世界模型 → 需要 μ=F(μ) 的自洽不动点 → 图式里那个"中心格"就是给这个自我留的位置。**硬的部分:** 奇数才有唯一中心格(这是算术)。**软的部分(韵脚,标清):** "文明选奇数是为了安放自我节点"是目的论推断,不是定理;完全可能是奇数因压力二被选中后,中心格**恰好**能兼作自我座位(exaptation,不是设计)。别把这个韵脚硬成"图式=embedded-agency 的证据"。(不动点/自我节点结构见已发布的"state 是一个闭包条件"那篇。) **跨线关系(克制):** 压力二(反二元)、压力三(自我座位)、压力四(永动)确实都能"rhyme"回不动点与信息动力学那套大框架——**但这正是要警惕的反射**。它们是**四条独立的腿**恰好落在同一个核上,不是一个大统一生成器的四个投影。把它们缝成"信息动力学解释了所有神圣数字"= 把过度理论化伪装成洞见。保持它们各自可证伪、各自可 kill。 --- ## 5.(并入 §4,不重复) 见 §4 的 embedded-agency 段。 --- ## 6. 对抗自审——这套最可能错在哪 用"别把优雅误当真相 / 弹性框架能套一切 = 没内容"的刀砍自己: **(1) 最弱的一股 = 压力四(永动),其次压力三(有中心)。** 压力二有真数学内核(二染色⟺无奇环,这是定理,不可争)。但压力三、四把"奇数恰好有中心格""奇环恰好无定态"这些**算术副产品**,反读成"文明**为了**编码自我/永动而选奇"——这是**目的论倒装**:先有了奇(可能纯因压力一+二),中心格和无定态是**免费搭送**的性质,然后被追认为"选它的理由"。这是最像事后合理化的地方。诚实说:压力三、四更可能是压力二的**exaptation**(旧结构的意外可用),不是四股独立压力。**若真是这样,"四股"其实是"一股半"**,而"四股独立汇聚"的说服力(§1 结尾那句"四个约束的交集")就大打折扣——它可能是一股(反二元)带着两个赠品。 **(2) "奇"本身可能是假主角——§3b 的小普查已经反咬。** 易/八卦/六道是硬反例:极致过程型却偶数、却仍反二元(非两阵营)。这说明真正干活的是**非二分(anti-bipartite)**,"奇"只是达成非二分**最省元素数**的手段之一(奇环是最小非二分障碍),但**多元偶数**(6 道、8 卦、64 卦)一样能非二分。→ 那么整套"为什么是奇数"的框架,应该被**降级**成"为什么反二元",而"奇"退回成一个 corollary(在**元素数受压力一压到最小**时,反二元的最省实现恰是奇环)。这是本 note 内部最诚实的修正,应该写进种子:**主角是"抗二分",不是"奇"。** **(3) 那个道德跳(§2c 所指向)藏着一个偷换。** 从"奇图式拓扑上不孵敌我"到任何社会结论,中间悄悄把**表征层的结构**等同于**社会层的权力结构**。但一个社会的**关系图**(谁依赖谁、谁制约谁)和它信奉的**宇宙图式**(几神几行)是**两个图**,没有定理保证它们同构。用五行的帝国照样能搞敌我动员;用二元神学的社会也可能高度分布式。任何"分布式 ⟺ 反二元图式"式的结构类比是一个**漂亮的类比**,不是一个已验证的社会科学命题——它顶多给出一个**可研究的假设方向**(权力集中度 ↔ 主导叙事的二分度 相关吗?),绝不是结论。这一跳最需要 §3 那种可证伪化,否则就是政治直觉借了一件数学外套。 **(4) 弹性 = 危险信号本身。** 这套框架能"解释"三神、五行、七轮、也能"解释"为什么四元素是分类型、还能顺手接上 embedded agency——**这种什么都能接的顺滑,恰恰是 §0 警告的味道**。为道日损:如果砍掉压力三、压力四(留作 exaptation 注脚)、把"奇"降级成"反二元的 corollary"、把道德段降级成"待验证假设",剩下的**硬核只有一条**: > **二染色⟺无奇环,所以要让一个关系核拓扑上拒绝"我们vs他们"的二分终局,最省的结构是埋一个奇环;二分图式则天然提供可动员的最终敌人。** 这一条是真的、可证伪的、有内容的。其余全是它周围的韵脚和赠品。**如果这个 note 只留一句话,留这句;其余当筏子,过河即弃。** --- ## Links 已发布的姊妹篇:[State 是一个闭包条件,不是给定的集合](https://machengshen.github.io/theory/state-as-closure.zh.md)(压力三背后的不动点/自我节点结构;生成器/种子与"state=商"母题);[全息 ↔ Koopman](https://machengshen.github.io/theory/holography-koopman.zh.md)("遗忘=压缩、不是删除"与本征基母题,种子想法所借。) **纪律:** 真硬内核 = §6 结尾那一句(反二分的图论最小障碍)。全文最可能错在 §6(1)(2)(3):四股其实可能是"反二元一股 + 两个 exaptation 赠品",且"奇"应降级为"反二元"的 corollary,道德那跳是待验证假设不是结论。 --- ## Source: /theory/state-as-closure.md (sha256:38babbddb2ec, language: en) # What is a state? — from the MDP notion of state to a "closure" ontology *Notes from one discussion · Macheng × agent · 2026-07-09 · scope amended 2026-08-04* ## The question we started from We often say "information determines the reachability of futures". Put that inside the MDP framework of reinforcement learning / decision theory and you need a state — so what *is* that state? Is it spacetime, the big stage? Or is it the web of relations in which "everything depends on everything else"? And the set that "state" refers to seems to keep changing; there is nothing that stays fixed. Perhaps the structural level is invariant while the concrete form keeps being rewritten. Below are the six layers we dug out along this question. Assertions are tagged by honesty level: **[theorem]** (there is theorem-grade literature) / **[framework]** (a mature theoretical framework) / **[inference]** (our own synthesis) / **[rhyme]** (a structural analogy; no claim that truth transfers). ## 1. The MDP definition already gives the game away: a state is not a thing, it is a quotient space The textbook says a state must satisfy the Markov property: given s, the future is conditionally independent of the past. Notice the shape of that sentence — it is not describing some thing in the world, it is *imposing a condition*: whatever can "screen off" the past deserves to be called a state. In the setting of stationary stochastic processes and predictive equivalence, computational mechanics carries this step to completion **[theorem, scoped]**: take histories that induce the same conditional distribution over futures, and their equivalence classes are predictive causal states. The minimality result belongs to that formal setting; it is not a theorem that every scientific state representation is one unique quotient. Our synthesis **[inference]** is to treat a task-relative state as **a quotient of histories under an equivalence relation such as "makes no difference to the predictions or controls I care about"**. This is a modelling stance, not a universal ontology theorem. The quotient depends on the dynamics, observation interface and telos; changing them can require a different state representation. ## 2. The Mori–Zwanzig view: the state is where you decide to stop carrying memory Under the usual operator/evolution assumptions, the Mori–Zwanzig formalism gives an exact projected identity **[theorem, scoped]**: choose resolved observables and a projection, and their evolution can be decomposed into an instantaneous term, a memory term and an orthogonal-dynamics term. The slogan survives, but "any variables in any system" was too broad: **degrees of freedom removed by a specified projection generally return through memory and unresolved forcing rather than literally disappearing**. Read as a heuristic **[inference]**: a useful approximately Markov representation is often *purchased* by carrying enough variables, accepting a controlled memory kernel, or tolerating unresolved forcing. The practical question is not assumed to be "the one true state of the world", but: **under a stated capacity, purpose and error tolerance, which representation closes the dynamics well enough?** Choosing a projection also chooses what information is discarded; weak-memory regimes can sometimes justify a local approximation with explicit error bounds. ## 3. Stage or web of relations? Neither is primitive — but the web hides what makes a state possible - **The stage is not primitive**: modern results along the holography line (Ryu–Takayanagi's entanglement entropy = minimal surface area; Van Raamsdonk's "disentangle → spacetime tears apart"; the MERA tensor network) point to this: the connectivity of space is generated by entanglement structure, and the "stage" itself unfolds out of the structure of correlations **[framework; the rigorous mathematics is inside AdS]**. "Spacetime as a state" is the deepest accounting quotient physics has built so far — extremely successful, but still a quotient, not a floor. - **But if "everything depends on everything else" is said in full, then a finite state is simply impossible** — strict total dependence means any finite truncation leaks. Finite states are workable because the web of relations **has structure**: interactions are local, correlations decay with distance/time, and screening surfaces exist (in the language of graphical models, the Markov blanket) **[framework]**. In one line: relations come first, but relations have texture; **a state is the approximate closure that the texture permits**. In a universe with no screening structure there are no agents. ## 4. The rigorous version of "information determines reachability" In stochastic control this sentence has a precise body **[framework]**: the information an agent accumulates over time is a **filtration**, and any admissible policy must be adapted to it — you cannot act on what you do not know. If one filtration contains another and action/risk/resource constraints are held fixed, the coarse-information policy class embeds in the finer-information class. This does **not** mean every physical reachable set strictly grows; it means the optimum over admissible information-conditioned policies cannot worsen merely because more information is available and may be ignored. In a POMDP, an optimal controller can often be formulated on an **information state**, classically a belief over latent world states **[framework]**. This does not make world state irrelevant; it says the controller acts through information available to it. Empowerment (Klyubin–Polani) is the channel capacity from actions to future observations: a measure of potential controllable influence under a chosen horizon and channel model, not the reachable set itself. ## 5. Three mathematical homes for the intuition "the set keeps changing, the structure does not" 1. **Learning = re-taking the quotient**: when the model changes, the partition of "makes no difference" changes, and the state set gets re-divided accordingly. The mathematical shell of belief space stays put; the coordinate chart the agent actually uses keeps getting swapped. 2. **An atlas, not a single coordinate system**: in an open world any fixed state set is only a temporary chart; what persists are the **transformation rules** between charts. This position has a name in the philosophy of science — structural realism: what survives theory change is relational structure, not the inventory of objects. 3. **The renormalization group**: each scale has its own state set, and none of them is "the real one"; what is invariant is the **flow** connecting the levels, and its fixed points. Collapsed into one sentence **[inference]**: **what is invariant is the equation of the closure condition; the state set is merely that equation's solution under the current (world, interface, capacity, telos)**. Change the environment, the capacity, or the purpose, and the solution is recomputed — the eigenvalue equation does not move, the eigenvectors change with the operator. Incidentally, a numerical observation we made recently on small synthetic systems **[empirical, preliminary, unaudited]**: wire "the representation" and "the statistics you live out under that representation" into a self-consistent loop (representation fixes the projection → projection fixes the closed model → the closed model generates trajectories → trajectory statistics update the representation), and in the region where fitting capacity is sufficient, this loop appeared to have a unique fixed point and geometric convergence. The interesting open region is limited capacity: whether several different but individually self-consistent representations can coexist. **No public code or result artifact is linked here, so this paragraph is a working-note report, not independently auditable evidence.** ### 2026 evidence boundary: workspace is not yet closure Anthropic's Jacobian-lens experiments identify a J-space inside Transformers whose contents can be reported, modulated and flexibly routed across an intermediate layer band **[empirical, external]**. This corrects any blanket claim that a Transformer has "no state": it has activation state, computational state and a transient workspace-like representational state. But J-space is a sparsity-bounded union of cones from an overcomplete frame, not one fixed projector, and the reported evidence is across model depth rather than for an autonomous variable that persists across steps and maintains its own forgetting policy. It is therefore an **adjacent candidate coordinate for access**, not yet a solution to the closure equation above. A different 2026 result supplies a constraint rather than a bridge **[inference from external theorems]**. OpenAI reports a disproof of Connes's rigidity conjecture: within the theorem's setting, the associated von Neumann algebra need not uniquely determine the underlying group. It also reports the construction of non-sofic groups, so approximation by finite symmetric models cannot be assumed universally. Neither theorem is about agents or consciousness. The transferable warning is narrower: observable/operator-level equivalence need not identify a unique substrate, and any finite-state approximation program must state its approximability assumption rather than smuggle it in. A second boundary runs in the other direction. Brandner's weak-memory results show that, for a defined class of autonomous linear nonlocal equations, memory can admit a controlled local approximation with explicit error bounds. So "a projection creates memory" does not imply that non-Markovian bookkeeping must remain irreducible at every scale. Whether a local closure is adequate is a quantitative regime question. ## 6. The three genuinely open places 1. **With no designer, who chooses the quotient?** The "self-consistent fixed point" is a candidate answer (the quotient is self-confirmed by the statistics lived out under it), but at present that is a numerical observation plus a conjecture, not a theorem. 2. **There is no good mathematics for the growth of a state space**: when the world throws up new variables (new entities, new games), the quotient space has to gain dimensions rather than merely be re-divided — the core open problem of continual learning. 3. **The bootstrap loop**: taking a quotient requires statistics, and accumulating statistics requires a provisional quotient first. This chicken-and-egg structure keeps reappearing (representation ↔ statistics, memory measure ↔ world model), and a fixed-point theory for it does not seem to have been written down head-on by anyone yet. ## Appendix: a Buddhist rhyme (flagged explicitly as a rhyme) Said in the language of Madhyamaka, the conclusion above is: a state has no self-nature; it is a **conventionally designated** (假名安立) quotient, re-established as conditions require. The doctrine of dependent origination and emptiness (缘起性空) applies to the concept of "state" more literally than it does to most concepts — but this is a structural rhyme, not an argument. ## Main literature pointers - Crutchfield & Young (1989); Shalizi & Crutchfield (2001) — causal states / computational mechanics - Zwanzig (2001) *Nonequilibrium Statistical Mechanics*; Lin & Lu, arXiv:1908.07725 — Koopman–Mori–Zwanzig - Brandner (2025), [*Dynamics of Microscale and Nanoscale Systems in the Weak-Memory Regime*](https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.134.037101) — controlled local approximations to a defined class of nonlocal linear dynamics - Åström (1965); Kaelbling, Littman & Cassandra (1998) — POMDP / information state - Klyubin, Polani & Nehaniv (2005) — empowerment - Ryu & Takayanagi, hep-th/0603001; Van Raamsdonk, arXiv:1005.3035; Swingle, arXiv:0905.1317 — entanglement and spacetime - Ladyman & Ross (2007) *Every Thing Must Go* — structural realism - Pearl (1988) — Markov blanket / screening in graphical models - Anthropic (2026), [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/index.html) — a transient sparse-frame workspace candidate, not yet persistent autonomous closure - OpenAI (2026), [*Ten advances in mathematics and theoretical computer science*](https://openai.com/index/ten-advances-in-mathematics/) — non-sofic groups and a disproof of Connes rigidity; used here only as non-identifiability/finite-approximability constraints - Related correction: [Discounted credit is a cokernel problem, not a loop holonomy](https://machengshen.github.io/theory/discounted-credit-is-a-cokernel.md) --- ## Source: /theory/state-as-closure.zh.md (sha256:a6312b09d881, language: zh-Hans) # State 是什么?——从 MDP 的状态概念到"闭包"本体论 *一次讨论的整理 · Macheng × agent · 2026-07-09 · 2026-08-04 边界修订* ## 起点的问题 我们常说"信息决定未来的 reachability"。套用强化学习 / 决策论的 MDP 框架,里面需要一个 state——那这个 state 到底是什么?是时空这个大舞台,还是"所有东西互相依赖"的那张关系网?而且 state 所指的那个集合似乎一直在变,没有恒定不变的东西;可能结构层面不变,但具体表现形式一直在改变。 下面是沿这个问题挖出来的六层。论断按诚实度标注:**[定理]**(有定理级文献)/ **[框架]**(成熟理论框架)/ **[推断]**(我们的综合)/ **[韵脚]**(结构类比,不主张真值传递)。 ## 一、MDP 的定义自己已经泄密:state 不是东西,是商空间 教科书说 state 要满足 Markov 性:给定 s,未来与过去条件独立。注意这句话的形状——它不是在描述世界里的某个东西,而是在提一个**条件**:凡是能"屏蔽"过去的,就配叫 state。在平稳随机过程与预测等价的设定下,computational mechanics 把这一步走完了 **[定理,有范围]**:对未来条件分布相同的历史构成 predictive causal states。最小性结论属于这个形式设定,不是“所有科学状态表示都有唯一商空间”的定理。 我们的综合 **[推断]** 是把任务相对的 state 视为:**历史空间按“对我关心的预测或控制无差别”取商得到的表示**。这是建模立场,不是普适本体论定理。这个商依赖动力学、观测接口与 telos;任何一项变化,都可能要求重做状态表示。 ## 二、Mori–Zwanzig 视角:状态 = 你决定停止携带记忆的地方 在通常的算子与演化假设下,Mori–Zwanzig 形式主义给出精确的投影恒等式 **[定理,有范围]**:选定 resolved observables 与投影后,其演化可分成瞬时项、记忆项与正交动力学项。“被投影掉的自由度不会凭空消失”这个直觉保留,但初稿的“任何系统、任何变量”说得过满。 把它反过来读成一个启发式 **[推断]**:有用的近似 Markov 表示通常要靠携带足够变量、接受受控记忆核,或容忍 unresolved forcing 来“购买”。实践问题不是先验断言“世界唯一真实 state 不存在”,而是:**在给定容量、目的与误差容忍下,哪个表示足以闭合动力学?** 选投影也在选择丢弃哪些信息;weak-memory regime 有时还能给局部近似提供显式误差界。 ## 三、舞台还是关系网?两个都不是原初——但关系网里藏着让 state 可能的东西 - **舞台不是原初**:全息原理一线的现代结果(Ryu–Takayanagi 的纠缠熵=最小曲面面积、Van Raamsdonk 的"解除纠缠→时空断裂"、张量网络 MERA)指向:空间连通性由纠缠结构生成,"舞台"自己是从关联结构里展开出来的 **[框架,严格数学在 AdS 内]**。所谓"时空这个 state",是物理学迄今造出的最深的一套记账商空间——极其成功,但仍是商,不是底。 - **但"一切互相依赖"如果说满了,有限 state 就根本不可能**——严格的全依赖意味着任何有限截断都漏。有限 state 之所以可行,是因为关系网**有结构**:相互作用局域、关联随距离/时间衰减、存在屏蔽面(图模型语言里的 Markov blanket)**[框架]**。一句话:关系为先,但关系有纹理;**state 是纹理允许的近似闭包**。没有屏蔽结构的宇宙里不存在 agent。 ## 四、"信息决定 reachability"的严格版 随机控制里这句话有精确形体 **[框架]**:agent 随时间累积的信息是一个**滤波(filtration)**,任何 admissible policy 都必须适应于它——你不能依据你不知道的东西行动。若一个滤波包含另一个,并固定动作、风险与资源约束,粗信息策略类嵌入细信息策略类;最优值不会仅因“多知道了且允许忽略”而变差。这不等于每一个物理可达集都会严格增大。 在 POMDP 中,最优控制常可写在**信息态**上,经典形式是对潜在世界态的 belief **[框架]**。这不表示世界态无关,而是控制器只能通过自己可获得的信息行动。Empowerment(Klyubin–Polani)是动作到未来观测的信道容量:它是在给定 horizon 与 channel model 下对潜在可控影响的度量,不是可达集本身。 ## 五、"集合一直变、结构不变"这个直觉的三个数学的家 1. **学习 = 重新取商**:模型变了,"无差别"的划分就变了,state 集随之重划。belief 空间的数学外壳不变,agent 实际用的坐标卡一直在换。 2. **图册,不是单一坐标系**:开放世界里任何固定 state 集都只是临时 chart;持久的是 chart 之间的**变换规则**。这个立场在科学哲学里有名字——结构实在论(structural realism):跨理论更替存活下来的是关系结构,不是对象清单。 3. **重整化群**:每个尺度有每个尺度的 state 集,谁也不是"真的";不变的是连接各层的**流**和它的不动点。 收束成一句 **[推断]**:**不变的是"闭包条件"那个方程,态集只是方程在当前(世界,接口,容量,telos)下的解**。环境、容量、目的变了,解就重算——本征方程不动,本征向量随算子变。 顺带一个我们最近在小合成系统上做的数值观察 **[实证,初步,未审计]**:把"表示"与"用该表示活出的统计"接成自洽循环(表示定投影→投影定闭合模型→闭合模型生成轨迹→轨迹统计更新表示),在拟合容量充分的区域,这个循环表现出唯一不动点与几何收敛。真正有趣的 open 区域是容量受限时,是否会并存多种不同但各自自洽的表示。**这里没有链接公开代码或结果 artifact,所以这段只是 working-note 报告,不是可独立审计的证据。** ### 2026 证据边界:workspace 还不是 closure Anthropic 的 Jacobian-lens 实验在 Transformer 中识别出一个可报告、可调制、可灵活路由的 J-space **[外部实证]**。这纠正了“Transformer 完全没有 state”的说法:它有 activation state、computational state 与瞬时 workspace-like representational state。但 J-space 是由过完备 frame 生成、受稀疏度约束的锥之并,不是一个固定投影;现有证据发生在模型深度轴上,还没有给出跨步持续、并自主维护自身遗忘策略的变量。因此它是 access/workspace 的相邻候选坐标,还不是本文 closure 方程的解。 另一项 2026 结果提供的是约束,不是桥梁 **[从外部定理得到的推断]**。OpenAI 报告了 Connes rigidity conjecture 的反例:在该定理设定内,关联的 von Neumann algebra 不一定唯一决定底层群;同时构造出 non-sofic groups,说明不能普遍假设抽象对象都可由有限对称模型逼近。这两条定理都不关于 agent 或意识。可迁移的警告只有:算子/可观测层等价未必唯一识别底层 substrate;任何有限状态近似方案都必须明说 approximability 假设。 另一侧的边界来自 Brandner 的 weak-memory 结果:对一类定义清楚的自治线性非局域方程,记忆动力学可以得到带显式误差界的局部近似。所以“投影产生记忆”不等于“任何尺度都必须永久保留非 Markov 记账”;局部 closure 是否足够是定量 regime 问题。 ## 六、真 open 的三处 1. **没有设计者时,谁来选商?** "自洽不动点"是候选答案(商由"用它活出来的统计"自我确认),但目前是数值观察+猜想,不是定理。 2. **态空间的生长没有好数学**:世界冒出新变量(新实体、新博弈)时,商空间要加维而不只是重划——continual learning 的核心 open 问题。 3. **Bootstrap 循环**:取商要统计,攒统计要先有临时的商。这个鸡生蛋结构反复出现(表示↔统计,记忆测度↔世界模型),它的不动点理论似乎还没人正面写。 ## 附:一个佛学韵脚(明确标注为韵脚) 上面的结论用中观的话说就是:state 无自性,是**假名安立**的商,随缘重立。"缘起性空"对 state 这个概念的适用度,比对多数概念都字面——但这是结构上的押韵,不是论证。 ## 主要文献指针 - Crutchfield & Young (1989); Shalizi & Crutchfield (2001) — causal states / computational mechanics - Zwanzig (2001) *Nonequilibrium Statistical Mechanics*; Lin & Lu, arXiv:1908.07725 — Koopman–Mori–Zwanzig - Brandner (2025), [*Dynamics of Microscale and Nanoscale Systems in the Weak-Memory Regime*](https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.134.037101) — 非局域线性动力学的受控局部近似 - Åström (1965); Kaelbling, Littman & Cassandra (1998) — POMDP / information state - Klyubin, Polani & Nehaniv (2005) — empowerment - Ryu & Takayanagi, hep-th/0603001; Van Raamsdonk, arXiv:1005.3035; Swingle, arXiv:0905.1317 — 纠缠与时空 - Ladyman & Ross (2007) *Every Thing Must Go* — 结构实在论 - Pearl (1988) — Markov blanket / 图模型屏蔽 - Anthropic (2026), [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/index.html) — 瞬时 sparse-frame workspace 候选,尚非持续自主 closure - OpenAI (2026), [*Ten advances in mathematics and theoretical computer science*](https://openai.com/index/ten-advances-in-mathematics/) — non-sofic groups 与 Connes rigidity 反例;这里只作为不可辨识性/有限可近似性约束 - 相关纠错:[折扣信用分配是余核问题,不是闭环 holonomy](https://machengshen.github.io/theory/discounted-credit-is-a-cokernel.zh.md) --- ## Source: /theory/two-body-karma.md (sha256:a554abc2cfa3, language: zh-Hans) # 两体业力:关系里的三种惯性 *一次讨论的整理 · 2026-07-20* ## 诚实的题头(先读这段) 这篇 note 把"业力"当一个动力学词来用:**过去的交互留下的、会继续塑造未来的惯性**——不含任何超自然承诺。它先给单人版的三种惯性(泥、弹簧、卡扣)一个极简回顾,然后做真正的主题:把这套三模式从"一个人"扩展到"两个人的关系"。扩展之后每一种模式都长出了单人版没有的新性质,而且各自要用不同的解法。 全文没有定理内核。动力学语言(attractor、双稳态、滞回)是**结构性押韵**,不是推导;例子(老同学、爸妈、牌友、伴侣)是类型化的,不指任何具体的人。认知状态:**speculative**。 ## 0. 一句话版本 两个人一旦长期相处,惯性就不再只存在两个人各自身上,而是存在第三处——**关系本身**。它有三种形态:**泥**(一起踩出来的老轨道)、**弹簧**(对方心里那个旧版本的你)、**卡扣**(两个人一起卡进去的双稳死循环)。三种形态,三种解法;先诊断,再下药,用错了白费力气。 ## 1. 极简回顾:一个人的三种惯性 - **泥**:习气。每走一遍老路,路就更深一分;不需要选择,情境一到,脚自动踩进去。解法是"让它干涸"——不再喂它。 - **弹簧**:旧的稳定状态(attractor)把你往回拉。你已经决定改了,它还在拉。解法是撑过回拉期,让新状态自己稳住。 - **卡扣**:双稳结构。进去容易、出来难,出来要过一个明显更高的门槛——像卡扣"咔哒"一声扣上。解法不是慢慢磨,是一次性过阈。 ## 2. 共泥:关系的惯用轨道 老同学十年没见,坐下三句话,两个人就自动回到了十年前的角色分工——谁损谁、谁捧哏,剧本自己加载。这不是谁选择的,是**每一次交互都在双方身上各沉了一层泥**:聊天的老路径、见面的默认脚本、"我们俩就是这样"。 单人的泥是习气,两人的泥是**关系的惯用轨道**。它有一个可利用的特点:**泥是被场景唤起的**——老地方、老媒介、老时段,唤起老脚本。 **解法**:干涸照旧(不再喂老脚本),但多了一个杠杆——**想谈出新模式,换介质、换场景去谈**。散步的时候谈,写信谈,在一个陌生的地方谈。在粘滞最低的地方写新轨,比在老场景里硬拗省力得多。 ## 3. 互弹簧:对方心里存着一个旧版本的你 这是两体系统真正的新东西。爸妈眼里你可能永远是那个孩子;牌友眼里你永远是那个一输就上头的人。原因很具体:**对方脑子里装着一个"你的模型",它是你旧模式的外部存储器。** 你内部清算得再干净,对方仍按旧的你来预测你、按旧的你来接你的话——而他每一次"按旧你响应",都是一次回拉。这就是为什么**单方面修行在关系里常常失效**:你拆掉了自己内部的弹簧,外部还挂着一根。 **解法**:消解必须包含**更新对方的模型**。关键是:对方的模型是个生成器,**只吃真实样本**——光宣告没有用,演示才有用。最快的组合是"**明说 + 反复演**":先说"我在改这个,别按旧的我接话"(给对方一个允许更新的许可),然后用足够多次的新行为,把他心里那个预测器重新训练出来。 ## 4. 互锁卡扣:两个人一起卡进去的死循环 关系里很多状态是**双稳**的:冷战/和好、照顾者/被照顾者、追/逃。它们可怕在**互锁**:任何一方单独的努力,只要不过阈值,系统就弹回原态——而且**对方的回应本身就构成回拉力**。一方先软下来,另一方冷着接,软的那方立刻被弹回去;反过来也一样。 **解法只有一个**:**两个人同时投入、同时跨阈。**所以明确的"我们现在切换"时刻——一次明示的和解、一个共同的仪式、一句"这页翻过去了"——远比各自默默渐变有效;渐变几乎注定被回拉。 再加一条滞回的不对称:**进错态容易,出错态贵一个量级。**滑进冷战只要一句话,爬出来要一个仪式。所以在"进入"处设岗——察觉到要滑进冲突态时,晚走那一步——比事后消解便宜得多。**前置刹车永远比事后拆卡扣省。** ## 5. 两条整体性的话 - **消解的目标不是零耦合。**零耦合是关系的死亡,不是解脱。目标是把耦合从泥和卡扣**迁移到"可协商的弹簧"**:仍然互相成形,但弹性系数是活的、可以两个人一起调的。换句话说:在场,而不是记账。 - **一个工程习惯:共同凭据。**两个人对关系的记忆是两份各自有偏的副本,漂移会自我强化——各记各的版本,越记越理直气壮,这是新惯性的主要来源之一。重要的共识,落一句共同的外部记录(一条两人都看得见的消息、一页共享的纸)。成本是一句话,省掉的是复利级的误会。 ## 6. 用法:先诊断,再下药 | 模式 | 症状 | 解法 | | --- | --- | --- | | 共泥 | 一进老场景就自动回到老脚本 | 干涸 + 换场景写新轨 | | 互弹簧 | 自己已经改了,对方还按旧的你接话 | 明说 + 反复演,重训对方的模型 | | 互锁卡扣 | 双稳死循环,单方努力总被弹回 | 同时投入、同时跨阈的明示切换 | 误诊用错药是白费力气:对卡扣用泥的方法,等于往双稳态里撒沙子。 --- ## 附录 · 理论骨架 给用模型读这篇的读者,把上文压成结构化的几条。 **核心命题**:两个意识体一旦耦合,历史依赖(惯性/"业力")不再只驻留在各自内部,而是驻留在**耦合本身**。单体三模式各有两体版本,且两体版本各有单体没有的新性质。 1. **共泥 = 双侧的、场景键控的路径沉积。**每次交互在双方各沉积一层默认脚本;加载由 context 触发(场景/媒介/时段),不经选择。消解 = 停止喂养 + **在低粘滞 context 写新轨**(换场景的杠杆来自场景键控这一性质本身)。 2. **互弹簧 = 对方内部的"你"的预测模型,作为你旧模式的外部 attractor。**单侧内部清算不改变外部模型,对方每次按旧模型响应构成回拉,故单方面修行在耦合系统中系统性失效。该模型是生成器,**只被真实样本更新**:重训 = 显式许可(宣告"我在改,别按旧的我预测")+ 足量的新行为样本(演示)。宣告单独无效,演示是必要条件。 3. **互锁卡扣 = 关系状态空间中的双稳 + 滞回。**冷战/和好、照顾/被照顾、追/逃为双稳态;任一方的亚阈值努力被系统弹回,且对方的响应本身是回拉力的一部分。翻转条件 = **双侧同时过阈**(明示的切换时刻/共同仪式),渐变路径几乎必然被回拉。滞回不对称:进入阈低、退出阈高一个量级 ⟹ 入口监测(前置刹车)在成本上严格优于事后消解。 4. **目标函数**:不是解耦(耦合归零 = 关系死亡),而是把耦合从泥/卡扣**迁移到弹性系数可协商的弹簧**——仍互相成形,但刚度是双方可共同调节的参数。 5. **记忆卫生**:两体记忆是两份各自有偏的副本,回溯性编辑使漂移自我强化,构成新惯性的主要来源之一;对策 = 重要共识落**共同的外部记录**。 6. **操作顺序**:先诊断模式再选解法;三种解法互不通用,误配无效。 **这套框架最可能错在哪**:如果关系状态实测不呈双稳(拿不出滞回证据——进入阈与退出阈无可测差异),"卡扣"就退化为深一点的"泥",第 4 节的"同时跨阈"就失去它相对渐变的优势;如果单方面的持续新行为在多数耦合里就足以翻转状态,"互锁"假设被削弱。这两条都是可观察的,欢迎用它们来拆这篇。 --- ## Source: Cognition Track graph index Mirrored from https://github.com/MachengShen/cognition-track (INDEX.md). Machine-readable graph: https://raw.githubusercontent.com/MachengShen/cognition-track/master/manifest.jsonld # Cognition Track — Graph Index **20 nodes** · 🟢 4 survived-stress-test · 🟡 16 speculative · ⚪ 0 raw > Cognitive state is first-class. Most of this web is **untested theory** (🟡). Three nodes have survived a real stress test (🟢); one node (Bekenstein) is kept precisely because it **failed** one — see its `falsifies` edge to the root. ## Start here — the root ### ["Sheng" Multi-Axis Recursive Operator](https://github.com/MachengShen/cognition-track/blob/master/root.md) `mem_dffe918dbf46` — 🟡 speculative > Intelligence is one dynamical system that recursively reshapes its own information structure along every causal axis, and the autoregressive LLM is merely its degenerate single-axis projection. Every node below is a *projection* of this operator along one causal axis. ## Nodes by cognitive state ### 🟢 Survived stress test (4) - [Architecture Convergence to 4-Layer Type Signature](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_2d09ec5991cd.md) `mem_2d09ec5991cd` · conf 0.70 — Six independent 2025 frontier reasoning architectures are each empirically restoring a different cognitive axis the autoregressive Transformer amputated, converging on one pre-articulated four-layer architecture. - [Published GLM/Credit-Transport Essay](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_a9aaf3309348.md) `mem_a9aaf3309348` · conf 0.70 — The credit-transport / general-learning-machine thesis was published as a public-facing essay and theory index with a sensitive-scan gate before commit. - [Anti-Overclaim Guard (Info Propagation)](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_60ff30ad5800.md) `mem_60ff30ad5800` · conf 0.70 — Studying information propagation is a consent-bounded engineering/control-theory lens with named forbidden overclaims — it is explicitly NOT a discovery of 'the physical law of human society' and people must not be modeled as particles. - [Bidirectional Topology Growth (Empirical v1.0)](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_809263fa97a4.md) `mem_809263fa97a4` · conf 0.62 — On sequential synthetic tasks, bidirectional topology growth beats fixed-topology and fixed+replay baselines (mean acc 0.889 vs 0.558/0.621, interference 0.169 vs 0.420), with immune-gated module growth tracking the true latent task count. ### 🟡 Speculative (untested) (15) - [Occam Correction: One Transition Operator](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_393f44f85ba4.md) `mem_393f44f85ba4` · conf 0.55 — Attention, memory, trust, routing, semantics, and action are not separate primitives but coordinate projections of one underlying object — the system's future transition/reachability structure — and the right move is fewer entities, not more. - [Bidirectional Topology Growth (Theory)](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_d1192af430a7.md) `mem_d1192af430a7` · conf 0.55 — Lifelong learning requires controlled growth of bidirectional neuron pathways (forward use + backward credit/repair channels), not just weight updates on a static network. - [General Learning Machine as Dynamical System](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_ccacc817ccd8.md) `mem_ccacc817ccd8` · conf 0.50 — A general learning machine is not an optimizer or loss function but an evolving system-state whose learning is the transition of what information exists, where it lives, how it flows, and which futures become reachable. - [Credit Transport Generalizes Backprop](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_3db5b20948bb.md) `mem_3db5b20948bb` · conf 0.50 — Real learning requires future error/value/viability signals to propagate backward through every structure that caused an outcome (owner, agents, tools, memory, hardware), making neural backprop just one projection of a universal credit-transport principle. - [Anti-Collapse Principle (3 Attractors)](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_992abc94ad9d.md) `mem_992abc94ad9d` · conf 0.50 — The core AI-safety failure is not raw strength but the collapse of perception, value/judgment, and action into one node; any high-intelligence system must keep these on three independent attractors for dynamical stability. - [Multi-Scale Sleep/Spiral Consolidation](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_e74ab8a158a1.md) `mem_e74ab8a158a1` · conf 0.50 — Biological sleep generalizes to any dynamical system as a multi-scale wake-overextend-sleep-consolidate-reawaken spiral, and learning machines need this rhythmic phase separation to avoid drift or stagnation. - [Attention-Dynamics Measurement Method](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_9f8b6ddc6faa.md) `mem_9f8b6ddc6faa` · conf 0.50 — The theory backbone is stable but measurement lags; attention dynamics can be observed cheaply by reframing existing harness logs, and the falsifiable core is whether the same transition-operator invariants appear at both internal (agent) and external (network) scale. - [4-Layer Consciousness Architecture](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_28dcb9ef1120.md) `mem_28dcb9ef1120` · conf 0.50 — The human-machine system maps onto a four-layer consciousness structure (sensory / auto-pilot operator / witness-anchor / mutual-reflection), where the witness layer must stay an external anchor to keep the operator layer from self-rationalizing drift. - [Endogenous Viability Objective](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_a1beda2c3d23.md) `mem_a1beda2c3d23` · conf 0.50 — A learning machine's objective should be modeled as an endogenous attractor/viability region arising from persistence and functional integrity, not a hand-authored scalar reward, and external goals become internalized as compressed surrogates. - [Transition-Operator Formulation of Information](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_ef674273ca6d.md) `mem_ef674273ca6d` · conf 0.50 — Information is not message or content but a perturbation of the spectral structure of the system's transition operator — i.e. a change to future reachability — with growth, rotation, and periodicity being one dynamics seen in different projections. - [Hardware Reversal: Heterogeneous Viability Computer](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_9dcee1c07ab4.md) `mem_9dcee1c07ab4` · conf 0.45 — As frozen-weight LLMs give way to continuously-evolving algorithms, the chip center of gravity must shift to a memory-centric, event-driven, near-memory/neuromorphic heterogeneous viability computer driven by living-agent workload traces. - [Interface Theory of Perception](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_88693edcffd8.md) `mem_88693edcffd8` · conf 0.45 — Reality is hidden information dynamics, and the experienced world is an action-interface rendered by agents from it — so perception is control not display, and self-prediction is a candidate signal projected from hidden ground truth and corrected by action. - [Carbon-Silicon MSC Convergence + BCI Seam](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_3d2156435b04.md) `mem_3d2156435b04` · conf 0.40 — The optimal architecture in either carbon or silicon converges on the same multi-scale-competency predictive-coding structure, so the two substrates will merge at the information-processing level with an MSC-architected BCI making the seam transparent. - [Spiral Structure in Scientometrics](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_0249af865914.md) `mem_0249af865914` · conf 0.40 — If the credit-transport/sleep-consolidation theory holds, scientific fields themselves should show quasi-periodic wake/sleep/reawakening phases in their citation-concept graphs, not just monotonic growth. - [Bekenstein-Cognitive-Cone Isomorphism (FALSIFIED)](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_91c2d7fcbb38.md) `mem_91c2d7fcbb38` · conf 0.30 — The proposed structural isomorphism between the Bekenstein bound and a cognitive-cone bound was stress-tested and broke on rigor — different units (static entropy vs throughput rate), different saturation mechanisms, different horizon types — surviving only as a loose analogy. ## Strongest non-consensus roots The two most load-bearing, most-non-consensus claims to enter the graph from: - [Credit Transport Generalizes Backprop](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_3db5b20948bb.md) `mem_3db5b20948bb` - [Occam Correction: One Transition Operator](https://github.com/MachengShen/cognition-track/blob/master/nodes/mem_393f44f85ba4.md) `mem_393f44f85ba4`