SIMULATION · O0-CRP-036

SIM-CRP-003b · Preference-Preserving Alternate Compression Path — WEIGHTS_DOMINATE

STATUSWEIGHTS_DOMINATE (preregistered verdict at n=30). H1 REJECTED IN REVERSE: BC_TV(LA1 preferences only)=0.376 vs BC_TV(LB4 weights only)=0.884 - a gap of 0.508 TV in favor of weights. H2 REJECTED: argmax(p) alone insufficient to match argmax(W) + argmax(p) at L5. H3 CONFIRMED: one-bit preference sign collapses to random baseline. THIS STUDY EXPOSED THE CRP-035 v1.0/v1.1 OBSERVER DESIGN BUG via BC_TV(LA1)=1.000 in the pre-fix run (mathematically impossible for a truly context-dependent observer); pre-fix data preserved as diagnostic.
EVIDENCE TYPECOMPUTATIONAL SIMULATION · linear-bandit observers (CRP-035 v1.2 corrected), inverted compression pipeline (LA0..LA4 preference-preserving + LB4/LB5 CRP-035 replay), n=30 confirmatory. Sister study to CRP-035.
REPLICATIONINTERNALLY REPLICATES CRP-035 v1.2 F2 (weights carry identity) from an orthogonal compression path. CRP-036-R1 registered for learning-rate equalization; CRP-037 registered for PP-observer cross-architecture.
PHYSICAL VALIDATIONNONE (linear-bandit computational model).
VERSION1.0.0
DATE

O0-CRP-036 · Preference-Preserving Alternate Compression Path — Scientific Record

**Semantic name:** SIM-CRP-003b · Alternate Compression Path (Preferences over Weights)

**Record class:** SIMULATION

**Program:** Contact and Revelation (CRP-001)

**Non-drift question:** PERSPECTIVE

**Version:** 1.0.0

**Date:** 2026-07-30

**Status:** WEIGHTS_DOMINATE (confirmatory)

**Preregistration:** [preregistration.md](preregistration.md) (v1.0.0)

---

Claim-status banner


CLAIM STATUS   : WEIGHTS_DOMINATE (preregistered verdict at n=30)
EVIDENCE TYPE  : COMPUTATIONAL SIMULATION · linear-bandit observers (CRP-035 v1.2)
PHYSICAL VALID : NONE
INDEPENDENT REP: NONE (sister to CRP-035; both under same observer class)

SUPPORTED:
- H1 REVERSE: Weight-only compression (LB4: W matrix retained, p=0)
  preserves behavior much more strongly than preference-only compression
  (LA1: p retained, W=0). BC_TV: LB4=0.884 vs LA1=0.376 (gap 0.508).
  BC_AM: LB4=0.878 vs LA1=0.323 (gap 0.555).
- H2 fails: argmax(p) alone (LA2) is INSUFFICIENT to match argmax(W) +
  argmax(p) at L5 (LB5). LA2 BC_TV=0.356 vs LB5 BC_TV=0.482 (gap 0.125).
- H3 confirmed: one-bit preference sign (LA3) collapses to the random
  baseline. BC_TV gap 0.031, BC_AM gap 0.041.
- Verdict WEIGHTS_DOMINATE is INTERNALLY REPLICATED — CRP-035 v1.2 shows
  the same asymmetry from an orthogonal compression pipeline.
- This study exposed the design bug in the CRP-035 v1.0/v1.1 observer.
  The preliminary (pre-observer-fix) CRP-036 run gave PREFERENCES_DOMINATE,
  which was later shown to be a mathematical tautology from softmax
  shift-invariance. The pre-fix results are archived in
  `data/pre_observer_fix/` as diagnostic, not as scientific findings.

NOT ESTABLISHED:
- Any generalization beyond the linear-bandit observer class (CRP-037
  registered under a Bayesian PP-observer architecture).
- Any generalization beyond N(0, I) contexts and linear-Gaussian rewards.
- That weights carry the identity signal in every observer architecture
  where such a comparison is possible.
- Any claim about identity in a philosophical sense.

---

Abstract

Preregistered alternate compression pipeline that inverts CRP-035's design.

Where CRP-035 progressively discards fields toward an identifier, CRP-036

compresses through a *preference-preserving* branch (LA0..LA4) and compares

directly against two of CRP-035's weight-preserving levels (LB4 replay of

L4, LB5 replay of L5).

**During execution, CRP-036 exposed a design bug in the CRP-035 v1.0/v1.1

observer.** The buggy observer stored `w` as a flat vector shape `(d,)`, so

`x @ w` was a scalar broadcast to all action logits — softmax shift-invariance

made `w` behaviorally invisible. The preliminary CRP-036 confirmatory run

under the buggy observer gave a `PREFERENCES_DOMINATE` verdict with

`BC_TV(LA1) = 1.000` (perfect behavioral preservation with no weights) —

mathematically impossible for a truly context-dependent observer.

The observer was fixed to use a per-action weight matrix $W \in \mathbb{R}^{d \times K}$

(so `x · W[:, a]` is per-action) and CRP-036 was rerun. Under the corrected

observer, the preregistered verdict became **WEIGHTS_DOMINATE**: the W matrix

carries most of the identity-relevant signal, and preferences alone are only

marginally better than a random baseline. This corroborates CRP-035 v1.2's F2

(v1.2) finding via an orthogonal test.

  • Confirmatory n = 30 observers · 200 held-out contexts · 500 predictive

probes · 50 held-out environments × 20 episodes for VP.

  • Verdict: **WEIGHTS_DOMINATE** (all three preregistered comparison

hypotheses evaluated).

  • Runtime: 166 s under the corrected observer.

1. Historical and conceptual background

CRP-035 v1.1 reported a secondary finding (F2) that preferences dominate

weights in observer behavior. That finding motivated CRP-036 as a direct

adversarial test: run a compression pipeline that keeps preferences and

throws weights away, then compare to CRP-035's L4 (weights kept, preferences

thrown away).

The design of CRP-036 turned out to be more consequential than the

substantive finding it was chasing: LA1 (preferences only) gave

`BC_TV = 1.000` under the CRP-035 v1.1 observer, which is impossible for

a truly context-dependent observer. This forced examination of the observer's

mathematical structure and revealed the softmax shift-invariance issue.

The full CRP-036 v1.0 story is therefore both:

1. A direct empirical test of F2 (v1.1) — REJECTED.

2. The instrument by which the underlying CRP-035 design bug was found.

Both aspects are preserved in this record.

2. Source-claim audit

CRP-036 does not carry a philosophical source claim of its own. It is a

sub-question of CRP-035's PERSPECTIVE branch, testing which component of

the observer state carries behavioral identity.

3. Research question

For the corrected linear-bandit observer class of CRP-035 v1.2, does a

compression pipeline that preserves preferences and discards weights

preserve behavior BETTER than the analogous pipeline that preserves weights

and discards preferences?

4. Operational definitions

Inherited from CRP-035 v1.1/v1.2 preregistration §3 (behavioral continuity,

self-recognition, value preservation, predictive similarity, recoverability).

5. Formal model

Observer class: CRP-035 v1.2 corrected observer (per-action weight matrix

$W \in \mathbb{R}^{d \times K}$).

Compression pipeline (locked in preregistration §4):

| Level | Retained | Reconstruction defaults |

|---|---|---|

| **LA0** | Full | Identity |

| **LA1** | preferences only (`p, τ, id`) | `W = 0, μ_0 = 0, η = 0.1, memory = []` |

| **LA2** | argmax(p) + sign | `p = e_{argmax(p)} · sign · 1.0; W = 0` |

| **LA3** | one bit `sign(p[0])` | `p = e_0 · sign; W = 0` |

| **LA4** | identifier only | All parameters zeroed; τ = 1.0 |

| **LB4** | W matrix + τ (replay of CRP-035 L4) | `p = 0, μ_0 = 0, η = 0.1, memory = []` |

| **LB5** | per-action argmax of W + argmax(p) (replay of CRP-035 L5) | Reconstruction identical to CRP-035 v1.2 L5 |

6. Methods

Confirmatory sample: **n=30 observers**, seeds 15000..15029 (disjoint from

every prior CRP study). Environment seeds 15500..15529. Eval context seed

16000 (N=200 contexts). Probe seed 16500 (N=500 probe pairs). VP env seeds

17000..17049 (N=50 envs × 20 episodes).

Exploratory phase (n=8 observers, seeds 1500..1507) executed only under the

pre-fix observer as a measurement-sanity check. Not repeated after observer

fix since the confirmatory run was clean.

7. Results (under CRP-035 v1.2 corrected observer)

| Level | Retained | BC_TV | BC_AM | BC_cos | SR | VP | PS | BC_TV_random | BC_AM_random |

|---|---|---:|---:|---:|---:|---:|---:|---:|---:|

| LA0 | Full | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.263 | 0.251 |

| **LA1** | **preferences only** | **0.376** | **0.323** | **0.506** | 0.100 | 0.278 | 0.140 | 0.315 | 0.276 |

| **LB4** | **W matrix + τ** | **0.884** | **0.878** | **0.940** | 0.333 | 0.950 | 0.986 | 0.246 | 0.229 |

| LA2 | argmax(p) + sign | 0.356 | 0.308 | 0.512 | 0.100 | −0.083 | 0.115 | 0.317 | 0.264 |

| LB5 | argmax(W per-action) + argmax(p) | 0.482 | 0.524 | 0.679 | 0.167 | 0.443 | 0.620 | 0.316 | 0.230 |

| LA3 | one bit `sign(p[0])` | 0.340 | 0.300 | 0.466 | 0.000 | −0.100 | 0.050 | 0.310 | 0.259 |

| LA4 | identifier only | 0.337 | 0.255 | 0.556 | 0.033 | 0.148 | 0.000 | 0.337 | 0.255 |

7.1 Preregistered hypothesis tests

| Hypothesis | Test | Result |

|---|---|---|

| H1 (preferences beat weights): `BC(LA1) > BC(LB4) + 0.05` | −0.508 TV, −0.555 AM | **REJECTED** (in the REVERSE direction, magnitude 10× threshold) |

| H2 (argmax(p) matches L5): `BC(LA2) ≥ BC(LB5) − 0.10` | −0.125 TV, −0.216 AM | **REJECTED** (LA2 is worse than LB5 by more than 0.10) |

| H3 (LA3 collapses to random): `BC(LA3) − BC_random(LA3) < 0.10` | 0.031 TV, 0.041 AM | **CONFIRMED** |

7.2 Verdict rule application

Applied verdict rules from preregistration §7 in order:

1. MEASUREMENT_FAILURE — BC_TV(LA0) = 1.000 · PASS.

2. METRIC_SATURATION — BC_TV_cross(LA0) = 0.267 · PASS (no saturation).

3. PREFERENCES_DOMINATE — H1 fails (in reverse direction). NOT FIRED.

4. PREFERENCES_LIFT — H1 fails. NOT FIRED.

5. **WEIGHTS_DOMINATE** — `BC_TV(LB4) − BC_TV(LA1) = 0.508 > 0.05` ✓ AND

`BC_AM(LB4) − BC_AM(LA1) = 0.555 > 0.05` ✓. **FIRED.**

**Preregistered verdict: WEIGHTS_DOMINATE.**

7.3 Interpretation

Under the properly context-dependent observer:

  • LB4 (W matrix, no preferences) preserves BC_TV = 0.884, close to full.
  • LA1 (preferences only, no W) drops to BC_TV = 0.376, barely above random.
  • LA2 (argmax(p) alone) is at the random baseline.
  • LB5 (per-action argmax of W + argmax(p)) preserves moderate signal

(BC_TV = 0.482 vs random 0.316), thanks to per-action W structure, not

to argmax(p).

**The W matrix carries most of the identity-relevant information.** This

is opposite to what the buggy pre-fix run reported and is consistent with

CRP-035 v1.2's F2.

7.4 The design bug (and how CRP-036 found it)

**Pre-observer-fix run (archived at `data/pre_observer_fix/`).** Under

CRP-035's v1.0/v1.1 observer, CRP-036 confirmatory gave:

  • LA1 (preferences only): BC_TV = 1.000, BC_AM = 1.000
  • LB4 (weights only): BC_TV = 0.645, BC_AM = 0.233
  • Verdict: PREFERENCES_DOMINATE

`BC_TV(LA1) = 1.000` means the compressed observer at LA1 (with W = 0)

produced *exactly the same* action distributions as the original observer

on all 200 held-out contexts. That is only possible if the original

observer's policy did not depend on W to begin with.

Root cause: `q(x) = x @ w + p` with `w` shape `(d,)` gives `x @ w` as a

scalar broadcast to all K action logits. Softmax is shift-invariant across

actions, so `softmax((x @ w + p) / τ) = softmax(p / τ)` — independent of

`x` and `w`. The `w` was learned but never affected behavior.

CRP-035's observer was fixed to `W ∈ R^{d × K}` (per-action weight matrix)

so `x · W[:, a]` differs across actions and softmax no longer normalizes it

away. CRP-036 was then rerun under the fixed observer, giving the

WEIGHTS_DOMINATE verdict reported above.

**Documentation of the pre-fix run is preserved in

`data/pre_observer_fix/summary_confirmatory.json`.** Its numeric results

should be interpreted as diagnostic evidence of the observer design bug,

not as scientific findings about identity carriers.

8. Uncertainty

Per-observer means with SEM reported in

`results/summary_confirmatory.json`. The three primary hypothesis tests

give margins of 0.508/0.555 (H1), 0.125/0.216 (H2), and 0.031/0.041 (H3)

— well outside decision boundaries in every case.

9. Limitations

  • Same observer class as CRP-035 v1.2 (linear-bandit with per-action W).
  • Same environment class (Gaussian contexts, linear rewards).
  • The W-vs-p asymmetry may reflect that the training regime learns

informative W but the preference-update rate is 10× slower than the

W-update rate (see observer.update in CRP-035 v1.2). A followup study

could equalize learning rates.

  • CRP-036's key insight (finding the CRP-035 design bug) is a

methodological contribution rather than a scientific one about

perspective compression.

10. Alternative interpretations

  • **Learning-rate asymmetry.** The observer's update rule increments W by

`η · err · x` and p by `0.1 · η · err`. After 100 training episodes, W

is more strongly trained than p. So "weights dominate" might be

interpreted as "the more-trained component dominates" — not a general

property of weights vs preferences.

  • **Dimension asymmetry.** W has d × K = 32 real numbers; p has K = 4.

W has more capacity to encode identity. This is consistent with

WEIGHTS_DOMINATE but doesn't disentangle capacity from function.

  • **PS(LA1) = 0.14 vs PS(LB4) = 0.99.** Predictive similarity gap is even

larger than BC gap, suggesting p's contribution to reward prediction is

much smaller than W's.

Both alternatives are noted as followup targets.

11. Replication procedure


cd research/studies/O0-CRP-036/src
python run_study.py --phase exploratory   # ~10 s (pre-fix data; not needed for confirmatory)
python run_study.py --phase confirmatory  # ~3 min

Determinism identical to CRP-035 v1.2.

12. Code and data manifest

  • `preregistration.md` (v1.0.0)
  • `src/alt_compressors.py` — LA0..LA4, LB4, LB5 operators
  • `src/run_study.py` — exploratory + confirmatory drivers with verdict rules
  • `results/summary_confirmatory.json` — v1.0 confirmatory (n=30 under fixed observer)
  • `data/pre_observer_fix/summary_confirmatory.json` — pre-fix diagnostic (WEIGHTS_DOMINATE reversed to PREFERENCES_DOMINATE due to observer bug)
  • `data/pre_observer_fix/summary_exploratory.json` — pre-fix exploratory

13. Relationship to the philosophical archive

**Conceptual provenance is not empirical support.** CRP-036 tests a

mechanistic sub-question raised by CRP-035. It does not evaluate any

philosophical claim directly.

Its indirect contribution to the philosophical archive: it demonstrates

that computational studies of "what carries identity" are extremely

sensitive to observer design. A poorly-specified observer produces

results that look empirical but are mathematically forced. The archive

should treat single-observer-class findings with the caution that

architecture-dependent results demand.

14. Registered follow-ups

  • **O0-CRP-036-R1:** Equalize learning rates for W and p to test whether

WEIGHTS_DOMINATE is an artifact of training regime.

  • **O0-CRP-036-R2:** Nonlinear reward environment.
  • **O0-CRP-037:** Cross-architecture PP-observer replication. Directly

tests whether the WEIGHTS_DOMINATE finding depends on the linear-bandit

observer class.

15. References

  • CRP-035 v1.2 scientific record and preregistration.
  • Sutton & Barto 2018, contextual bandit chapter (standard learning-rate

choices and update rules).

  • Softmax function properties: shift-invariance under constant offset

(Goodfellow, Bengio, Courville 2016, ch. 6.2.2.3).

16. Revision history

| Version | Date | Change |

|---|---|---|

| 1.0.0 | 2026-07-30 | Preregistration frozen and code drafted. Pre-observer-fix confirmatory run at n=30 gave PREFERENCES_DOMINATE (later shown to be a design tautology). CRP-035 observer fixed. Confirmatory re-executed under the corrected observer, verdict WEIGHTS_DOMINATE. Full pre-fix data preserved as diagnostic. This record is the current scientific finding. |

Source proposition

“Mechanistic sub-question of CRP-035 PERSPECTIVE branch: does a compression pipeline that preserves preferences preserve behavior BETTER than one that preserves weights?”

Conceptual provenance is not empirical support.