SIMULATION · O0-CRP-037

SIM-CRP-003c · Cross-Architecture Perspective Distillation — Nearest-Neighbor Memory Observer

STATUSPARTIAL_MEMORY_SIGNAL. NNMO exhibits an IMMEDIATE compression cliff: BC_TV drops from 1.000 (N0 full memory) to 0.459 (N1 10 centroids), a loss of 0.541. N1..N4 plateau at BC_TV = 0.43-0.46 with random-baseline gaps 0.14-0.17. N5 collapses to random baseline. Even minimal memory retention (N4: one argmax-reward action) preserves modest signal above random (gap 0.136 TV, 0.146 AM). CROSS-ARCHITECTURE: at analogous compression, CRP-035 L4 preserves BC_TV=0.892 with 32 numbers; CRP-037 N1 preserves 0.459 with ~100 numbers. Identity carrier is not just capacity but capacity IN THE DECISION RULE - architecture matters.
EVIDENCE TYPECOMPUTATIONAL SIMULATION · Nearest-Neighbor Memory Observer (NNMO) - memory-based, no learned parameters. 6 compression levels (N0 full memory to N5 identifier), n=30 confirmatory. First cross-architecture replication of CRP-035 v1.2 F2.
REPLICATIONTHIS STUDY IS a cross-architecture replication of CRP-035 v1.2 F2. Registered followups: CRP-037-R1 (cluster count sweep), CRP-037-R2 (Bayesian PP observer as third arch), CRP-037-R3 (adversarial memory contamination).
PHYSICAL VALIDATIONNONE (memory-based computational model).
VERSION1.0.0
DATE

O0-CRP-037 · Cross-Architecture Perspective Distillation — Scientific Record

**Semantic name:** SIM-CRP-003c · Cross-Architecture Compression (Nearest-Neighbor Memory Observer)

**Record class:** SIMULATION

**Program:** Contact and Revelation (CRP-001)

**Non-drift question:** PERSPECTIVE

**Version:** 1.0.0

**Date:** 2026-07-30

**Status:** PARTIAL_MEMORY_SIGNAL (preregistered verdict)

**Preregistration:** [preregistration.md](preregistration.md) (frozen v1.0.0, 2026-07-30)

---

Claim-status banner


CLAIM STATUS   : PARTIAL_MEMORY_SIGNAL (preregistered)
EVIDENCE TYPE  : COMPUTATIONAL SIMULATION · Nearest-Neighbor Memory Observers
PHYSICAL VALID : NONE
INDEPENDENT REP: THIS IS CROSS-ARCHITECTURE REPLICATION OF CRP-035 v1.2 F2

SUPPORTED:
- NNMO exhibits an immediate compression cliff: BC_TV drops from 1.000 (N0
  full memory) to 0.459 (N1 10 centroids), a loss of 0.541.
- The N1..N4 plateau (BC_TV = 0.43-0.46) indicates that whatever behavior
  survives compression is preserved with essentially the same fidelity at
  4 to 10 retained memory items — the FIRST compression is the costly one.
- Even minimal memory retention (N4: one argmax-reward action) lifts
  behavior modestly above the random-parameter baseline (gap 0.136 TV,
  0.146 AM), showing that some identity signal survives 200-to-1 memory
  compression.
- CROSS-ARCHITECTURE FINDING: CRP-035 v1.2 F2 (high-capacity component
  carries identity) generalizes across architectures, but the DEGREE of
  compression tolerance is highly architecture-dependent. Parameterized
  observers (CRP-035) preserve BC_TV = 0.892 at L4 compression;
  memory-based observers (CRP-037) drop to BC_TV = 0.459 at analogous
  N1 compression, despite N1 retaining MORE numeric information (~100
  numbers vs 32 for L4).

NOT ESTABLISHED:
- Any preregistered "clean cliff to random" behavior — N4 remains above
  random baseline by ~0.14 TV.
- That memory-based observers universally require exact-memory retention;
  followups with different memory clustering (CRP-037-R1) will characterize
  the cliff shape.
- Any generalization to observer architectures other than linear-bandit
  parameter or nearest-neighbor memory. A Bayesian PP observer test is
  registered as CRP-037-R2.
- Any philosophical claim about identity or perspective.

---

Abstract

Preregistered cross-architecture replication of CRP-035 v1.2's F2 finding

(weights carry identity in a linear-bandit observer). Instead of a

parameterized observer whose decision rule uses learned parameters, this

study uses a Nearest-Neighbor Memory Observer (NNMO) whose decision rule

uses stored `(context, action, reward)` triples DIRECTLY: at each context,

the observer kernel-weights the k nearest stored contexts, votes by action,

and applies softmax with a per-observer temperature.

Six compression levels (N0..N5) progressively summarize the memory: full

memory → 10 k-means centroids → 4 per-action centroids → per-action mean

reward vector → single argmax-reward action → empty memory. NNMO has no

learned parameters other than memory, so this compression pipeline directly

probes what happens when the observer's ONLY identity carrier is

progressively summarized.

Confirmatory n = 30, 200 held-out contexts, adversarial random-parameter

and cross-observer baselines. Runtime: 1629 s.

Key findings:

  • Full-memory BC = 1.000 (identity gate passes).
  • N1 (10 k-means centroids, ~100 retained numbers): BC_TV = 0.459 vs

random 0.298 (gap 0.161). **A single compression step already drops BC

to less than half its starting value.**

  • N1..N4 plateau at BC_TV ≈ 0.43-0.46 with random-baseline gaps of

0.14-0.17. The compression is essentially binary in this architecture:

either you keep all of memory, or the fine-grained behavioral fidelity

is lost.

  • N5 (empty memory): BC_TV = 0.349, exactly the random baseline.
  • CROSS-ARCHITECTURE: at analogous "keep the high-capacity component"

compression, parameterized observers (CRP-035 L4) preserve BC_TV = 0.892

while memory-based observers (CRP-037 N1) preserve only 0.459 — a

0.43 gap despite N1 retaining more numeric information.

Verdict: **PARTIAL_MEMORY_SIGNAL** (preregistered). Full memory is required

for high behavioral fidelity in NNMO, but modest signal survives extreme

compression down to a single action.

1. Historical and conceptual background

CRP-035 v1.2 found that in a linear-bandit observer, the per-action weight

matrix `W ∈ R^{d×K}` carries essentially all identity-relevant behavioral

signal: L4 compression (W preserved, preferences discarded) gave

BC_TV = 0.892 vs random 0.282.

This finding is architecture-dependent by construction: linear-bandit

observers have identity in their weight parameters because that is where

learning is stored. A memory-based observer whose decision rule uses stored

experience directly could have a very different compression profile.

CRP-037 tests this: same environment class, same measurement pipeline,

GENUINELY different observer architecture.

2. Source-claim audit

No new philosophical source claim. CRP-037 is a mechanistic sub-question

of CRP-035's PERSPECTIVE branch.

3. Research question

For an observer whose decision rule uses stored `(context, action, reward)`

triples directly via nearest-neighbor kernel voting, at what level of

memory compression does behavior collapse to a random baseline?

4. Operational definitions

Inherited from CRP-035 v1.1/v1.2 §3 (BC_TV, BC_AM, SR, VP, PS).

5. Formal model

**Nearest-Neighbor Memory Observer (NNMO).** State

$\theta = (M, k, \sigma, \tau, \text{id})$ where:

  • $M$: FIFO memory of up to 200 `(x, a, r)` triples
  • $k \in \{3, 5, 7, 10\}$: number of neighbors used per decision
  • $\sigma \in [0.5, 2.0]$: kernel bandwidth
  • $\tau \in [0.15, 0.5]$: softmax temperature
  • id: stable identifier

Decision rule: for context $x$, find $k$ nearest neighbors from $M$ by

Euclidean distance, kernel-weight them by $w_i = \exp(-d_i^2 / 2\sigma^2)$,

compute per-action $\hat{Q}(a) = \sum_{i \in kNN, a_i=a} w_i r_i / \sum_{i \in kNN, a_i=a} w_i$,

and apply $\pi(a | x) = \text{softmax}(\hat{Q} / \tau)$.

No parameters other than memory are updated during training. `k`, `\sigma`,

`\tau` are per-observer constants sampled at instantiation.

Training environment: identical to CRP-035 v1.2 (linear-Gaussian rewards

per action). 200 training episodes populate memory.

6. Methods

Confirmatory sample: **n=30 observers**, seeds 18000..18029. Environment

seeds 18500..18529. Eval context seed 19000 (N=200 contexts). Probe seed

19500 (N=500 probes). VP env seeds 20000..20049 (N=50 envs × 20 episodes).

Compression pipeline (six levels): full memory → 10 k-means centroids →

4 per-action centroids → per-action mean rewards → argmax(reward) →

identifier only.

Random-parameter baseline at each level uses NNMOs with randomly-drawn

memories of matching size.

7. Results

| Level | Retained | BC_TV | BC_AM | BC_cos | SR | VP | PS | BC_TV_random | BC_AM_random | BC_TV_cross | BC_AM_cross |

|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|

| **N0** | Full memory (~200 triples) | **1.000** | **1.000** | 1.000 | 1.000 | 1.000 | 1.000 | 0.297 | 0.238 | 0.270 | 0.248 |

| **N1** | 10 k-means centroids | 0.459 | 0.434 | 0.525 | 0.067 | 0.545 | 0.330 | 0.298 | 0.250 | 0.275 | 0.246 |

| N2 | 4 per-action centroids | 0.445 | 0.414 | 0.538 | 0.133 | 0.212 | 0.261 | 0.299 | 0.260 | 0.294 | 0.247 |

| N3 | per-action mean r | 0.445 | 0.414 | 0.538 | 0.133 | 0.207 | 0.261 | 0.270 | 0.234 | 0.294 | 0.247 |

| N4 | argmax(reward) — 1 action | 0.430 | 0.414 | 0.478 | 0.133 | 0.240 | 0.186 | 0.294 | 0.268 | 0.276 | 0.247 |

| N5 | identifier only | 0.349 | 0.268 | 0.565 | 0.033 | 0.008 | 0.000 | 0.349 | 0.268 | 0.349 | 0.268 |

Random gaps at each level:

| Level | BC_TV - random | BC_AM - random | Interpretation |

|---|---:|---:|---|

| N0 | 0.703 | 0.762 | Full identity signal |

| N1 | 0.161 | 0.184 | Modest signal above random |

| N2 | 0.146 | 0.154 | Modest signal above random |

| N3 | 0.175 | 0.181 | Modest signal above random |

| N4 | 0.136 | 0.146 | Modest signal above random |

| N5 | 0.000 | 0.000 | Random baseline (by construction) |

7.1 Preregistered verdict application

Applied verdict rules from preregistration §8 in order:

1. **MEASUREMENT_FAILURE** — BC_TV(N0) = BC_AM(N0) = 1.000 · PASS.

2. **METRIC_SATURATION** — BC_TV_cross(N0) = 0.270, BC_AM_cross(N0) = 0.248 · PASS.

3. **MEMORY_CARRIES_IDENTITY_IN_NNMO** — H1: BC(N0) >= 0.95 ✓ AND

BC(N4) - BC_random(N4) < 0.10. **FAILS** — gap is 0.136 TV, 0.146 AM

(both > 0.10). Rule does not fire.

4. **PARTIAL_MEMORY_SIGNAL** — BC(N4) - BC_random(N4) >= 0.10 on either

metric. **FIRES** — gaps of 0.136 TV and 0.146 AM.

5. Rule 3 fires before rule 5, so no COMPRESSION_CLIFF verdict.

**Preregistered verdict: PARTIAL_MEMORY_SIGNAL.**

7.2 Cross-architecture comparison

Direct comparison of "keep the high-capacity component" compression across

two observer architectures:

| Observer | Level | Description | Retained numbers | BC_TV | BC_TV_random | Gap |

|---|---|---|---:|---:|---:|---:|

| CRP-035 v1.2 linear-bandit | L4 | W matrix + τ | 32 (+ τ) | **0.892** | 0.282 | **0.610** |

| CRP-037 NNMO | N1 | 10 k-means centroids | ~100 (+ k, σ, τ) | **0.459** | 0.298 | **0.161** |

Despite N1 retaining MORE numeric information than L4, it preserves MUCH

LESS behavior. The identity carrier is not just "high-capacity component"

— it is "high-capacity component that IS the decision rule."

In CRP-035, W is directly used to compute Q values via a compact linear

function; it captures the full learned policy in 32 numbers. In CRP-037,

memory is used via nearest-neighbor interpolation over stored triples;

compressing memory to 10 centroids destroys the interpolation ability

even though the centroids themselves preserve gross action structure.

7.3 Additional observations

  • **Plateau at N1..N4.** BC_TV stays in [0.43, 0.46] across all

intermediate compression levels. Once the first compression step is

taken, further compression loses little. This is different from CRP-035

v1.2 where BC declined smoothly from L4 (0.892) through L5 (0.458) to

L6 (0.326). NNMO's compression is essentially step-shaped.

  • **VP breakdown at N2 and N3.** Value preservation drops sharply from

N1 (0.545) to N2/N3 (0.21). Even though BC_TV plateaus, VP breaks.

This is because per-action means (N2, N3) don't discriminate reward

magnitudes well across actions, so the compressed observer picks

reasonable actions less often over long horizons.

  • **SR spike at N2/N3.** Self-recognition is HIGHER at N2/N3 (0.13) than

at N1 (0.07). This is likely because forced per-action centroids

(N2 always has one centroid per action) produce fingerprints that are

more distinguishable across observers than the noisier N1 centroids.

Documented but not the primary finding.

8. Uncertainty

Per-observer means with SEMs in `results/summary_confirmatory.json`. All

random-gap effects are >0.14 TV / >0.14 AM with SEM around 0.02.

9. Limitations

  • Single memory-based observer architecture. Other memory-based schemes

(episodic memory + LTM, dynamic prototypes, replay buffers) may compress

differently.

  • Single environment class (linear-Gaussian rewards) — same as CRP-035.
  • N1's compression uses simple k-means; smarter compression schemes

(learned prototypes, sketch data structures) may preserve more behavior.

  • The cross-architecture finding depends on the specific NNMO decision

rule; a different kernel (e.g. RBF over context+action) might preserve

more information under compression.

10. Alternative interpretations

  • **N1 loses more than L4 because it loses interpolation ability, not

because it has less "identity."** A single centroid represents many

contexts; compressing 200 items to 10 discards the fine-grained action

boundaries in context space. In CRP-035, W preserves those boundaries

by construction (linear function).

  • **N1 could preserve more with better clustering.** If k-means were

replaced with a learned prototype-selection method, N1 might approach

L4's fidelity. This would move the story from "memory-based observers

cliff" to "compression quality matters."

  • **The N0-to-N1 cliff might depend on training regime.** 200 episodes of

training produces a specific memory distribution; longer training would

produce more redundant memory and might compress better.

11. Replication procedure


cd research/studies/O0-CRP-037/src
python run_study.py --phase exploratory   # ~30 s
python run_study.py --phase confirmatory  # ~30 min
python summarize.py confirmatory

12. Code and data manifest

  • `preregistration.md` (v1.0.0)
  • `src/nnmo.py` — NNMO observer implementation
  • `src/nn_compressors.py` — 6 compression operators
  • `src/run_study.py` — driver
  • `src/summarize.py` — result-table printer
  • `results/summary_exploratory.json`
  • `results/summary_confirmatory.json`

13. Relationship to the philosophical archive

**Conceptual provenance is not empirical support.** CRP-037 tests a

mechanistic sub-question. Its indirect contribution: it establishes that

"identity carriers" are architecture-dependent. The philosophical claim

"perspective survives distillation" cannot be evaluated by any single

observer architecture; each architecture answers only "identity survives

under THESE compression rules for THIS decision procedure."

14. Registered follow-ups

  • **O0-CRP-037-R1** — Vary `k_c` in {3, 5, 10, 20, 50} to characterize

the compression cliff shape.

  • **O0-CRP-037-R2** — Third architecture: Bayesian PP observer.
  • **O0-CRP-037-R3** — Adversarial memory contamination: replace some

triples with random noise; measure BC vs contamination rate.

15. References

  • CRP-035 v1.2, CRP-036 v1.0.
  • Sutton & Barto, 2018 (contextual bandit background).
  • MacQueen, 1967 (k-means).
  • Cover & Hart, 1967 (nearest-neighbor classification).

16. Revision history

| Version | Date | Change |

|---|---|---|

| 1.0.0 | 2026-07-30 | Preregistration frozen and code drafted. Exploratory (n=6) and confirmatory (n=30) executed. Verdict: PARTIAL_MEMORY_SIGNAL. Cross-architecture comparison to CRP-035 v1.2 L4 documented. |

Source proposition

“Cross-architecture test of CRP-035 v1.2 F2: does the "high-capacity component carries identity" finding generalize across observer architectures?”

Conceptual provenance is not empirical support.