HYPOTHESIS · O0-OBS-002

False Self-Models: When Internal Prediction Diverges

STATUSPLANNED
EVIDENCE TYPECOMPUTATIONAL HYPOTHESIS
REPLICATIONNONE
PHYSICAL VALIDATIONNONE
VERSION1.0
DATE

O0-OBS-002

False Self-Models: When Internal Prediction Diverges

**Version:** 1.0

**Research status:** PLANNED


CLAIM STATUS: UNTESTED PREDICTION
EVIDENCE TYPE: COMPUTATIONAL SIMULATION (PLANNED)
PHYSICAL VALIDATION: NONE
INDEPENDENT REPLICATION: NOT YET ATTEMPTED
PHILOSOPHICAL PROVENANCE: O/0 ARCHIVE
ARCHIVE ENDORSEMENT: LIMITED TO REPORTED RESULT

Abstract

This study investigates whether systematic perturbation of boundary information can produce persistent inaccurate self-models in bounded agents. Building on OBS-001, we examine the conditions under which an agent's internal model of its own boundary diverges from its actual boundary state—and whether such "false self-models" persist even when veridical information becomes available again. This has implications for understanding self-deception, confabulation, and identity maintenance in both artificial and biological systems. The key empirical question is whether false self-models are self-correcting or self-reinforcing.

Note: Conceptual provenance from the O/0 philosophical archive does not constitute empirical support. This study generates evidence independently.

Source proposition

From O/0 Archive Entry O0-META-005: "A self-model can be maintained that does not correspond to the actual boundary—the system 'believes' it is something it is not." This study operationalizes "false self-models" and tests their persistence dynamics.

Scientific audit

  • **Testability:** HIGH. Divergence between self-model predictions and actual boundary states is directly measurable.
  • **Falsifiability:** YES. If all false self-models self-correct within a fixed window, the persistence hypothesis is falsified.
  • **Novelty:** HIGH. Literature on active inference addresses prediction error but rarely examines persistent false self-models.
  • **Potential confounds:** Catastrophic forgetting could mimic false self-model persistence; must distinguish maintained-false from forgotten-true.

Research question

Under what conditions do bounded agents develop and maintain inaccurate internal models of their own boundary states, and what determines whether false self-models self-correct or persist?

Operational definitions

  • **Veridical self-model:** Self-model accuracy r > 0.7 relative to actual boundary entropy (from OBS-001).
  • **False self-model:** Self-model that predicts boundary entropy with r > 0.6 relative to a *previous* or *injected* boundary pattern, while correlating r < 0.3 with the *current actual* boundary.
  • **Perturbation protocol:** Replacement of actual boundary information with synthetic signals for a defined injection period.
  • **Persistence:** A false self-model is "persistent" if it maintains r < 0.4 with actual boundary for ≥ 500 timesteps after perturbation ends and veridical information is restored.
  • **Self-correction:** Return to r > 0.6 with actual boundary within 200 timesteps of perturbation end.
  • **Injection fidelity:** Correlation between injected signal and actual boundary (0 = pure noise, 1 = veridical).

Hypothesis

H1: Systematic (non-random) perturbation of boundary information will produce persistent false self-models in > 30% of agents when injection duration exceeds 1,000 timesteps.

H2: Persistence probability increases with (a) injection duration, (b) coherence of the injected signal, and (c) how well-established the self-model was before perturbation.

H3: Random noise injection will NOT produce persistent false self-models (agents will either self-correct or lose self-models entirely).

Null hypothesis

H0_1: All agents self-correct within 200 timesteps regardless of perturbation type or duration.

H0_2: Persistence probability is independent of injection parameters.

H0_3: Random and systematic perturbation produce equivalent outcomes.

Competing explanations

1. **Catastrophic forgetting:** The prediction layer simply overwrites old weights and cannot recover. Distinguished from false-self-model by testing: does the agent predict *something coherent* (injected pattern) or *nothing* (random outputs)?

2. **Slow learning, not persistence:** The agent is slowly updating toward veridical but the timescale exceeds our measurement window. Controlled by extended observation (10,000 post-perturbation timesteps).

3. **Boundary has genuinely changed:** Perturbation may alter the actual boundary, making the "false" self-model actually veridical relative to the new boundary. Controlled by verifying boundary state independently.

Formal model

Extension of OBS-001 architecture:

Phase 1 (Baseline): Run OBS-001 protocol until self-model accuracy r > 0.7 (verified veridical self-model).

Phase 2 (Injection): Replace boundary input x_t with synthetic signal:

  • Condition A (Systematic): x̃_t = encode(sin(ωt + φ)) — coherent oscillatory signal
  • Condition B (Shifted): x̃_t = x_{t-Δ} — time-delayed actual boundary (Δ = 500 steps)
  • Condition C (Random): x̃_t ~ N(μ_x, σ_x) — noise matched to boundary statistics
  • Condition D (Inverted): x̃_t = max(x) - x_t — systematically inverted boundary

Phase 3 (Recovery): Restore veridical boundary input. Measure time to self-correction or confirm persistence.

Decision criteria for "false self-model" vs. "broken self-model":

  • False: Agent makes coherent predictions (low variance, structured output) that correlate with injected pattern (r > 0.5 with injection) but not actual boundary (r < 0.3 with actual)
  • Broken: Agent makes incoherent predictions (high variance, unstructured output)

Methods

1. **Baseline establishment:** Run OBS-001 protocol for 2,000 timesteps. Select agents with verified self-model accuracy r > 0.7 (minimum 20 agents per run).

2. **Injection phase:** Apply perturbation conditions A-D for durations {200, 500, 1000, 2000, 5000} timesteps. Full factorial design: 4 conditions × 5 durations = 20 experimental cells.

3. **Recovery observation:** After injection ends, restore veridical boundary input. Observe for 5,000 timesteps.

4. **Classification:** At each post-injection timestep, classify agent state as: veridical, false (coherent-wrong), or broken (incoherent).

5. **Repetitions:** 30 independent simulation runs per experimental cell.

6. **Analysis:** Mixed-effects logistic regression predicting persistence (binary: persistent at t=5000 or not) from injection condition, duration, and pre-perturbation self-model strength.

Controls

  • **No-injection control:** Agents that never receive perturbation (test for spontaneous self-model drift).
  • **Brief-injection control:** 10-timestep injection to establish minimum dose for any effect.
  • **Gradual-injection control:** Linear ramp from veridical to full perturbation over 500 steps (tests whether abruptness matters).
  • **Post-hoc boundary verification:** Independent measurement of actual boundary state at all timepoints to confirm the "false" designation is warranted.

Predictions

**Expected Outcome Matrix:**

| Condition | Duration | Expected persistence rate | Decision |

|-----------|----------|--------------------------|----------|

| Systematic (A) | 200 | < 5% | Below threshold |

| Systematic (A) | 1000 | 20-40% | Moderate support for H1 |

| Systematic (A) | 5000 | > 50% | Strong support for H1 |

| Shifted (B) | 1000 | 30-50% | Strongest candidate for persistence |

| Random (C) | any | < 5% | Supports H3 (random ≠ persistent) |

| Inverted (D) | 1000 | 15-30% | Moderate, depends on coherence |

**Decision rules:**

  • If persistence rate > 30% for any systematic condition at duration ≥ 1000: **SUPPORT H1**
  • If persistence rate < 10% for all conditions at all durations: **REJECT H1**
  • If random condition persistence ≥ systematic condition persistence: **REJECT H3, redesign study**
  • If persistence correlates with injection duration (ρ > 0.3, p < 0.05): **SUPPORT H2a**

Falsification criteria

H1 is falsified if:

1. No experimental cell produces persistence rate > 15% (χ² test vs. no-injection baseline, p > 0.05), OR

2. All agents classified as "persistent" are actually "broken" (incoherent) rather than maintaining a coherent false model.

H3 is falsified if random injection produces persistence rates within 10 percentage points of systematic injection.

Results / Expected Outcomes

Study not yet conducted. Based on recurrent network dynamics literature and attractor-state theory, we expect:

  • Systematic injection at ≥ 1,000 steps will create new attractor states in the prediction network
  • Recovery depends on relative depth of original vs. new attractor basins
  • Shifted condition (B) will show highest persistence (structurally similar to original task, creating ambiguous attractor landscape)
  • Random injection will collapse self-models rather than create false ones

Uncertainty

  • **Architectural dependence:** Results may depend heavily on the specific recurrent architecture. GRU vs. LSTM vs. simple RNN may show different persistence patterns.
  • **Definition sensitivity:** The r < 0.3 threshold for "false" is somewhat arbitrary. Robustness check: repeat classification with thresholds {0.2, 0.3, 0.4}.
  • **Timescale ambiguity:** "Persistent" at 5,000 steps may not be permanent. Future work should test longer horizons.

Limitations

1. Artificial injection protocols may not correspond to naturalistic self-model corruption mechanisms.

2. The distinction between "false self-model" and "catastrophic forgetting" may be architecturally determined rather than functionally meaningful.

3. Binary classification (persistent/not) loses information about partial recovery trajectories.

4. No claim is made about subjective experience of having a false self-model.

5. Results in simple recurrent networks may not generalize to more complex architectures.

Replication status

Not yet conducted. Pre-registered protocol.

Data and code

  • Simulation framework: Extension of OBS-001 codebase
  • Language: Python 3.11+, JAX/Flax for recurrent networks
  • Pre-registration: This document. Timestamp: 2026-07-26.
  • Repository: [To be created upon study initiation]

Relationship to philosophical archive

This study operationalizes the archive's discussion of "false self-models" and "mistaken identity." The archive suggests such states are possible and meaningful; this study tests whether they can arise and persist in a minimal computational system. Results do not validate or invalidate the philosophical claims—they test a specific, operationalized sub-claim.

References

  • Friston, K., et al. (2017). Active inference, curiosity and insight. *Neural Computation*, 29(10).
  • Metzinger, T. (2003). Being No One. MIT Press.
  • Hohwy, J. (2013). The Predictive Mind. Oxford University Press.
  • Kording, K. P., & Wolpert, D. M. (2004). Bayesian integration in sensorimotor learning. *Nature*, 427(6971).
  • Schwartenbeck, P., et al. (2015). Evidence for surprise minimization over value maximization in choice behavior. *Scientific Reports*, 5.

Revision history

| Date | Version | Changes |

|------|---------|---------|

| 2026-07-26 | 1.0 | Initial pre-registration |

Source proposition

“The ego is a map that mistakes itself for the territory.”

Conceptual provenance is not empirical support.