O0-OBS-002
False Self-Models: When Internal Prediction Diverges
**Version:** 1.0
**Research status:** PLANNED
CLAIM STATUS: UNTESTED PREDICTION
EVIDENCE TYPE: COMPUTATIONAL SIMULATION (PLANNED)
PHYSICAL VALIDATION: NONE
INDEPENDENT REPLICATION: NOT YET ATTEMPTED
PHILOSOPHICAL PROVENANCE: O/0 ARCHIVE
ARCHIVE ENDORSEMENT: LIMITED TO REPORTED RESULT
Abstract
This study investigates whether systematic perturbation of boundary information can produce persistent inaccurate self-models in bounded agents. Building on OBS-001, we examine the conditions under which an agent's internal model of its own boundary diverges from its actual boundary state—and whether such "false self-models" persist even when veridical information becomes available again. This has implications for understanding self-deception, confabulation, and identity maintenance in both artificial and biological systems. The key empirical question is whether false self-models are self-correcting or self-reinforcing.
Note: Conceptual provenance from the O/0 philosophical archive does not constitute empirical support. This study generates evidence independently.
Source proposition
From O/0 Archive Entry O0-META-005: "A self-model can be maintained that does not correspond to the actual boundary—the system 'believes' it is something it is not." This study operationalizes "false self-models" and tests their persistence dynamics.
Scientific audit
- **Testability:** HIGH. Divergence between self-model predictions and actual boundary states is directly measurable.
- **Falsifiability:** YES. If all false self-models self-correct within a fixed window, the persistence hypothesis is falsified.
- **Novelty:** HIGH. Literature on active inference addresses prediction error but rarely examines persistent false self-models.
- **Potential confounds:** Catastrophic forgetting could mimic false self-model persistence; must distinguish maintained-false from forgotten-true.
Research question
Under what conditions do bounded agents develop and maintain inaccurate internal models of their own boundary states, and what determines whether false self-models self-correct or persist?
Operational definitions
- **Veridical self-model:** Self-model accuracy r > 0.7 relative to actual boundary entropy (from OBS-001).
- **False self-model:** Self-model that predicts boundary entropy with r > 0.6 relative to a *previous* or *injected* boundary pattern, while correlating r < 0.3 with the *current actual* boundary.
- **Perturbation protocol:** Replacement of actual boundary information with synthetic signals for a defined injection period.
- **Persistence:** A false self-model is "persistent" if it maintains r < 0.4 with actual boundary for ≥ 500 timesteps after perturbation ends and veridical information is restored.
- **Self-correction:** Return to r > 0.6 with actual boundary within 200 timesteps of perturbation end.
- **Injection fidelity:** Correlation between injected signal and actual boundary (0 = pure noise, 1 = veridical).
Hypothesis
H1: Systematic (non-random) perturbation of boundary information will produce persistent false self-models in > 30% of agents when injection duration exceeds 1,000 timesteps.
H2: Persistence probability increases with (a) injection duration, (b) coherence of the injected signal, and (c) how well-established the self-model was before perturbation.
H3: Random noise injection will NOT produce persistent false self-models (agents will either self-correct or lose self-models entirely).
Null hypothesis
H0_1: All agents self-correct within 200 timesteps regardless of perturbation type or duration.
H0_2: Persistence probability is independent of injection parameters.
H0_3: Random and systematic perturbation produce equivalent outcomes.
Competing explanations
1. **Catastrophic forgetting:** The prediction layer simply overwrites old weights and cannot recover. Distinguished from false-self-model by testing: does the agent predict *something coherent* (injected pattern) or *nothing* (random outputs)?
2. **Slow learning, not persistence:** The agent is slowly updating toward veridical but the timescale exceeds our measurement window. Controlled by extended observation (10,000 post-perturbation timesteps).
3. **Boundary has genuinely changed:** Perturbation may alter the actual boundary, making the "false" self-model actually veridical relative to the new boundary. Controlled by verifying boundary state independently.
Formal model
Extension of OBS-001 architecture:
Phase 1 (Baseline): Run OBS-001 protocol until self-model accuracy r > 0.7 (verified veridical self-model).
Phase 2 (Injection): Replace boundary input x_t with synthetic signal:
- Condition A (Systematic): x̃_t = encode(sin(ωt + φ)) — coherent oscillatory signal
- Condition B (Shifted): x̃_t = x_{t-Δ} — time-delayed actual boundary (Δ = 500 steps)
- Condition C (Random): x̃_t ~ N(μ_x, σ_x) — noise matched to boundary statistics
- Condition D (Inverted): x̃_t = max(x) - x_t — systematically inverted boundary
Phase 3 (Recovery): Restore veridical boundary input. Measure time to self-correction or confirm persistence.
Decision criteria for "false self-model" vs. "broken self-model":
- False: Agent makes coherent predictions (low variance, structured output) that correlate with injected pattern (r > 0.5 with injection) but not actual boundary (r < 0.3 with actual)
- Broken: Agent makes incoherent predictions (high variance, unstructured output)
Methods
1. **Baseline establishment:** Run OBS-001 protocol for 2,000 timesteps. Select agents with verified self-model accuracy r > 0.7 (minimum 20 agents per run).
2. **Injection phase:** Apply perturbation conditions A-D for durations {200, 500, 1000, 2000, 5000} timesteps. Full factorial design: 4 conditions × 5 durations = 20 experimental cells.
3. **Recovery observation:** After injection ends, restore veridical boundary input. Observe for 5,000 timesteps.
4. **Classification:** At each post-injection timestep, classify agent state as: veridical, false (coherent-wrong), or broken (incoherent).
5. **Repetitions:** 30 independent simulation runs per experimental cell.
6. **Analysis:** Mixed-effects logistic regression predicting persistence (binary: persistent at t=5000 or not) from injection condition, duration, and pre-perturbation self-model strength.
Controls
- **No-injection control:** Agents that never receive perturbation (test for spontaneous self-model drift).
- **Brief-injection control:** 10-timestep injection to establish minimum dose for any effect.
- **Gradual-injection control:** Linear ramp from veridical to full perturbation over 500 steps (tests whether abruptness matters).
- **Post-hoc boundary verification:** Independent measurement of actual boundary state at all timepoints to confirm the "false" designation is warranted.
Predictions
**Expected Outcome Matrix:**
| Condition | Duration | Expected persistence rate | Decision |
|-----------|----------|--------------------------|----------|
| Systematic (A) | 200 | < 5% | Below threshold |
| Systematic (A) | 1000 | 20-40% | Moderate support for H1 |
| Systematic (A) | 5000 | > 50% | Strong support for H1 |
| Shifted (B) | 1000 | 30-50% | Strongest candidate for persistence |
| Random (C) | any | < 5% | Supports H3 (random ≠ persistent) |
| Inverted (D) | 1000 | 15-30% | Moderate, depends on coherence |
**Decision rules:**
- If persistence rate > 30% for any systematic condition at duration ≥ 1000: **SUPPORT H1**
- If persistence rate < 10% for all conditions at all durations: **REJECT H1**
- If random condition persistence ≥ systematic condition persistence: **REJECT H3, redesign study**
- If persistence correlates with injection duration (ρ > 0.3, p < 0.05): **SUPPORT H2a**
Falsification criteria
H1 is falsified if:
1. No experimental cell produces persistence rate > 15% (χ² test vs. no-injection baseline, p > 0.05), OR
2. All agents classified as "persistent" are actually "broken" (incoherent) rather than maintaining a coherent false model.
H3 is falsified if random injection produces persistence rates within 10 percentage points of systematic injection.
Results / Expected Outcomes
Study not yet conducted. Based on recurrent network dynamics literature and attractor-state theory, we expect:
- Systematic injection at ≥ 1,000 steps will create new attractor states in the prediction network
- Recovery depends on relative depth of original vs. new attractor basins
- Shifted condition (B) will show highest persistence (structurally similar to original task, creating ambiguous attractor landscape)
- Random injection will collapse self-models rather than create false ones
Uncertainty
- **Architectural dependence:** Results may depend heavily on the specific recurrent architecture. GRU vs. LSTM vs. simple RNN may show different persistence patterns.
- **Definition sensitivity:** The r < 0.3 threshold for "false" is somewhat arbitrary. Robustness check: repeat classification with thresholds {0.2, 0.3, 0.4}.
- **Timescale ambiguity:** "Persistent" at 5,000 steps may not be permanent. Future work should test longer horizons.
Limitations
1. Artificial injection protocols may not correspond to naturalistic self-model corruption mechanisms.
2. The distinction between "false self-model" and "catastrophic forgetting" may be architecturally determined rather than functionally meaningful.
3. Binary classification (persistent/not) loses information about partial recovery trajectories.
4. No claim is made about subjective experience of having a false self-model.
5. Results in simple recurrent networks may not generalize to more complex architectures.
Replication status
Not yet conducted. Pre-registered protocol.
Data and code
- Simulation framework: Extension of OBS-001 codebase
- Language: Python 3.11+, JAX/Flax for recurrent networks
- Pre-registration: This document. Timestamp: 2026-07-26.
- Repository: [To be created upon study initiation]
Relationship to philosophical archive
This study operationalizes the archive's discussion of "false self-models" and "mistaken identity." The archive suggests such states are possible and meaningful; this study tests whether they can arise and persist in a minimal computational system. Results do not validate or invalidate the philosophical claims—they test a specific, operationalized sub-claim.
References
- Friston, K., et al. (2017). Active inference, curiosity and insight. *Neural Computation*, 29(10).
- Metzinger, T. (2003). Being No One. MIT Press.
- Hohwy, J. (2013). The Predictive Mind. Oxford University Press.
- Kording, K. P., & Wolpert, D. M. (2004). Bayesian integration in sensorimotor learning. *Nature*, 427(6971).
- Schwartenbeck, P., et al. (2015). Evidence for surprise minimization over value maximization in choice behavior. *Scientific Reports*, 5.
Revision history
| Date | Version | Changes |
|------|---------|---------|
| 2026-07-26 | 1.0 | Initial pre-registration |