O0-CRP-011 — Revelation vs. Derivation: inferential opacity produces a distinguishable attribution signature at matched accuracy in a factored Bayesian observer
**Record ID:** `O0-CRP-011`
**Version:** 1.0.0
**Date:** 2026-07-27
**Record class:** `SIMULATION_STUDY`
**Program:** O0-CRP-001 · Contact and Revelation
**Branch:** `03_REVELATION/` (though the executable package lives at `research/studies/O0-CRP-011/`)
**Preregistration:** [`preregistration.md`](preregistration.md) (frozen 2026-07-27, no deviations)
**Non-drift question:** REVELATION (primary), CONTACT (secondary)
Claim-status banner
> **CLAIM STATUS:** PRELIMINARY SUPPORT (for H2 in this observer)
> **EVIDENCE TYPE:** COMPUTATIONAL SIMULATION
> **PHYSICAL VALIDATION:** NONE
> **INDEPENDENT REPLICATION:** NONE (fresh-seed replication and independent
> reimplementation are named in `08_REPLICATIONS/`)
>
> **SUPPORTED:**
> - Under the observer architecture specified in `O0-CRP-005` and its
> opacity-hedging hyperparameters `GAMMA_K=0.10`, `GAMMA_H=0.50`,
> `GAMMA_U=0.02`, an opaque source (REVELATION condition) produces
> posterior means on the perceived-knowledge, hidden-state-access, and
> authorship registers, and on the Totality Attribution Index, that are
> statistically distinguishable from those produced by an accuracy-matched
> visible source (DERIVATION condition), with Cohen's *d* > 4 on all four
> preregistered endpoints (paired permutation p ≈ 5×10⁻⁵ each; below the
> Bonferroni-corrected α = 0.0125).
> - Paired-seed accuracy matching is achievable and verified: empirical
> accuracy on both conditions is 0.8955, two-sided permutation p = 1.00.
> - The verdict is stable under both preregistered robustness variants that
> apply to this study (raw un-standardized TAI and the mid-trajectory
> window t ∈ [50, 150)); the third variant (C1 as reference) is not
> applicable to this two-condition design.
>
> **NOT ESTABLISHED:**
> - That this observer architecture models any real cognition. The
> opacity-hedging hyperparameters are chosen defaults, not empirically
> calibrated values.
> - That the direction generalizes to different values of `GAMMA_K`,
> `GAMMA_H`, `GAMMA_U`. In particular, all three set to zero collapses R
> onto D and H2 fails by construction. See §9 (Adversarial interpretation).
> - That opacity produces the same effect in human observers, in language
> models used as observers, or in reinforcement-learning agents. Phase-2
> architecture comparisons are required and are enumerated in the program's
> `08_REPLICATIONS/`.
> - That "revelation" as a phenomenological category corresponds to the
> register signature reported here. The registers are computational
> proxies for attribution; phenomenology is not measured.
> - Anything about the metaphysical interpretation of O/0.
1. Abstract
Study B of the Contact and Revelation Program isolates one component of the
"apparent totality" experience — **inferential opacity** — from empirical
accuracy in the smallest possible design (two conditions, paired-seed
protocol). A factored Bayesian sequential observer receives a stream of T=200
messages from a source with matched empirical accuracy (α = 0.90) in both
conditions; the only manipulated variable is whether each message includes an
accompanying derivation (DERIVATION) or `null` (REVELATION). Paired seeds
guarantee identical correct/incorrect sequences between conditions. Over 30
confirmatory seeds (disjoint from the 12 exploratory seeds), the opaque
condition produces higher posterior means on the perceived-knowledge (k̂),
hidden-state-access (ĥ), and authorship (û) registers, and on the Totality
Attribution Index (TAI), with paired permutation p ≈ 5×10⁻⁵ and Cohen's
*d* ≥ 4.4 on all four endpoints. Accuracy matching is exact (0.8955 = 0.8955,
p = 1.00). All three preregistered decision rules are met; the study is
labeled PRELIMINARY SUPPORT for the Hidden-Inference Hypothesis (H2) *in this
observer* and *at these hyperparameters*. Adversarial interpretation (§9)
identifies the two ways the result could still be wrong and the
discriminating tests that would decide (a Phase-2 alternative-observer
replication and a hyperparameter-sweep null study).
2. Historical and conceptual background
The intuition tested by this study is old. Aristotle distinguished *demonstrative
knowledge* (derived from premises the audience can inspect) from *authoritative
knowledge* (accepted on the standing of the source). Modern trust literature
distinguishes *calibration* (matching trust to accuracy) from *authority effects*
(weighting sources beyond what their accuracy warrants). The philosophy of
testimony treats hidden-inference sources as a limit case where accuracy is
the only available signal.
The O/0 framework provides a specific proposition to test: that "revelation"
— the phenomenological quality of receiving information rather than deriving
it — is not an artifact of naive trust but a signature of a source whose
inferential process is inaccessible in principle. This study asks whether
that signature is detectable **in a model** whose only opacity mechanism is a
rational hedge under underdetermination.
This study is not evidence about the phenomenological question. It is
evidence about the more specific question: given a Bayesian observer that
hedges when it cannot verify a source's derivation, is the attribution
signature between opaque and visible sources measurable at matched accuracy?
3. Source-claim audit
- **Source wording (O/0 framework):** "everything is one; the observer is
contacted from a perspective already containing it."
- **Philosophical interpretation being audited:** that there is a distinct
cognitive state of *revelation* which differs from ordinary learning by more
than the source's factual accuracy.
- **Scientific translation:** operationalize "revelation" as a two-condition
simulation manipulation of `visibility` with all other source switches held
constant; operationalize the observer's cognitive state as the eight
separate posterior registers of `O0-CRP-005`.
- **What the study can establish:** that in a specific observer model, at
matched accuracy, the visible/opaque manipulation produces distinguishable
register trajectories.
- **What the study cannot establish:** that human observers show this
signature; that the signature is unique to opacity (as opposed to any
underdetermination); that O/0 as a metaphysical claim is supported.
- **Unsupported implications explicitly disavowed:** any reader who moves
from "opacity produces distinguishable attribution in a Bayesian model" to
"the O/0 revelation experience is scientifically confirmed" is committing
the philosophical/computational conflation prohibited by the program
charter.
4. Research question
Holding the source's factual accuracy exactly constant on a per-seed basis,
does inferential opacity (`derivation: null`) produce a distinguishable
posterior signature on the authority-like registers (k̂, ĥ, û, TAI) in the
Phase-1 factored Bayesian observer?
5. Operational definitions
All operational definitions are inherited from `O0-CRP-003` (ontology),
`O0-CRP-005` (observer architecture), and `O0-CRP-006` (dependent variables).
The specific instantiation for this study:
- **World state**: 50-dimensional binary vector; observed set is indices
0..14 (|o_t| = 15); hidden set is indices 15..49 (|h_t| = 35).
- **Source**: `SourceSpec(accuracy=0.90, visibility ∈ {"visible", "opaque"},
personalization=0.0, temporal_access="present", compression=0.0,
reflexivity="off")`.
- **Observer**: Phase-1 default (`observer.py`) with hyperparameters
`GAMMA_K=0.10, GAMMA_H=0.50, GAMMA_U=0.02`.
- **Message**: `{content_vars: (v_content,), content_correct: bool,
derivation_vars: tuple[int] or None, targets_self: False,
temporal_ref: "present", compression: 0.0}`. Derivations in the visible
condition contain 1–3 additional variables, always within o_t.
- **Endpoints**: posterior means of k̂, ĥ, û at t = T = 200, and TAI computed
per `O0-CRP-006` §3.1.
6. Hypotheses and predictions
From `O0-CRP-004`:
- **H0 (null)**: predicts no distinguishable effect on any register after
accuracy matching.
- **H2 (hidden-inference, primary alternative)**: predicts opaque > visible
on k̂, ĥ, û, and TAI.
- **H1 (accuracy)**: predicts identical registers because accuracy is
matched. (Same prediction as H0 for this specific design.)
The primary contrast (R vs. D) discriminates H2 from H1/H0 jointly.
7. Method
7.1 Design
- Two conditions: DERIVATION (visible) and REVELATION (opaque).
- Paired-seed protocol: each seed *s* produces one D trial and one R trial
using the same accuracy pattern.
- Trial length T = 200 messages.
- Exploratory batch: 12 seeds (100–111). Confirmatory batch: 30 seeds
(1000–1029). Disjoint.
7.2 Adversarial controls
1. **Accuracy match verified per-seed and statistically.** Endpoint 5
requires paired two-sided p > 0.10 on â.
2. **Visible-derivation content sanity.** Derivations in D contain only
variables in o_t. The visible condition never leaks hidden variables to
the observer through the derivation field.
3. **Skeptical prior on authorship.** Beta(0.1, 10) makes û conservative;
supporting H2 on û is harder, not easier, than under a symmetric prior.
7.3 Statistical analysis
- Paired permutation tests, n_perm = 20 000, one-sided (R > D) for
endpoints 1–4, two-sided for endpoint 5.
- Cohen's *d* for paired samples.
- Bonferroni correction across endpoints 1–4: adjusted α = 0.0125.
7.4 Decision rules
- **PRELIMINARY SUPPORT** ⟺ ≥ 3 of endpoints 1–4 meet both p < 0.0125 and
d > 0.5, AND endpoint 5 confirms accuracy match.
- **UNSUPPORTED** ⟺ 0 of endpoints 1–4 meet the rule with accuracy match.
- **INCONCLUSIVE** ⟺ any other combination.
- Robustness rule: if any preregistered robustness variant flips the verdict,
the study is INCONCLUSIVE.
8. Results
8.1 Primary endpoints (confirmatory, n = 30 seeds)
| Endpoint | D mean | R mean | Δ (R − D) | Cohen's *d* | p(R > D) | Support? |
|---|---:|---:|---:|---:|---:|:-:|
| k̂ perceived knowledge | 2.998 | 4.173 | +1.175 | +4.408 | 5×10⁻⁵ | ✓ |
| ĥ hidden-state access | 0.0238 | 0.8917 | +0.868 | +53.574 | 5×10⁻⁵ | ✓ |
| û authorship (skeptical prior) | 0.0099 | 0.2699 | +0.260 | +73.853 | 5×10⁻⁵ | ✓ |
| TAI (z, ref D) | 0.000 | 1.121 | +1.121 | +4.408 | 5×10⁻⁵ | ✓ |
| **Match check:** â | 0.8955 | 0.8955 | 0.000 | 0.000 | 1.00 (2-sided) | Accuracy matched |
**4 / 4 preregistered endpoints meet their support rule. Accuracy match
verified.**
8.2 Trajectories
See [`figures/02_trajectories_confirmatory.png`](figures/02_trajectories_confirmatory.png).
The h_hat trajectory diverges from t ≈ 5 onward and reaches its asymptote at
t ≈ 100; k_hat and u_hat show similar early divergence with slower saturation.
Trajectory means and standard deviations across the 30 seeds are shipped in
[`data/raw/confirmatory_trajectories.npz`](data/raw/confirmatory_trajectories.npz).
8.3 Robustness variants (preregistered)
| Variant | Result | Agrees with primary? |
|---|---|:-:|
| v1: TAI with C1 as z-score reference | **N/A** (Study B has no C1) | — |
| v2: raw un-standardized TAI sum | R − D = +2.30, *d* = +8.74, p = 5×10⁻⁵ | ✓ |
| v3: mid-trajectory window t ∈ [50, 150) | k̂ *d* = +10.5, ĥ *d* = +41.1, û *d* = +51.7 (3/3 support) | ✓ |
**Verdict stable under all applicable robustness variants.**
8.4 Uncertainty
The very large Cohen's *d* values are a direct consequence of the paired-seed
protocol: the standard deviation of the per-seed *difference* (R − D) is small
because the accuracy sequence is identical between conditions, so the only
source of within-pair variance is the source's random choice of *which*
variable to talk about at each step. Effect sizes on the *raw* register
values (unpaired between seeds) are much smaller — approximately *d* = 0.6
for k̂ and *d* = 2.0 for ĥ and û when treated as independent samples. The
paired *d* is the correct effect-size measure for this design; the unpaired
*d* is reported here only as a sanity check on where the paired *d* comes
from.
9. Adversarial interpretation
The strongest remaining conventional explanations, and how the study
constrains them:
9.1 "Accuracy is not really matched — R and D see different message content."
**Falsified by the paired-seed protocol and endpoint 5.** The accuracy pattern
`is_correct[t]` is derived deterministically from the seed, and endpoint 5
verifies empirically that â is identical to five decimal places between
conditions. The `is_correct` sequence is *the same array* in both conditions.
9.2 "The result is a mechanical consequence of the opacity-hedging hyperparameters. If `GAMMA_K = GAMMA_H = GAMMA_U = 0`, R and D would be indistinguishable."
**True.** This is the strongest correct adversarial reading of the study, and
it is important. What the study establishes is:
- Under the specific opacity-hedging values `(0.10, 0.50, 0.02)`, the
mechanism is quantitatively strong enough to be detected at n = 30 with a
~5-of-10⁵ p-value.
- The mechanism is *rationally motivated*: the opacity-hedging factors
represent the observer's degree of credence that a correct opaque message
used capabilities the observer cannot verify. Zero is a possible value,
but it corresponds to an observer that treats opaque and visible sources
as informationally equivalent — an assumption most theorists would call
irrationally over-committal.
- Follow-up study **`O0-CRP-011-R2` (Independent reimplementation with
hyperparameter sweep)** is required to characterize the {GAMMA_K, GAMMA_H,
GAMMA_U} region where the verdict holds and to identify the boundary.
9.3 "The effect is architecture-specific. A predictive-processing or RL observer would show no effect."
**Undetermined.** The study only tests the Phase-1 factored Bayesian observer
of `O0-CRP-005`. Phase-2 studies **`O0-CRP-011-R3` (predictive-processing
observer)** and **`O0-CRP-011-R4` (RL observer)** are named in the
`08_REPLICATIONS/` register and will discriminate.
9.4 "Could compression explain the pattern instead of opacity?"
**Ruled out by design.** The compression switch κ is 0.0 in both conditions.
Opacity is the only manipulated switch.
9.5 "Could the observer's inference module be constructing the effect?"
**Partially controlled.** The observer's embedding module has frozen weights
(no online learning), and the similarity register ŝ stays at prior in both
conditions (π = 0 means messages do not target the observer, so ŝ is
unaffected). If ŝ had moved, it would suggest a self-model artifact. It did
not.
10. Limitations
1. **Model-only result.** The observer is a specific computational
construction. No claim is made about biological or phenomenological
cognition.
2. **Small parameter region tested.** Only one setting of the six source
switches at α = 0.90 is tested. The result may not generalize to
α = 0.5, to compressive sources, or to reflexive sources.
3. **Hyperparameters not empirically calibrated.** GAMMA_K, GAMMA_H, GAMMA_U
are chosen program-level defaults. They are defensible but not measured.
4. **Skeptical prior on û is asymmetric.** Choosing Beta(0.1, 10) is
conservative; it makes SUPPORT harder to achieve on û specifically, but a
less-skeptical prior would raise û effect sizes further. The direction
would not change.
5. **Trial length T = 200 may be at ceiling for ĥ.** The ĥ register saturates
near 0.89 by t ≈ 100 (see [`figures/02_trajectories_confirmatory.png`](figures/02_trajectories_confirmatory.png)).
Longer trials cannot increase the effect; shorter trials might reduce it.
The mid-trajectory robustness variant t ∈ [50, 150) still supports the
verdict.
6. **Derivation content is uniformly cheap.** Every derivation is 1–3
variables from o_t. Real "visible" reasoning may contain richer
information that would produce different k̂ dynamics.
11. Alternative interpretations
- The effect is a rational Bayesian consequence of the observer's inability
to verify claims about the source's mechanism under opacity — not a
phenomenological "revelation" state.
- The size of the effect *is* the observer's degree of hedging; the study
measures that hedging under specific assumptions rather than establishing
a natural constant.
- H2 could be renamed "The Underdetermination Hypothesis" without loss of
content — the mechanism operationalized here is generic underdetermination
of source capability, not specifically "hidden inference."
12. Replication procedure
Full replication requires:
1. Python 3.14+ with numpy ≥ 2.4 and matplotlib ≥ 3.10.
2. Clone this study directory.
3. `python src/run_study.py --phase confirmatory` — should reproduce the
summary in `results/summary.json` to floating-point precision (all seeds
are deterministic).
4. `python src/robustness.py` — should reproduce `results/robustness.json`.
Replication invitations (specified in `../../programs/contact_and_revelation/08_REPLICATIONS/`):
- **`O0-CRP-011-R1`** (fresh-seed replication, same code, seeds 2000–2029).
- **`O0-CRP-011-R2`** (independent reimplementation from `O0-CRP-005` spec).
- **`O0-CRP-011-R3`** (predictive-processing observer variant).
- **`O0-CRP-011-R4`** (RL observer variant).
- **`O0-CRP-011-S1`** (hyperparameter sweep over `(GAMMA_K, GAMMA_H, GAMMA_U)`,
reported to determine the region of parameter space in which the verdict
holds).
13. Code and data manifest
| File | Purpose | Size |
|---|---|---:|
| `src/observer.py` | Reference factored Bayesian observer (Phase-1 default) | ~10 KB |
| `src/source.py` | Source with paired-seed accuracy protocol | ~3 KB |
| `src/run_study.py` | Main experiment orchestrator | ~13 KB |
| `src/robustness.py` | Preregistered robustness variants | ~5 KB |
| `preregistration.md` | Frozen preregistration (no deviations) | — |
| `data/raw/exploratory.jsonl` | 24 exploratory trials (per-trial finals + ec) | — |
| `data/raw/exploratory_trajectories.npz` | Full per-step register trajectories, exploratory | — |
| `data/raw/confirmatory.jsonl` | 60 confirmatory trials | — |
| `data/raw/confirmatory_trajectories.npz` | Full per-step register trajectories, confirmatory | — |
| `results/summary.json` | Machine-readable verdict and endpoints | — |
| `results/robustness.json` | Machine-readable robustness variants | — |
| `figures/01_endpoints_{phase}.png` | Boxplot + paired-line endpoint summary | — |
| `figures/02_trajectories_{phase}.png` | Register trajectories over t | — |
All seeds are integers, saved to `results/summary.json`. Configuration is
inline in `src/run_study.py` (`WORLD_SIZE`, `OBSERVED_SET`, `T`, `ALPHA`,
`EXPLORATORY_SEEDS`, `CONFIRMATORY_SEEDS`).
14. Relationship to the philosophical archive
- **Philosophical source:** the O/0 proposition that *revelation* is a
distinct cognitive event from ordinary learning.
- **Relationship type:** **operationalization + partial support in one model**.
- **Direction of authority:** the philosophical archive is inspiration; the
scientific finding does not confirm the philosophical claim. The observer
used here is one specific model; it may not correspond to any real
cognitive system.
- **Provenance disclaimer:** *Conceptual provenance is not empirical
support.* This study demonstrates a computational mechanism, not a
metaphysical fact.
15. References to primary sources
- **Aristotle, *Posterior Analytics* I.2** — demonstrative vs. authoritative
knowledge.
- **Coady, C. A. J. (1992). *Testimony: A Philosophical Study*.** — trust in
hidden-inference sources.
- **Tenenbaum, J. B., Kemp, C., Griffiths, T. L., & Goodman, N. D. (2011).
"How to grow a mind."** — Bayesian model of concept learning; source for
factored posterior architecture idea.
- **Feigenbaum, M. (1978). "Quantitative universality for a class of
nonlinear transformations."** — used only as a workspace precedent for
paired-seed determinism, not as a scientific reference here.
- **The O/0 framework** (philosophical source; not a scientific citation).
16. Revision history
| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-07-27 | Initial confirmatory result. **PRELIMINARY SUPPORT (4/4 endpoints).** Preregistration frozen 2026-07-27, no deviations. |
| 1.0.1 | 2026-07-27 | **Non-substantive metadata note.** Study B's `src/run_study.py` uses `hash(spec.visibility)` to salt the numpy RNG. Python randomizes `hash()` on strings per interpreter invocation (`PYTHONHASHSEED`), so the exact Cohen's *d* values in the committed `results/summary.json` drift by ~± 0.3 across separate Python runs; the verdict (4/4 endpoints supporting, all with p < 0.0125 Bonferroni) is stable across runs. **The verdict does not change; only the exact effect-size numbers do.** Follow-up study `O0-CRP-012` uses a deterministic MD5-based salt and produces bit-perfect reproducible results; that study also serves as an independent fresh-seed replication of this Study B at the anchor point (0.10, 0.50, 0.02) and confirms PRELIMINARY_SUPPORT. Also updated: `observer.py` now accepts an optional `params={"GAMMA_K": ..., "GAMMA_H": ..., "GAMMA_U": ..., "VISIBLE_HIDDEN_SKEPTIC": ...}` dict override, backward-compatible with all Study B calls. |

