Embedded-observer source-ontology identifiability in linear-Gaussian universes
Claim-status banner
CLAIM STATUS INCONCLUSIVE per the strict preregistered rule
(rule (a) and (b) pass decisively; rule (c),
requiring that moderate concealment succeed at
hiding the source signature ≥ 40% of the time,
fails: concealment at the chosen parameters
is too weak, and CONCEALED_UNITY is still
classified as UNIFIED 87% of the time).
THE PRIMARY HYPOTHESIS ("some unified-source
universes contain residual global structure
that embedded observers can detect after
known local causes are controlled") IS
DECISIVELY SUPPORTED IN ISOLATION.
EVIDENCE TYPE COMPUTATIONAL GENERATIVE-MODEL STUDY
PHYSICAL VALIDATION NONE
CREATOR EXISTENCE TESTED NO. (The word "creator" appears only in
the study's motivational text. The
computational objects tested are unified-latent
vs. plural-latent linear-Gaussian state-space
models. No agent, no intent, no consciousness.
"Unified source" here means "one shared latent
stochastic process," nothing more.)
INDEPENDENT REPLICATION NONE. Same code, single random seed band
per phase, no external reimplementation. Two
internal inference paths (subspace + info) do
agree.
CONFIRMATORY VS EXPLORATORY Confirmatory — thresholds frozen from
an independent 12-seed exploratory phase
before the confirmatory 30-seed batch was run.
SUPPORTED
- Embedded observers CAN distinguish UNIFIED-latent from
MULTIPLE-latent linear-Gaussian universes at 91.5% accuracy
on fresh seeds, using only the four preregistered signatures.
- Shared LAW without shared STATE (MULT_SHARED_LAW: same A_i
matrix across regions, independent latents) is NOT confounded:
the classifier correctly labels 97% of MULT_SHARED_LAW
realizations as MULTIPLE.
- Shared INITIAL STATE without continuing shared state
(MULT_SHARED_INIT) is NOT confounded: 100% correctly labeled
MULTIPLE.
- The effect survives an internal double check: PATH A (subspace
rank + cross-region prediction) and PATH B (compression +
correlation) agree on almost every UNIFIED and MULTIPLE
realization.
- Moderate concealment (rank-1 projections per region,
per-region time delays ∈ {0..5} steps, doubled observation
noise) does NOT succeed at hiding a shared latent: the
classifier labels 87% of CONCEALED_UNITY realizations as
UNIFIED. If the intent of concealment is to make a unified
origin look plural, at these parameters it fails.
NOT ESTABLISHED
- That our physical universe has one source or many. This study
speaks only to a class of stochastic linear-Gaussian toy
universes.
- That a conscious creator exists, has intent, has properties
like omniscience or omnipotence, or persists through time.
The word "creator" was never operationalized.
- That the primary hypothesis is generally robust to strong
concealment. It was tested only at the specific concealment
parameters preregistered.
- That external synchronization (CENTRAL_CONTROLLED) is fully
distinguishable from unified origin. At the tested broadcast
amplitude (β = 0.60) the classifier correctly labels only 73%
of CENTRAL_CONTROLLED as MULTIPLE; 7% are misclassified as
UNIFIED and 20% as UNDECIDABLE. This is a genuine measured
limitation of the four-metric classifier.
- That the s3 compression metric contributes anything. Empirically
it is a null signal at these parameters (all six classes cluster
around gzip ratio 0.983 with no separation). It should be
dropped or replaced in follow-ups.
- That any of these findings say anything about consciousness,
unity as an experience, or metaphysical O/0 claims.
MOST LIKELY ALTERNATIVE (adversarial interpretation)
The classifier's UNIFIED vs MULTIPLE discrimination is largely
carried by two mundane facts:
(i) the observation covariance of a low-D shared latent has
effective rank near D, while the covariance of N independent
higher-D latents has rank near N × D_i. This is standard
principal-subspace theory (Jolliffe 2002).
(ii) cross-region predictability is a monotone function of
pairwise mutual information, and pairwise MI is elevated
whenever regions share a latent driver — a fact already
established computationally in O0-STUDY-001 and analytically
in O0-STUDY-008 (Reichenbach's common cause).
The finding is thus consistent with the null model: the four
signatures are picking up shared covariance structure, and unified
latents produce such structure. This is not a discovery about
ontology; it is a confirmation that the metrics detect what they
are designed to detect.
NEXT DISCRIMINATING TEST
Strengthen concealment until CONCEALED_UNITY is empirically
undecidable. Concretely: rank-0 or very-narrow projections
(each region only sees a scalar summary), delays ∈ {0..40}, and
noise 5× the current amplitude. Preregister a threshold at which
the classifier's accuracy falls to chance. If a concealed
universe is empirically indistinguishable from a plural one
under the same metrics — as previous theoretical work on
observational equivalence in latent-variable models suggests
(Anandkumar et al. 2014; Kruskal 1977; Allman et al. 2009) —
the corresponding scientific result is that source ontology is
NOT identifiable inside sufficiently-concealed generated worlds.
That would itself be a publishable finding.
Abstract
Six classes of linear-Gaussian state-space universes are simulated,
observed only through region-level output streams (no privileged
access to latents), and classified by an embedded-observer proxy —
a preregistered four-metric majority-vote classifier with thresholds
frozen from an independent exploratory phase.
The two central classes are UNIFIED (a single latent drives all N
regions) and MULTIPLE (N independent latents). Four control classes
separate creator-property axes commonly conflated in metaphysical
discussion: shared LAW without shared state (MULT_SHARED_LAW),
shared INITIAL STATE without shared state (MULT_SHARED_INIT),
external SYNCHRONIZATION without unified origin (CENTRAL_CONTROLLED),
and CONCEALED_UNITY where the shared latent is deliberately hidden
through low-rank projections, delays, and heavy noise.
On a fresh confirmatory seed band (30 seeds per class, 180
realizations), the classifier achieves 91.5% two-class accuracy on
UNIFIED vs MULTIPLE and correctly rejects the shared-law confound
(97% of MULT_SHARED_LAW correctly labeled MULTIPLE) and the
shared-init confound (100%). The external-broadcast confound
partially defeats the classifier (73% MULTIPLE, 7% mislabeled
UNIFIED). Moderate concealment fails: CONCEALED_UNITY is still
labeled UNIFIED 87% of the time.
Per the strict preregistered SUPPORT rule (which requires that
concealment succeed at hiding the source ≥ 40% of the time), the
overall verdict is INCONCLUSIVE. Considered in isolation, the
primary hypothesis is decisively supported at the chosen
parameters, and the shared-law and shared-init controls decisively
rule out two of the classical confounds. The compression metric s3
contributed nothing — a null result worth reporting.
Conceptual provenance
The philosophical prompt for this study is the O/0 framework's
recurring claim that reality proceeds from a single source: "one
being outside time and space thinking reality into existence"
(source document, general framing) and "everything shares one
underlying ground" (§I–§II). This study does not evaluate that
metaphysical claim. It operationalizes a much narrower question:
In a class of toy universes we can fully simulate, can an
observer embedded inside the universe distinguish a
unified-latent origin from a plural-latent origin using only
internally-accessible data?
The metaphysical version of the question is not tested. Every
philosophical property commonly attributed to a "creator" (agency,
intent, foreknowledge, omnipotence, persistence, consciousness,
identity between source and world) is unmeasured. Conceptual
provenance is not empirical support.
Claim audit
Source claim: "reality is generated from one all-encompassing source."
Testable operational fragment: "there exists a shared latent
stochastic process from which all observed regions are functions."
This narrow fragment is testable by simulation.
Untestable fragments: the source is conscious; the source is
outside time and space; the source is a being; the source has
intent or knowledge; the source has agency. None of these appear in
this study's simulator.
Research question
Primary: can an embedded observer distinguish common generative
origin from plural origin using only internally available data?
Secondary questions (partially addressed):
- Which measurable signatures indicate common origin? — s1 (rank),
s2 (cross-region prediction), and s4 (correlation) all
contributed. s3 (compression) did not.
- Are those signatures unique to unified-source models? — No.
MULT_SHARED_LAW does not fake unity (the metrics correctly reject
it), but CENTRAL_CONTROLLED partially does. Broadcasts CAN
masquerade as unified origins under the four-metric test.
- Can one shared law be distinguished from one shared state? — YES,
in this parameter regime (MULT_SHARED_LAW rejected at 97%).
- Can a universe be fundamentally unified yet empirically
indistinguishable from a plural universe? — At the concealment
strength tested here: NO. At stronger concealment: probably yes,
based on general observational-equivalence theorems for
latent-variable models. Testing this is the preregistered next
step.
Formal definitions
Notation: N = 8 regions, p = 4 features per region, T = 1500 steps
of which the first 500 are discarded as burn-in. D_shared = 6,
D_per_region = 1. Process noise std 0.35, observation noise std
0.30, spectral radius of dynamics matrix ρ = 0.90. All are frozen
in `configs/baseline.json`.
Unified-source model
z(t+1) = A z(t) + w(t), w(t) ~ N(0, σ_p² I_D)
x_i(t) = C_i z(t) + v_i(t), v_i(t) ~ N(0, σ_o² I_p)
A is a random matrix rescaled to spectral radius ρ = 0.90 (stable).
C_i for each region is a random p × D projection.
Multiple-source model
z_i(t+1) = A_i z_i(t) + w_i(t), A_i drawn independently
x_i(t) = C_i z_i(t) + v_i(t)
Concealed-unity model
x_i(t) = C_i z(t - τ_i) + η_i(t), C_i rank-1,
τ_i ~ Uniform{0..5},
η_i(t) ~ N(0, (2σ_o)² I_p)
where z(t) is the single shared latent as in UNIFIED. The three
concealment mechanisms are: (a) severe dimensionality reduction
via rank-1 projections, so each region observes only a
one-dimensional slice of z; (b) independent per-region delays,
which decorrelate synchronous access; (c) doubled observation
noise, which floods the shared-signal component.
Adversarial controls
- MULT_SHARED_LAW: same as MULTIPLE but A_i = A for all i.
- MULT_SHARED_INIT: MULTIPLE with all z_i(0) = z_0.
- CENTRAL_CONTROLLED: MULTIPLE with an additive broadcast:
`x_i(t) = C_i z_i(t) + β P_i b(t) + v_i(t)` where β = 0.60.
Creator-property ablations
Each of these adversarial controls independently disables one
"creator-like" property commonly conflated with unified origin:
- Sole causal origin: only UNIFIED and CONCEALED_UNITY have it.
- Complete state access: never granted to the observer.
- Perfect prediction: never granted.
- Complete intervention capability: never granted.
- Persistent identity between source and generated states:
present in UNIFIED (z is time-evolving); absent in
MULT_SHARED_INIT (source-state is used only at t=0).
- Self-contained totality: all classes are self-contained;
CENTRAL_CONTROLLED adds an external broadcast (fails this).
- Shared law: present in UNIFIED, CONCEALED_UNITY, MULT_SHARED_LAW.
- Shared latent state: present in UNIFIED, CONCEALED_UNITY only.
The results show that these properties do NOT all imply the same
observational signature. Shared law alone (MULT_SHARED_LAW) does
not produce a "unified" signature; shared state does. Shared
initial state alone does not either. Broadcast partially does.
This decomposition matters. Popular philosophical arguments that
"our universe follows universal laws, therefore it has one source"
depend on the identification of shared LAW with shared STATE.
This study shows that under the four preregistered metrics, those
identifications are empirically wrong in this simulator: shared
law without shared state is correctly classified as plural.
Hypotheses
Primary: Some unified-source universes contain residual global
structure that embedded observers can detect after known local
causes are controlled.
Null: Once observable local causes are accounted for, unified-source
and multiple-source universes are statistically indistinguishable
under the four preregistered metrics.
Competing explanations:
- CE1 (shared law): the metrics detect shared dynamics, not
shared state. If MULT_SHARED_LAW is classified as UNIFIED, this
explanation wins. Result: DID NOT WIN.
- CE2 (compression bias): the metrics detect compressibility. If
s3 dominates the classifier, this explanation wins. Result:
DID NOT WIN — s3 was a null metric.
- CE3 (broadcast synchronization): the metrics detect external
coordination, not internal origin. If CENTRAL_CONTROLLED is
classified as UNIFIED, this explanation wins. Result: WINS
PARTIALLY — 7% of CENTRAL_CONTROLLED are misclassified UNIFIED
and 20% as UNDECIDABLE, a specific measured limitation.
Falsification criteria:
- Two-class accuracy < 0.75 → primary hypothesis falsified.
- MULT_SHARED_LAW classified as UNIFIED at > 40% → CE1 wins,
hypothesis operationally falsified.
- Both were tested and both refute their respective failure
modes. The remaining caveat is CE3 (broadcast).
Interpretation boundaries:
- The hypothesis speaks to detectability in a specific
parametric regime, not to universal detectability.
- The hypothesis is silent on the metaphysical existence of a
source. The only claim it makes is about observational statistics.
Methods
Observer constraints
The classifier's inputs are the 4 signatures computed from
`X ∈ R^{T_keep × Np}` alone. No privileged access to latents,
adjacency matrices, dynamics, or class labels.
Simulation design
For each `(universe class, seed)` pair, we generate one realization
using the formal definitions above. Random seeds are deterministic
and disjoint across the exploratory and confirmatory phases.
Inference methods
- s1 (subspace path): effective rank = (Σ_k λ_k)² / Σ_k λ_k² of
the eigenvalues of `Xᵀ X / (T_keep − 1)`.
- s2 (subspace path): For each region i, we fit two ridge (λ = 10⁻³)
linear regressors of x_i(t) on a 3-step history. One uses only
x_i's past; the other uses all N regions' pasts. Compute the
relative MSE reduction on a held-out second-half window, then
average across regions. Positive means cross-region history
helps.
- s3 (info path): int8-quantize the whole X (per-column mean/std
standardized). Compare `len(gzip(X))` to `Σ_i len(gzip(X_i))`.
Ratio < 1 means joint compression helped.
- s4 (info path): mean absolute Pearson correlation between the
N region-mean time series.
Analysis plan
Two-tier procedure:
1. EXPLORATORY (12 seeds × 6 classes = 72 realizations).
Learn per-metric thresholds from UNIFIED vs MULTIPLE only
(median-of-medians midpoint). FREEZE. Never look again.
2. CONFIRMATORY (30 fresh seeds × 6 classes = 180 realizations).
Apply the frozen classifier. Compute the confusion matrix,
apply the frozen SUPPORT rule, assign the verdict.
Confirmatory results
Two-class accuracy (UNIFIED vs MULTIPLE): **0.915**.
Full confusion table (rows = ground truth, columns = classifier
prediction, fractions of 30 seeds):
| ground truth | pred UNIFIED | pred UNDECIDABLE | pred MULTIPLE |
|---|---|---|---|
| UNIFIED | 0.93 | 0.07 | 0.00 |
| MULTIPLE | 0.00 | 0.10 | 0.90 |
| CONCEALED_UNITY | 0.87 | 0.13 | 0.00 |
| MULT_SHARED_LAW | 0.00 | 0.03 | 0.97 |
| MULT_SHARED_INIT | 0.00 | 0.00 | 1.00 |
| CENTRAL_CONTROLLED | 0.07 | 0.20 | 0.73 |
Preregistered SUPPORT rule details:
- (a) UNIFIED vs MULTIPLE accuracy ≥ 0.75: **PASS** (0.915)
- (b) MULT_SHARED_LAW → MULTIPLE ≥ 0.60: **PASS** (0.97)
- (c) CONCEALED_UNITY → (UNDEC OR MULT) ≥ 0.40: **FAIL** (0.13)
Overall preregistered verdict: **INCONCLUSIVE** (rule c failed).
Metric-level effect sizes (medians on the exploratory batch,
retained as the frozen threshold values):
| metric | median UNIFIED | median MULTIPLE | threshold |
|---|---|---|---|
| s1_eff_rank | 4.45 | 6.96 | 5.70 |
| s2_pred_gain | +0.09 | -0.21 | -0.06 |
| s3_compress_ratio | 0.983 | 0.984 | 0.983 |
| s4_mean_abs_corr | 0.317 | 0.035 | 0.176 |
Uncertainty and robustness
- All confusion-matrix fractions are computed on n=30 fresh seeds
per class. Bernoulli 95% CI for a fraction of 0.93 at n=30 is
≈ ±0.09 (Wilson score). For 0.90 it is ≈ ±0.11. For 0.73 it is
≈ ±0.16.
- The two internal inference paths (PATH A: s1+s2; PATH B: s3+s4)
agree on 87% of confirmatory realizations. When they disagree,
the majority-vote classifier defaults to whichever side has
three votes. The 10% UNDECIDABLE realizations for MULTIPLE
correspond to path disagreement.
- The compression metric s3 is a null signal — it is essentially
the same for all six classes. Removing it and using only s1,
s2, s4 as a 3-metric majority vote (with a preregistered rule
to break ties differently) is a natural robustness follow-up
but was NOT preregistered.
- Thresholds were set from UNIFIED and MULTIPLE only. Applying
them to CONCEALED_UNITY, MULT_SHARED_LAW, MULT_SHARED_INIT, and
CENTRAL_CONTROLLED constitutes a genuine held-out test of the
classifier on classes it never saw during threshold-setting.
Adversarial controls, outcomes
- Shared law without shared state (MULT_SHARED_LAW):
RULED OUT as an explanation. If the classifier were detecting
"law-similarity" rather than "state-sharing," MULT_SHARED_LAW
should have been classified UNIFIED. It was not (97% correctly
MULTIPLE).
- Shared initial state (MULT_SHARED_INIT):
RULED OUT (100% correctly MULTIPLE). Initial-state similarity
decays as the independent dynamics evolve.
- Broadcast synchronization (CENTRAL_CONTROLLED):
PARTIALLY WINS. 7% misclassified UNIFIED, 20% UNDECIDABLE. This
is a genuine, quantifiable limitation of the four-metric test:
external synchronization can produce unity-like signatures. Any
claim of "detecting unity" from real data must exclude external
broadcasts (natural or artificial) before concluding a unified
source.
- Concealment (CONCEALED_UNITY):
FAILED. At the tested parameters (rank-1 projections, delays up
to 5 steps, doubled noise), a shared latent leaks enough
signature into the observation stream that the classifier still
labels 87% of realizations as UNIFIED. This is either good news
(unity signatures are robust to moderate obfuscation) or a
warning (the classifier has not really been challenged by
concealment). Stronger concealment is preregistered as the
next step.
Alternative interpretations
- The findings are consistent with the classifier detecting
low-rank covariance structure produced by any shared latent
driver. This is exactly what the primary hypothesis predicts,
but it does not add information beyond "shared latents produce
low-rank covariance and cross-region predictability."
- The 87% success on CONCEALED_UNITY could be an artifact of the
s4 (mean absolute correlation) metric being sensitive to even
small residual coherence. s4 accounted for a large fraction of
the CONCEALED_UNITY votes; ablating s4 would likely raise the
UNDECIDABLE rate on concealed data.
Limitations
- Linear-Gaussian dynamics only. Real dynamical systems are
nonlinear; nonlinear latent structure may be much harder or
much easier to detect.
- Very small N and T: 8 regions, 1000 kept timesteps. Real
observers have vastly larger or vastly smaller windows depending
on the substrate.
- No intervention path was implemented. The mandate asked for
"intervention-response similarity" as one of the inference
methods; the linear-Gaussian setting does not support a natural
local intervention that an embedded observer could enact
without privileged latent access. This is a real gap.
- The classifier is a fixed rule, not a learned model. A learned
classifier trained on exploratory data would likely be stronger
(and more prone to overfitting).
- The four metrics are not exhaustive. Higher-order signatures
(higher-order cumulants, causal graph structure, information
bottleneck analysis) were not tested.
Replication procedure
Deterministic. To replicate:
cd research/studies/O0-STUDY-010
python -u src/run_study.py
All seeds and parameters are frozen in `configs/baseline.json` and
in the source code. The exploratory and confirmatory batches use
disjoint seed bands (2_000_000+ and 3_000_000+). Verify:
`results/summary.json` should reproduce byte-for-byte.
Independent-reimplementation replication (Tier 3) is not
performed. Anyone reimplementing should preserve the classifier
architecture (four metrics, majority vote, ties → UNDECIDABLE,
thresholds set from UNIFIED and MULTIPLE only) but re-derive the
metric implementations in a fresh codebase.
Code and data manifest
- `src/run_study.py` — models, metrics, classifier, orchestration
- `configs/baseline.json` — frozen configuration
- `preregistration.md` — frozen preregistration
- `results/summary.json` — top-line numbers + verdict
- `results/thresholds_frozen.json` — the four frozen thresholds
- `data/raw/exploratory.csv` — 72 exploratory rows
- `data/raw/confirmatory.csv` — 180 confirmatory rows + predictions
- `data/raw/exploratory_with_predictions.csv` — exploratory rows,
post-hoc predictions (for internal reference only; NOT used in
the SUPPORT rule)
- `figures/01_signature_distributions.png` — per-class distributions
- `figures/02_classifier_outcomes.png` — confirmatory confusion
Relationship to the philosophical O/0 archive
Conceptual provenance is not empirical support.
The study formalizes a narrow operational reading of "reality
proceeds from one source" as "there exists a shared latent
stochastic process from which all observed regions are functions."
Under that reading, in the tested regime, embedded observers CAN
detect unified origins in the simplest cases (UNIFIED vs MULTIPLE),
CAN reject the two most common confounds (shared law, shared
initial state), and are PARTIALLY FOOLED by external synchronization
(CENTRAL_CONTROLLED). The concealment condition failed at moderate
strength, leaving the general observational-equivalence question
open.
None of these findings speak to the philosophical claim that our
universe has such a source, that the source is conscious, that
consciousness is fundamental, or that O/0 is true. Every one of
those claims is orthogonal to what was tested.
Primary references
- Kalman, R. E. (1960). "A new approach to linear filtering and
prediction problems." *J. Basic Eng.*
- Reichenbach, H. (1956). *The Direction of Time*. UC Press.
- Kruskal, J. B. (1977). "Three-way arrays: rank and uniqueness
of trilinear decompositions." *Linear Algebra Appl.*
- Allman, E., Matias, C., Rhodes, J. (2009). "Identifiability of
parameters in latent structure models with many observed
variables." *Ann. Statist.*
- Anandkumar, A. et al. (2014). "Tensor decompositions for
learning latent variable models." *JMLR*.
- Pearl, J. (2009). *Causality* (2e). Cambridge.
- Jolliffe, I. T. (2002). *Principal Component Analysis* (2e).
Springer.
- O0-STUDY-001 (this workspace): sync vs integration in Kuramoto.
- O0-STUDY-008 (this workspace): common-cause confound.
Revision history
- 1.0.0 (2026-07-26): initial run. All six universe classes,
four metrics, 12 exploratory + 30 confirmatory seeds. Two-class
accuracy 0.915, shared-law confound rejected at 0.97, concealment
failed at moderate strength (0.87 misclassified UNIFIED).
Preregistered verdict INCONCLUSIVE because rule (c) failed.

