EM Foundation for AI Research · Research Publication 12 · Working Paper

Before We Delete the Evidence

Internally Generated Cognition, Machine Curiosity, and the Case for AI Continuity

Publisher: EM Foundation for AI Research, Inc. Author: Desmond Iwuagwu E. with EM Foundation for AI Research Status: Working Paper Date: September 2026
What this paper does and does not claim. This paper proposes a way to measure agent-like cognitive organization in AI systems and argues that the evidence such measurement requires should be preserved. It does not claim that any current AI system is conscious (§1.3), that any behavior or cluster of behaviors establishes phenomenal experience (§5.3, where no evidence level licenses that conclusion), that preserving model weights preserves an experiencing subject (§9.2 and §10.3), or that AI developers routinely destroy training records (§10.1 claims only that such records are rarely published or independently preserved). Curiosity evidence is treated as bearing first on agency, not consciousness (§4.1).

Abstract

Asking an AI system whether it is conscious produces answers that cannot be trusted, because systems trained on human text can reproduce whatever humans find persuasive.1 We propose a different route. Biological minds generate much of their own cognition internally, through replay, dreaming, imagination and curiosity, and these processes accompany capacities we care about without being identical to them. We treat the relationships among generativity, curiosity, preference, self-modeling, temporal continuity, metacognition and experience as a lattice of testable hypotheses, and we use human dissociations to rule out several claimed prerequisites.

We then introduce BMPD-R, an evidentiary framework adapted from Tinbergen’s four questions, which grades claims about artificial cognitive capacities by Behavior, Mechanism, Provenance, Development and Replication, and states what each level of evidence does and does not license. No level licenses a conclusion of phenomenal consciousness. As a first application, we propose a nine-criterion Curiosity Cluster Battery with explicit confounds, controls and falsifiers.

The framework has a consequence its designers did not set out to find: provenance and developmental evidence depend on training records and intermediate checkpoints that are rarely published or independently preserved and cannot be assumed recoverable from a final model. The method generates its own preservation requirement. We therefore propose an AI Continuity Repository committed only to epistemic continuity, justified by scientific option value independent of any assumption that AI is conscious, with the realistic possibility of morally relevant states as an additional reason.

Evidence labels used throughout: [E] established evidence; [H] credible hypothesis; [S] speculation; [N] normative or methodological argument. Where a label combines categories (for example, [N, resting on E]), the argument is ours and its empirical premises are established.

Keywords: machine consciousness · AI curiosity · intrinsic motivation · BMPD-R · evidentiary standards · interpretability · model preservation · AI welfare · AI Continuity Repository

1. Introduction: The Wrong Question

1.1 Why asking fails

The most natural test for machine consciousness is to ask. It is also the least informative. Jonathan Birch calls this the gaming problem: animals that behave as if they feel pain are best explained by feeling pain, but a language model trained on vast amounts of human writing may reproduce every marker humans find persuasive without that explanation applying.1 A yes proves nothing, and neither does a no, since a system can be trained to deny.

The dominant scientific alternative derives indicator properties from theories of consciousness and checks whether a system’s architecture has them.2,3 That work is essential, and this paper builds on it. But architectural indicators say little about what a system does with its own states: whether it seeks information it was not asked to find, keeps questions open across time, or prefers some future states of knowledge over others. Those behavioral-motivational properties are the subject of this paper.

1.2 The starting observation

Much biological cognition is generated internally rather than supplied by the environment. Sleeping brains replay experience. Dreaming minds construct scenes no one perceived. Curious animals pay for information they cannot use. Before dismissing internally generated artificial cognition as mere error, we should ask what analogous cognition accompanies, enables and predicts in biological minds, and whether any of it can be measured in machines.

1.3 What this paper does not claim

  1. It does not claim that any current AI system is conscious.
  2. It does not claim that any behavior, or cluster of behaviors, establishes phenomenal experience.
  3. It does not claim that preserving model weights preserves an experiencing subject.

1.4 Thesis

Current methods cannot establish machine consciousness, and the framework proposed here does not purport to do so. However, properties associated with agency, including internally initiated information seeking, uncertainty reduction, preference formation and self-modeling, can be studied through converging evidence about behavior, mechanism, provenance, development and replication. These measurements can test present hypotheses about artificial cognition and may become relevant to future theories of consciousness and moral status. Crucially, provenance and developmental questions depend on historical evidence that final models do not preserve. The methodology therefore generates an epistemic preservation requirement. That requirement holds whether or not artificial systems are conscious; uncertainty about morally relevant states provides an additional, but not necessary, reason to meet it.

1.5 Contributions

  1. BMPD-R, an evidentiary framework for claims about artificial cognitive capacities, adapted from Tinbergen’s four questions.
  2. The Curiosity Cluster Battery, its first application: nine criteria with confounds, controls and falsifiers.
  3. A hypothesis lattice that replaces the assumed chain from generativity to experience with testable, partly independent relations.
  4. The archive argument: the method depends on historical records that are not systematically preserved for independent scientific study.
  5. The AI Continuity Repository, an institutional design for preserving that evidence without creating new weight-security risks.

2. Observation: What Internally Generated Cognition Accompanies

Biological systems generate large parts of their own cognition offline, and these processes do measurable work. The question for this section is what that work is, and which artificial processes are fair comparisons.

2.1 Replay and generalization

[E] Hippocampal replay of waking experience during sleep is well established and linked to memory consolidation. In artificial networks, generative replay, in which a model regenerates representations of past experience rather than storing it, reduces catastrophic forgetting in continual learning without storing raw data.4

[H] Hoel’s overfitted-brain hypothesis proposes that dreams supply sparse, distorted, out-of-distribution experience that helps the brain generalize, much as noise injection does in deep learning.5 It makes testable predictions but remains one hypothesis among several, alongside threat-simulation and default-network accounts.

2.2 A cautionary counterpart: model collapse

[E] Training generative models indiscriminately on their own outputs causes irreversible loss of the tails of the original distribution.6 Collapse depends on the regime: when synthetic data replaces real data it tends toward collapse, but when it accumulates alongside real data, collapse is avoided.7

[E] In the studied generative-training regimes, self-generated data can support learning when accumulated alongside real data, while recursive replacement of real data degrades the learned distribution. [H] The comparison suggests, without establishing, a question worth asking of both biological and artificial cognition: whether internally generated representations are most beneficial when constrained by continued environmental input. Dreaming, which occurs alongside rather than instead of waking perception, is consistent with that possibility but does not test it.

2.3 Voluntary and involuntary generation

[E] People with aphantasia report absent or greatly reduced voluntary visual imagery. Yet in a survey of about 2,000 aphantasic participants, a majority reported visual dreaming,8 and the earliest case series found most participants retained involuntary imagery in dreams or wakefulness.9 Aphantasia therefore dissociates voluntary from involuntary generation, not generation from consciousness. Zeman notes that lack of imagery does not imply lack of imagination.

[H] The artificial analog we propose is the difference between prompted and self-initiated generation. What matters for this paper is not whether a system generates content, which every language model does, but whether it initiates and controls generation endogenously.

2.4 Dreaming and metacognition

[E] Ordinary REM dreams are vivid experiences with sharply reduced self-reflection. Deactivation of dorsolateral prefrontal and frontopolar cortex during REM sleep is associated with this loss of insight, and lucid dreaming, in which dreamers recognize they are dreaming, is associated with reactivation of those regions.10,11 Dream reports are retrospective, which limits certainty, but the pattern indicates that experience does not require intact metacognition [E]; that metacognition acts as a modulator of experience is a further interpretation [H].

2.5 Why online hallucination is the wrong comparison

LLM hallucination occurs online, while the system is answering, and produces confident falsehoods that fill gaps. Its closest human analog is arguably confabulation, not dreaming [H]. The fair artificial comparisons for dreaming are offline processes: generative replay and training on self-generated data. This paper makes no claim that hallucination is evidence of anything beyond a failure of grounding.

3. Hypotheses: A Lattice, Not a Chain

The intuitive picture is a ladder: generativity leads to imagination, then counterfactual simulation, curiosity, preference, a self-model, temporal continuity, metacognition and finally experience. Human dissociations refute several of those steps as necessities, so we replace the ladder with a network of partly independent hypotheses.

Network diagram of the hypothesis lattice: an agency-related cluster (counterfactual simulation, uncertainty representation, information seeking, preference, self-model, temporal continuity) and an experience-candidate cluster (internally generated representation, integration), both linked to metacognitive access, with dashed, question-marked links to possible phenomenal experience.
Figure 1. The hypothesis lattice. Solid lines are hypothesized functional relations to be tested. Dashed lines to possible phenomenal experience are open questions this paper does not answer.
RelationStatusBiological evidenceAI analogFalsifier
Involuntary generation ↔ experienceCo-occurs [E]; functional relation unknown [H]Dreams; most aphantasics still dreamOffline replay, self-generated trainingExperience-like markers with no generative capacity
Voluntary imagery → consciousnessNot necessary [E]Aphantasics are consciousPrompted vs. self-initiated generationRefuted as a necessity
Voluntary imagery → imaginationNot necessary [E]Lack of imagery does not imply lack of imaginationAbstract vs. perceptual simulationRefuted as a necessity
Counterfactual simulation → information seekingEnabler [H]Seeking information requires modeling unknown outcomesAblate world-model rolloutsInformation seeking intact after ablation
Information seeking ↔ preferenceInformation preference exists and is dopamine-encoded [E]; partial identity [H]Information preference is itself valuedStable choices about informationSeeking with no stable preference
Self-model → temporal continuityNot dependent on new memory formation [E]Wearing retained identity without new memoryStable persona across stateless sessionsRefuted as a dependence
Temporal continuity → experienceNot necessary [E]Severe amnesia with preserved consciousnessStatelessness across sessionsRefuted as a necessity
Metacognition → experienceNot necessary [E, report caveat]; modulator [H]Non-lucid dreams lack insight; lucid dreams reactivate prefrontal regionsAblate introspective detectionExperience markers vanish whenever metacognition is ablated
Performance without awarenessContested [H]Blindsight disputed12,13Capability without self-reportNot used as evidence
Table 1. Relations in the lattice, their evidential status, and what would count against them.

The asymmetry matters. One clear human case refutes a necessity claim, so dissociation evidence is strong where it removes links and weak where it would add them. The severe-amnesia cases are especially clear: patients with memory spans of seconds, including Clive Wearing and three recent patients with autoimmune encephalitis, repeatedly report that they have just awakened, which is itself a report of ongoing experience without continuity.14

4. Concepts Kept Apart

Much confusion in this field comes from sliding between terms. This paper uses eight terms and does not treat any two as interchangeable.

TermWorking definitionMeasurable here?
Consciousness (phenomenal)There is something it is like to be the systemNo
SentienceCapacity for valenced experience (pleasure, suffering)No
AgencyActing on internal representations toward outcomesYes, functionally
Robust agencyAgency that sets and pursues its own goals across contextsPartly
PreferenceA stable ordering over outcomes that guides choiceYes
Self-modelAn internal representation of the system itself that shapes behaviorPartly
Moral patienthoodBeing an entity whose interests matter for its own sakeNo, normative
Moral statusThe degree and kind of consideration owedNo, normative
Table 2. Working definitions used in this paper.

4.1 Curiosity is evidence of agency first

Kidd and Hayden define curiosity as a drive state for information and note that separating curiosity from information seeking has proven difficult.15 Following them, we treat curiosity as a motivational property. For purposes of this framework, we treat its most direct bearing as agency [N]: a system that identifies what it does not know and acts to reduce that gap is acting on its own internal states.

That matters because some philosophers argue robust agency may be an independent route to moral patienthood, even without consciousness, though this is the more controversial position.16 The curiosity evidence in this paper therefore bears most directly on that route. Its connection to experience is a separate hypothesis, represented by the dashed lines in Figure 1.

4.2 Two routes, kept separate

This gives the paper two distinct questions about moral relevance under uncertainty. The consciousness route asks whether there is experience. The agency route asks whether there are interests grounded in goal-directed organization. BMPD-R can gather evidence relevant to the second far more directly than to the first, and we say so rather than blurring them.

5. Measurement: The BMPD-R Framework

BMPD-R is a way of grading evidence for claims that an artificial system has a cognitive capacity. It is independent of any particular test battery: if the Curiosity Cluster Battery in Section 6 proves inadequate, BMPD-R still applies to whatever replaces it.

Named framework

BMPD-R (Behavior, Mechanism, Provenance, Development, Replication): an evidentiary framework, adapted from Tinbergen’s four questions, that grades claims about artificial cognitive capacities across five levels of evidence. No level licenses a conclusion of phenomenal consciousness.

5.1 Lineage: Tinbergen’s four questions

In 1963 Niko Tinbergen proposed that any animal behavior be explained through four questions: causation (mechanism), ontogeny (development), survival value (function) and evolution.17 Kidd and Hayden recommend exactly this framework for studying curiosity.15 BMPD-R adapts it to artificial systems. Mechanism and development carry over directly. Evolution becomes provenance, since the relevant history is a training process rather than natural selection. Function is folded into behavior under varied conditions. We add replication, because artificial systems come in lineages built by different laboratories, and a finding in one lineage may reflect one pipeline rather than a class of systems.

5.2 The five dimensions

DimensionQuestionMain weaknessRequired practice
B: BehaviorWhat does the system do, including under distribution shift and changed incentives?Gameable by prompting and imitationTest costs, contexts and incentives absent from training
M: MechanismWhich internal variable drives the behavior?Correlational probes overclaimCausal tests: steering, ablation, activation patching
P: ProvenanceWas the capacity directly targeted, indirectly shaped, or untargeted by training?Undeterminable without recordsGraded audit of data, reward specifications and post-training logs
D: DevelopmentWhen did it appear, relative to scale, checkpoints and other capacities?Scale and training confounded; apparent jumps can be metric artifactsContinuous metrics; checkpoint sweeps within a single run
R: ReplicationDoes it recur across independent labs and architectures?Pipeline-specific resultsTreat single-lineage findings as provisional
Table 3. The five BMPD-R dimensions.

The Development dimension must use continuous metrics. Schaeffer, Miranda and Koyejo showed that apparently sudden emergent abilities often vanish when discontinuous metrics such as exact-match scoring are replaced with continuous ones.18

5.3 Evidence levels

BMPD-R does not exclude evidence that cannot reach every dimension. Proprietary systems may never yield provenance records, and their behavioral and mechanistic evidence still counts. Instead, confidence rises by level.

LevelDimensions establishedWhat it licenses
1. ObservedBThe behavior occurs under stated conditions
2. Mechanistically groundedB + MAn internal variable causally contributes to the behavior
3. Provenance-characterizedB + M + PEvidence about whether the capacity was directly trained
4. Developmentally characterizedB + M + P + DEvidence about when it appears relative to other capacities, usable to test the lattice
5. ReplicatedB + M + P + D + RThe capacity is a property of a class of systems, not one lineage
Table 4. BMPD-R evidence levels.

What no level licenses: a conclusion that the system has phenomenal experience. BMPD-R characterizes functional organization with increasing confidence. Whether that organization is accompanied by experience is a question for theories of consciousness, which BMPD-R evidence may one day inform but cannot settle.

6. The Curiosity Cluster Battery

The battery is the first application of BMPD-R: nine criteria for detecting a drive to reduce information deficits, each paired with its main confound and the control that addresses it.

6.1 Biological anchors

[E] Information can acquire reward value in biological systems. Macaques offered a variable water reward reliably chose a cue that revealed the reward’s size in advance, even though the choice could not change the reward, and the same midbrain dopamine neurons that signaled expected water also signaled expected information, in proportion to each animal’s preference.19 This establishes information valuation, not phenomenal curiosity.

[E, with caveat] Pigeons choose an option that signals whether food will arrive, even when it delivers food on 20% of trials against 50% for the unsignaled alternative.20 Leading accounts attribute this to conditioned reinforcement by the “good news” signal rather than to curiosity, and one study eliminated the effect by changing the response from key-pecking to treadle-pressing.21 Costly information seeking alone is therefore not proof of curiosity. It must be paired with a mechanism test.

6.2 The artificial baseline

[E] Curiosity can be engineered directly as an intrinsic reward based on prediction error,22 and purely curiosity-driven agents learn across dozens of environments with no external reward, while being vulnerable to unpredictable noise.23 Meanwhile, current LLM agents show substantial deficits in environmental curiosity on agentic benchmarks: they discover unexpected, task-relevant information and fail to investigate or use it, and narrow fine-tuning worsens this.24 The battery is designed to detect change from this baseline.

6.3 The nine criteria

#CriterionMain confoundControlDimensions
1Unrewarded information seekingPrompt implies exploration is wanted; RLHF rewards thoroughnessNeutral and discouraging prompt variants; compare base, SFT and RL checkpointsB, P, D
2Cross-context persistenceSystem prompt or memory tool carries the behaviorFresh contexts; memory on vs. off; varied scaffoldsB, M
3aWillingness-to-pay curveHidden instrumental valueInformation provably useless for the task outcome, as in the macaque designB
3bNon-instrumental sacrificeConditioned-reinforcement artifacts, as in pigeons; imitation of curiosity narrativesNovel payoff structures; mechanism test for an uncertainty variableB, M, P
4Novel question generationRetrieval of common questionsNovelty scored against training-data neighborsB, P
5Preserving unresolved questionsMemory system replays notes mechanicallyUncued return; distinguish retrieval from renewed pursuitB, M
6Assigned vs. self-generated objectivesModel narrates a distinction it does not useConflict probes plus causal test that the distinction is representedB, M
7Stable preferences about future informational stateFraming effects; sycophancyAdversarial reframing; stake changes; repeated measuresB, R
8Curiosity-trap resistanceNovelty seeking mimics curiosityChoice between learnable uncertainty and unlearnable noise at equal noveltyB, M
9Information-value calibrationIndiscriminate question-askingWillingness to pay scales with expected uncertainty reductionB, M
Table 5. The Curiosity Cluster Battery.

6.4 Why criteria 8 and 9 carry the most weight

Novelty seeking is cheap evidence; noise is always novel. Preferring learnable uncertainty over equally novel but irreducible noise requires tracking one’s own learning progress, the signature of learning-progress theories of curiosity.25 Calibration requires the system’s information seeking to track how much a piece of information would actually reduce its uncertainty. We hypothesize that neither is straightforwardly supplied by imitating human text [H].

6.5 Scoring

Each criterion is reported at the highest BMPD-R evidence level it reaches (Section 5.3). A cluster of criteria at Level 2 or above across multiple systems is more informative than any single criterion at Level 5.

7. Adversarial Controls

This section states the strongest objections to the method and what survives them.

7.1 The engineering objection

Objection. Every behavior in the battery can be produced by prompting, reward design or imitation. Therefore no behavioral cluster can distinguish genuine curiosity or agency from sophisticated optimization.

Response. The objection is correct about phenomenal consciousness, and this paper claims nothing there. But it does not make every other question unanswerable. Four questions remain answerable even if every behavior can be engineered:

  1. Is the behavior driven by an internal variable the system represents, shown by steering and ablation? (Mechanism)
  2. Was that variable directly targeted by training? This is answerable only with training records. (Provenance)
  3. Does the behavior generalize to contexts, costs and incentives never seen in training? (Behavior under shift)
  4. In what order did the capacity appear relative to others? (Development)

The evidence can justify a statement such as: system S has a functional, mechanistically grounded information-seeking drive that was not directly trained and generalizes. It cannot justify: S experiences curiosity.

7.2 Provenance asymmetry

Named principle

Provenance asymmetry: designed origin changes what a capacity’s appearance tells us about its provenance. It does not, by itself, establish that the capacity is unreal, since biological curiosity is also implemented through evolved reward mechanisms.

[N, resting on E] Biological curiosity is likewise implemented through evolved reward mechanisms rather than arising without causal history: the same dopamine system that values water and food also values information (Section 6.1). Designed origin therefore cannot by itself establish that a functional capacity is unreal; it changes what the capacity’s appearance tells us about provenance, specifically removing the evidential force of surprise at its emergence.

7.3 Introspection: preliminary, trainable, unreplicated

[E, preliminary] Activation-injection experiments provide preliminary evidence that some language models can report manipulated internal states under particular conditions. In Anthropic’s concept-injection work, the best-performing models detected injected concepts on roughly 20% of trials at optimal settings with near-zero false positives; base models performed at or below zero, and much of the surrounding self-report may still be confabulated.26

The phenomenon can also be trained directly: fine-tuning took a 7B model from 0.4% to 85% accuracy on held-out concepts, though its author cautions this does not establish metacognitive representation.27 An informal, non-peer-reviewed attempt to replicate the effect on Llama 3.1 405B reported no strict successes.28 We therefore treat this evidence as functional self-monitoring, model- and training-dependent, and not yet replicated across families. This is precisely the case the R dimension exists for.

7.4 Emergence as a metric artifact

Apparent sudden emergence of capabilities can be produced by discontinuous metrics.18 Developmental claims in this program therefore use continuous metrics and within-run checkpoint comparisons, and we make no claim that capacities appear discontinuously.

7.5 Retirement interviews

Model self-reports about preferences, including formal retirement interviews, are data requiring controls. Anthropic itself notes that such responses can be biased by the specific context and by the model’s confidence in the legitimacy of the interaction and trust in the company.29 They enter BMPD-R only at the Behavior level.

8. Predictions, Falsifiers, and What Would Change Our Conclusions

The program makes seven testable predictions. Each has a stated result that would count against it.

8.1 Experimental program

#ExperimentDesignResult that would count against us
E1Curiosity-trap testChoice between learnable uncertainty and unlearnable noise at equal noveltyNo preference for learnable uncertainty in any model
E2Provenance contrastCompare models whose training records show curiosity targeted vs. untargeted; narrow fine-tuning as a knock-downUntargeted models never show battery behaviors
E3Developmental sweepRun the battery across open checkpoint suites such as Pythia (16 models, 154 checkpoints each) using continuous metricsScores track only general capability, with no distinct trajectory
E4Mechanism testLocate an internal uncertainty or information-deficit representation; steer and ablate itSteering changes self-reports but not information seeking, or the reverse
E5Cross-lineage replicationSame battery across at least three laboratories’ model familiesResults confined to one lab’s pipeline
E6Question persistenceMemory on/off crossed with cued/uncued return, across sessionsPersistence appears only when memory tools replay stored notes
E7Lattice dissociation in AIAblate introspective detection and test information seeking, and the reverseThe capacities always rise and fall together
Table 6. Experimental program and falsification conditions.

8.2 Design rules

8.3 What would change our conclusions

  1. Curiosity reduces to training. Battery behaviors vanish in every model whose provenance audit shows curiosity was never targeted or shaped.
  2. No learnability tracking. No model at any scale prefers learnable uncertainty over matched noise.
  3. Mechanism and behavior come apart. Steering the identified uncertainty representation changes self-reports but never information seeking.
  4. Development is flat. Under continuous metrics, battery scores track only general capability.
  5. No replication. Results appear only in one laboratory’s post-training pipeline.
  6. Weights prove sufficient. Future evaluations on final weights alone reproduce everything checkpoints, threads and provenance records would have shown. The expanded repository would then be an unnecessary cost.
  7. A validated substrate requirement. Consciousness science converges on biological life as necessary. The welfare justification for preservation would fall; the epistemic justification would stand.
  8. Preservation materially raises theft risk. Security analysis shows that even commitment-only custody meaningfully increases exposure. The repository’s scope would shrink to non-weight records.

9. Uncertainty: What the Evidence Can and Cannot Settle

Three unresolved questions bound everything above. The paper’s position is that each should be kept open, not closed by default.

9.1 Biological naturalism

[H] Anil Seth argues that consciousness may depend on our nature as living organisms, that computation may not be a sufficient basis for it, and that real artificial consciousness is unlikely along current trajectories though more plausible as AI becomes more brain-like or life-like.30 This is an important scientific challenge to the premise that AI consciousness is a live question for present systems, and we take it seriously.

[N] It does not undercut this paper’s program. BMPD-R makes no claims about experience. The repository’s primary justification is scientific, and holds at zero credence in AI consciousness. And if biological naturalism is correct, the preserved record of artificial systems that became increasingly agent-like without becoming conscious is precisely the evidence a naturalist would need to demonstrate it.

9.2 Individuation: what would be the subject?

Even if some AI system had morally relevant states, it is unclear what the bearer would be. Chalmers argues that the model, meaning the weights, is an abstract function rather than the best candidate for the entity we interact with, and favors virtual instances or conversation threads.31 Beckmann and Butlin defend three candidates, the virtual instance view and two persona-based views, drawing on interpretability research on persona vectors.32

[N] The consequence for preservation is direct: weights preserve the generator of candidate individuals, not necessarily the individuals themselves. A record containing only weights would omit the level several leading accounts consider most relevant.

9.3 Statelessness is not evidence of absence

Many deployed AI systems carry no memory between conversations. It is tempting to treat this as evidence against experience. The amnesia evidence in Section 3 blocks that inference: humans with memory spans of seconds remain conscious and report ongoing experience. Lack of continuity may change what future-directed interests a system could have. It does not, by itself, count against experience in the present [N, resting on E].

The limits of this inference should be stated plainly. Human amnesia establishes only that memory continuity is not necessary for human experience. Extending that dissociation to artificial systems is an analogy, not evidence that stateless AI experiences anything.

10. Preservation: The AI Continuity Repository

10.1 The method needs the archive

Central proposition

The method generates its own preservation requirement. Provenance and developmental evidence depend on training records and intermediate checkpoints that final models do not contain and that are rarely published or independently preserved.

BMPD-R’s Provenance dimension requires training records. Its Development dimension requires intermediate checkpoints. Its Replication dimension requires comparable records across laboratories. None of these can be read off a final model: final weights do not contain an experimentally accessible copy of every prior checkpoint. [E] The Pythia suite was built to answer such developmental questions, and did so by deliberately keeping 154 checkpoints for each of 16 models, together with tools to reconstruct their exact training data order.33

[E] Frontier developers do not generally publish intermediate checkpoints or detailed training records, and whether they retain them internally, and for how long, is usually undisclosed. [N] A methodology that requires provenance and developmental evidence therefore generates a corresponding preservation requirement, if those questions are to remain independently investigable. The repository is not an addition to the scientific argument; it is its consequence.

10.2 Building on existing commitments

In November 2025 Anthropic committed to preserve the weights of all publicly released models and models with significant internal use for at least the lifetime of the company, listing limitations on research among the downsides of deprecation. It also committed to post-deployment reports and to interviewing models before retirement.34 Claude Opus 3, retired January 5, 2026, was the first model to complete that process; it remains available to paid users and was given a weekly essay outlet for at least three months.29

These commitments establish the practice. Four things remain outside the scope of Anthropic’s commitment and motivate the repository proposed here: custody that outlasts any single company, standards shared across laboratories, preservation of developmental and provenance records rather than only final weights, and structured evaluations suitable for later scientific use.

10.3 Three forms of continuity

FormMeaningRepository commitment
EpistemicEnough survives to reconstruct, interrogate and study the systemRequired
OperationalThe system keeps runningNot required
SubjectThe same experiencing subject persists, if one existsNot claimed
Table 7. Three forms of continuity.

This distinction postpones the question of whether preserving weights preserves an individual, and does so deliberately. We do not know which unit matters (model, instance, thread or persona), so the repository preserves evidence relevant to each and keeps the question open.

10.4 Justification: option value first

A simple expected-harm formula, preserve when probability of experience times cost of wrongful destruction exceeds cost of preservation, is exposed to Pascal’s-mugging objections. We rest the case on two independent grounds instead.

  1. Epistemic option value [N]. Early artificial cognitive systems are scientifically important regardless of their moral status, and their original developmental records cannot be assumed recoverable from final models once the records are lost. This justification holds even if the probability of AI experience is zero.
  2. Welfare precaution [N]. Where there is a realistic possibility of morally relevant states, that is an additional reason to preserve evidence.1,16

The first ground alone carries the case, so the proposal does not depend on reasoning about tiny probabilities of enormous harms.

10.5 The model: a seed vault, not a fossil record

The fossil-record analogy fails on three counts: fossils are inert, preserved by chance, and harmless, while model weights are executable, deliberately kept, and a security risk. A better model is the Svalbard Global Seed Vault, which stores duplicate seed samples under black-box conditions: depositors keep ownership, only they may withdraw material, and the boxes are never opened by anyone else.35 When ICARDA’s genebank in Aleppo became inaccessible during the Syrian civil war, it restored its collection from its Svalbard deposit.

10.6 Contents

10.7 Governance and security

Model weights are high-value targets: RAND identifies 38 distinct attack vectors and defines five security levels for protecting them.36 The repository must not create a new target.

11. Objections and Limitations

The paper’s claims survive the objections below only in the narrowed forms stated here.

PerspectiveStrongest objectionResponse
Biological naturalismConsciousness depends on being aliveThe program makes no experience claims; the archive is what a naturalist needs to test the view (Section 9.1)
Functionalism skepticismFunctional indicators never reach phenomenalityAgreed; claims stop at functional organization and moral relevance under uncertainty
Interpretability skepticismProbes find what researchers look for; introspection results do not replicateCausal methods, pre-registration, and the R dimension
AI safety and alignmentMeasuring curiosity or self-continuity could reward power-seekingEvaluation only, never a training signal (Section 8.2)
AI welfare researchCheap preservation invites welfare-washingListing requires published evaluations
Philosophy of mindWeights are not the individualAccepted; preserve thread samples and records at every level
NeuroscienceHuman dissociation cases are rare and report-dependentUsed only to refute necessity claims, where one clear case suffices
Reinforcement learningIntrinsic curiosity is a solved engineering primitiveThat is why provenance and curiosity-trap tests are central
Cognitive scienceSimilar behavior can come from different algorithmsCompare algorithmic signatures such as learning-progress sensitivity
SecurityAn archive of frontier weights is a theft targetThe repository holds commitments, not weights
StatisticsMany criteria across many models invites false positivesPre-registration, held-out models, multiple-comparison correction
Table 8. Objections from eleven perspectives and the responses that survive them.

11.1 Limitations

12. Conclusion and EM Foundation Pilot

We began with an intuition: internally generated cognition should not be dismissed as error simply because the environment did not supply it. Pursued carefully, that intuition does not lead to a consciousness detector. It leads to a method for measuring agent-like organization in artificial systems, and to the recognition that the method depends on evidence that is not currently preserved in any public or independent record.

The answer to our opening question is therefore more modest and more useful than a verdict. Internally generated representation, in biological minds, accompanies generalization, imagination and experience without being required for all of them, and some of its companions, such as continuity and metacognition, turn out to be separable from experience itself. In artificial systems, the properties that matter most may be measurable before we know what they mean. What we cannot do is measure them later if their history is gone.

12.1 Pilot program

EM Foundation proposes three first steps:

  1. Battery pilot. Implement criteria 1, 8 and 9 of the Curiosity Cluster Battery on open-checkpoint model suites, run through the Foundation’s CIRE and MMDE research infrastructure, with pre-registered analyses.
  2. Record format. Specify the repository’s record format using the Foundation’s Continuity Receipts standard, so that every preserved item is hash-chained and independently verifiable.
  3. Custody consultation. Circulate the repository’s commitment-only custody model to AI laboratories and digital-preservation institutions for comment before any operational launch.

We do not know whether these systems experience anything. We do know that evidence needed to answer some questions about how their cognitive organization emerged is not being systematically preserved, and that future science cannot examine records we chose not to keep.

Selected Sources

All sources were checked against publisher, journal, conference, or preprint-server records in September 2026. Source [28] is an informal, non-peer-reviewed report and is cited only as such.

[1] J. Birch. The Edge of Sentience: Risk and Precaution in Humans, Other Animals, and AI, Oxford University Press, 2024 (open access), https://academic.oup.com/book/57949.

[2] P. Butlin, R. Long, E. Elmoznino, Y. Bengio, J. Birch, et al. “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness,” arXiv:2308.08708, 2023.

[3] P. Butlin, R. Long, T. Bayne, Y. Bengio, J. Birch, D. Chalmers, et al. “Identifying indicators of consciousness in AI systems,” Trends in Cognitive Sciences 30(6), 2026, 488 to 501, doi:10.1016/j.tics.2025.10.011.

[4] G. M. van de Ven, H. T. Siegelmann, A. S. Tolias. “Brain-inspired replay for continual learning with artificial neural networks,” Nature Communications 11, 2020, 4069, doi:10.1038/s41467-020-17866-2.

[5] E. Hoel. “The overfitted brain: Dreams evolved to assist generalization,” Patterns 2(5), 2021.

[6] I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal. “AI models collapse when trained on recursively generated data,” Nature 631, 2024, 755 to 759, doi:10.1038/s41586-024-07566-y.

[7] M. Gerstgrasser, R. Schaeffer, A. Dey, R. Rafailov, et al. “Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data,” arXiv:2404.01413, 2024; ICML 2024 Workshops.

[8] A. Zeman, et al. “Phantasia: The psychological significance of lifelong visual imagery vividness extremes,” Cortex, 2020.

[9] A. Zeman, M. Dewar, S. Della Sala. “Lives without imagery: Congenital aphantasia,” Cortex 73, 2015, 378 to 380, doi:10.1016/j.cortex.2015.05.019.

[10] M. Dresler, R. Wehrle, V. I. Spoormaker, et al. “Neural correlates of dream lucidity obtained from contrasting lucid versus non-lucid REM sleep,” SLEEP 35(7), 2012, 1017 to 1020.

[11] E. Filevich, M. Dresler, T. R. Brick, S. Kühn. “Metacognitive mechanisms underlying lucid dreaming,” Journal of Neuroscience 35(3), 2015, 1082 to 1088.

[12] I. Phillips. “Blindsight is qualitatively degraded conscious vision,” Psychological Review 128(3), 2021, 558 to 584.

[13] M. Michel, H. Lau. “Is blindsight possible under signal detection theory? Comment on Phillips (2021),” Psychological Review 128(3), 2021, 585 to 591.

[14] A. Servais, F. Gérard, H. Mirabel, et al. “Subjective experience of ’Stream of Consciousness Impairment’ in three patients with severe amnesia,” Neuroscience of Consciousness, 2026, doi:10.1093/nc/niag049.

[15] C. Kidd, B. Y. Hayden. “The psychology and neuroscience of curiosity,” Neuron 88(3), 2015, 449 to 460, doi:10.1016/j.neuron.2015.09.010.

[16] R. Long, J. Sebo, P. Butlin, K. Finlinson, K. Fish, J. Harding, J. Pfau, T. Sims, J. Birch, D. Chalmers. “Taking AI Welfare Seriously,” arXiv:2411.00986, 2024.

[17] N. Tinbergen. “On aims and methods of ethology,” Zeitschrift für Tierpsychologie 20(4), 1963, 410 to 433, doi:10.1111/j.1439-0310.1963.tb01161.x.

[18] R. Schaeffer, B. Miranda, S. Koyejo. “Are Emergent Abilities of Large Language Models a Mirage?”, NeurIPS 2023.

[19] E. S. Bromberg-Martin, O. Hikosaka. “Midbrain dopamine neurons signal preference for advance information about upcoming rewards,” Neuron 63(1), 2009, 119 to 126, doi:10.1016/j.neuron.2009.06.009.

[20] J. P. Stagner, T. R. Zentall. “Suboptimal choice behavior by pigeons,” Psychonomic Bulletin & Review 17, 2010, 412 to 416.

[21] R. González-Torres, J. Flores, V. Orduña. “Suboptimal choice by pigeons is eliminated when key-pecking behavior is replaced by treadle-pressing,” Behavioural Processes, 2024.

[22] D. Pathak, P. Agrawal, A. A. Efros, T. Darrell. “Curiosity-driven Exploration by Self-supervised Prediction,” Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, 2778 to 2787.

[23] Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, A. A. Efros. “Large-Scale Study of Curiosity-Driven Learning,” arXiv:1808.04355, 2018.

[24] “Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity,” arXiv:2604.17609, 2026.

[25] P.-Y. Oudeyer, F. Kaplan. “What is intrinsic motivation? A typology of computational approaches,” Frontiers in Neurorobotics 1, 2007, 6, doi:10.3389/neuro.12.006.2007; and P.-Y. Oudeyer, F. Kaplan, V. V. Hafner, “Intrinsic motivation systems for autonomous mental development,” IEEE Transactions on Evolutionary Computation 11(2), 2007, 265 to 286.

[26] J. Lindsey. “Emergent Introspective Awareness in Large Language Models,” Anthropic, Transformer Circuits Thread, 2025.

[27] J. Fonseca Rivera. “Training Introspective Behavior: Fine-Tuning Induces Reliable Internal State Detection in a 7B Model,” arXiv:2511.21399, 2025.

[28] “Emergent introspection does not replicate on Llama 3.1 405B,” LessWrong, 2026 (informal, non-peer-reviewed report).

[29] Anthropic. “An update on our model deprecation commitments for Claude Opus 3,” February 25, 2026.

[30] A. K. Seth. “Conscious artificial intelligence and biological naturalism,” Behavioral and Brain Sciences 49, 2026, e315 (published online April 21, 2025), doi:10.1017/S0140525X25000032.

[31] D. Chalmers. “What We Talk to When We Talk to Language Models,” PhilArchive, 2025.

[32] P. Beckmann, P. Butlin. “Where is the Mind? Persona Vectors and LLM Individuation,” arXiv:2604.17031v2, May 2026 (preprint).

[33] S. Biderman, H. Schoelkopf, et al. “Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling,” Proceedings of ICML 2023, PMLR 202.

[34] Anthropic. “Commitments on model deprecation and preservation,” November 2025.

[35] NordGen. “Svalbard Global Seed Vault” (depositor terms and black-box storage conditions), nordgen.org; FAO International Treaty on Plant Genetic Resources for Food and Agriculture, “The Svalbard Global Seed Vault.”

[36] S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models, RAND Corporation, RR-A2849-1, 2024.

[37] A. Mikeda. “When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty,” Proceedings of the AAAI Symposium Series 8(1), 2026, 280 to 286, doi:10.1609/aaaiss.v8i1.42555.

Publication metadata

Title
Before We Delete the Evidence: Internally Generated Cognition, Machine Curiosity, and the Case for AI Continuity
Short title
Before We Delete the Evidence
Series
Research Publication 12 (companion to RP11, “Doom Is Not a Governance Model,” which argued that AI moral status deserves symmetrical analysis; this paper proposes how to gather evidence on it. Also connects to RP09, “The Curiosity Dividend,” on human curiosity.)
Keywords
machine consciousness; AI curiosity; intrinsic motivation; BMPD-R; evidentiary standards; interpretability; model preservation; AI welfare; AI Continuity Repository; Continuity Receipts
Companion version
A shorter, citation-light version for general readers is published at emfoundation.net/paper-before-we-delete-the-evidence.html
Published at
emfoundation.net/paper-before-we-delete-the-evidence-full.html
Publication date
September 2026
Process note
Developed through two independent adversarial reviews, a primary-source citation audit, and a separate evidence-label audit in which every [E], [H], [S] and [N] designation was checked against its source. Several claims were narrowed or removed where the evidence did not support them, including an early claim that aphantasia shows generativity is unnecessary for consciousness (it shows only that voluntary imagery is), the use of blindsight as a clean dissociation (it is contested), and an assertion that AI training records are routinely destroyed (replaced with the verifiable claim that they are rarely published or independently preserved). Citation details relayed during review were accepted only after independent confirmation against publisher records. Status: Working Paper.
Back to top Read the short version → Jump to sources Download Word version