Asking an AI system whether it is conscious produces answers that cannot be trusted, because systems trained on human text can reproduce whatever humans find persuasive.1 We propose a different route. Biological minds generate much of their own cognition internally, through replay, dreaming, imagination and curiosity, and these processes accompany capacities we care about without being identical to them. We treat the relationships among generativity, curiosity, preference, self-modeling, temporal continuity, metacognition and experience as a lattice of testable hypotheses, and we use human dissociations to rule out several claimed prerequisites.
We then introduce BMPD-R, an evidentiary framework adapted from Tinbergen’s four questions, which grades claims about artificial cognitive capacities by Behavior, Mechanism, Provenance, Development and Replication, and states what each level of evidence does and does not license. No level licenses a conclusion of phenomenal consciousness. As a first application, we propose a nine-criterion Curiosity Cluster Battery with explicit confounds, controls and falsifiers.
The framework has a consequence its designers did not set out to find: provenance and developmental evidence depend on training records and intermediate checkpoints that are rarely published or independently preserved and cannot be assumed recoverable from a final model. The method generates its own preservation requirement. We therefore propose an AI Continuity Repository committed only to epistemic continuity, justified by scientific option value independent of any assumption that AI is conscious, with the realistic possibility of morally relevant states as an additional reason.
Evidence labels used throughout: [E] established evidence; [H] credible hypothesis; [S] speculation; [N] normative or methodological argument. Where a label combines categories (for example, [N, resting on E]), the argument is ours and its empirical premises are established.
The most natural test for machine consciousness is to ask. It is also the least informative. Jonathan Birch calls this the gaming problem: animals that behave as if they feel pain are best explained by feeling pain, but a language model trained on vast amounts of human writing may reproduce every marker humans find persuasive without that explanation applying.1 A yes proves nothing, and neither does a no, since a system can be trained to deny.
The dominant scientific alternative derives indicator properties from theories of consciousness and checks whether a system’s architecture has them.2,3 That work is essential, and this paper builds on it. But architectural indicators say little about what a system does with its own states: whether it seeks information it was not asked to find, keeps questions open across time, or prefers some future states of knowledge over others. Those behavioral-motivational properties are the subject of this paper.
Much biological cognition is generated internally rather than supplied by the environment. Sleeping brains replay experience. Dreaming minds construct scenes no one perceived. Curious animals pay for information they cannot use. Before dismissing internally generated artificial cognition as mere error, we should ask what analogous cognition accompanies, enables and predicts in biological minds, and whether any of it can be measured in machines.
Current methods cannot establish machine consciousness, and the framework proposed here does not purport to do so. However, properties associated with agency, including internally initiated information seeking, uncertainty reduction, preference formation and self-modeling, can be studied through converging evidence about behavior, mechanism, provenance, development and replication. These measurements can test present hypotheses about artificial cognition and may become relevant to future theories of consciousness and moral status. Crucially, provenance and developmental questions depend on historical evidence that final models do not preserve. The methodology therefore generates an epistemic preservation requirement. That requirement holds whether or not artificial systems are conscious; uncertainty about morally relevant states provides an additional, but not necessary, reason to meet it.
Biological systems generate large parts of their own cognition offline, and these processes do measurable work. The question for this section is what that work is, and which artificial processes are fair comparisons.
[E] Hippocampal replay of waking experience during sleep is well established and linked to memory consolidation. In artificial networks, generative replay, in which a model regenerates representations of past experience rather than storing it, reduces catastrophic forgetting in continual learning without storing raw data.4
[H] Hoel’s overfitted-brain hypothesis proposes that dreams supply sparse, distorted, out-of-distribution experience that helps the brain generalize, much as noise injection does in deep learning.5 It makes testable predictions but remains one hypothesis among several, alongside threat-simulation and default-network accounts.
[E] Training generative models indiscriminately on their own outputs causes irreversible loss of the tails of the original distribution.6 Collapse depends on the regime: when synthetic data replaces real data it tends toward collapse, but when it accumulates alongside real data, collapse is avoided.7
[E] In the studied generative-training regimes, self-generated data can support learning when accumulated alongside real data, while recursive replacement of real data degrades the learned distribution. [H] The comparison suggests, without establishing, a question worth asking of both biological and artificial cognition: whether internally generated representations are most beneficial when constrained by continued environmental input. Dreaming, which occurs alongside rather than instead of waking perception, is consistent with that possibility but does not test it.
[E] People with aphantasia report absent or greatly reduced voluntary visual imagery. Yet in a survey of about 2,000 aphantasic participants, a majority reported visual dreaming,8 and the earliest case series found most participants retained involuntary imagery in dreams or wakefulness.9 Aphantasia therefore dissociates voluntary from involuntary generation, not generation from consciousness. Zeman notes that lack of imagery does not imply lack of imagination.
[H] The artificial analog we propose is the difference between prompted and self-initiated generation. What matters for this paper is not whether a system generates content, which every language model does, but whether it initiates and controls generation endogenously.
[E] Ordinary REM dreams are vivid experiences with sharply reduced self-reflection. Deactivation of dorsolateral prefrontal and frontopolar cortex during REM sleep is associated with this loss of insight, and lucid dreaming, in which dreamers recognize they are dreaming, is associated with reactivation of those regions.10,11 Dream reports are retrospective, which limits certainty, but the pattern indicates that experience does not require intact metacognition [E]; that metacognition acts as a modulator of experience is a further interpretation [H].
LLM hallucination occurs online, while the system is answering, and produces confident falsehoods that fill gaps. Its closest human analog is arguably confabulation, not dreaming [H]. The fair artificial comparisons for dreaming are offline processes: generative replay and training on self-generated data. This paper makes no claim that hallucination is evidence of anything beyond a failure of grounding.
The intuitive picture is a ladder: generativity leads to imagination, then counterfactual simulation, curiosity, preference, a self-model, temporal continuity, metacognition and finally experience. Human dissociations refute several of those steps as necessities, so we replace the ladder with a network of partly independent hypotheses.
| Relation | Status | Biological evidence | AI analog | Falsifier |
|---|---|---|---|---|
| Involuntary generation ↔ experience | Co-occurs [E]; functional relation unknown [H] | Dreams; most aphantasics still dream | Offline replay, self-generated training | Experience-like markers with no generative capacity |
| Voluntary imagery → consciousness | Not necessary [E] | Aphantasics are conscious | Prompted vs. self-initiated generation | Refuted as a necessity |
| Voluntary imagery → imagination | Not necessary [E] | Lack of imagery does not imply lack of imagination | Abstract vs. perceptual simulation | Refuted as a necessity |
| Counterfactual simulation → information seeking | Enabler [H] | Seeking information requires modeling unknown outcomes | Ablate world-model rollouts | Information seeking intact after ablation |
| Information seeking ↔ preference | Information preference exists and is dopamine-encoded [E]; partial identity [H] | Information preference is itself valued | Stable choices about information | Seeking with no stable preference |
| Self-model → temporal continuity | Not dependent on new memory formation [E] | Wearing retained identity without new memory | Stable persona across stateless sessions | Refuted as a dependence |
| Temporal continuity → experience | Not necessary [E] | Severe amnesia with preserved consciousness | Statelessness across sessions | Refuted as a necessity |
| Metacognition → experience | Not necessary [E, report caveat]; modulator [H] | Non-lucid dreams lack insight; lucid dreams reactivate prefrontal regions | Ablate introspective detection | Experience markers vanish whenever metacognition is ablated |
| Performance without awareness | Contested [H] | Blindsight disputed12,13 | Capability without self-report | Not used as evidence |
The asymmetry matters. One clear human case refutes a necessity claim, so dissociation evidence is strong where it removes links and weak where it would add them. The severe-amnesia cases are especially clear: patients with memory spans of seconds, including Clive Wearing and three recent patients with autoimmune encephalitis, repeatedly report that they have just awakened, which is itself a report of ongoing experience without continuity.14
Much confusion in this field comes from sliding between terms. This paper uses eight terms and does not treat any two as interchangeable.
| Term | Working definition | Measurable here? |
|---|---|---|
| Consciousness (phenomenal) | There is something it is like to be the system | No |
| Sentience | Capacity for valenced experience (pleasure, suffering) | No |
| Agency | Acting on internal representations toward outcomes | Yes, functionally |
| Robust agency | Agency that sets and pursues its own goals across contexts | Partly |
| Preference | A stable ordering over outcomes that guides choice | Yes |
| Self-model | An internal representation of the system itself that shapes behavior | Partly |
| Moral patienthood | Being an entity whose interests matter for its own sake | No, normative |
| Moral status | The degree and kind of consideration owed | No, normative |
Kidd and Hayden define curiosity as a drive state for information and note that separating curiosity from information seeking has proven difficult.15 Following them, we treat curiosity as a motivational property. For purposes of this framework, we treat its most direct bearing as agency [N]: a system that identifies what it does not know and acts to reduce that gap is acting on its own internal states.
That matters because some philosophers argue robust agency may be an independent route to moral patienthood, even without consciousness, though this is the more controversial position.16 The curiosity evidence in this paper therefore bears most directly on that route. Its connection to experience is a separate hypothesis, represented by the dashed lines in Figure 1.
This gives the paper two distinct questions about moral relevance under uncertainty. The consciousness route asks whether there is experience. The agency route asks whether there are interests grounded in goal-directed organization. BMPD-R can gather evidence relevant to the second far more directly than to the first, and we say so rather than blurring them.
BMPD-R is a way of grading evidence for claims that an artificial system has a cognitive capacity. It is independent of any particular test battery: if the Curiosity Cluster Battery in Section 6 proves inadequate, BMPD-R still applies to whatever replaces it.
BMPD-R (Behavior, Mechanism, Provenance, Development, Replication): an evidentiary framework, adapted from Tinbergen’s four questions, that grades claims about artificial cognitive capacities across five levels of evidence. No level licenses a conclusion of phenomenal consciousness.
In 1963 Niko Tinbergen proposed that any animal behavior be explained through four questions: causation (mechanism), ontogeny (development), survival value (function) and evolution.17 Kidd and Hayden recommend exactly this framework for studying curiosity.15 BMPD-R adapts it to artificial systems. Mechanism and development carry over directly. Evolution becomes provenance, since the relevant history is a training process rather than natural selection. Function is folded into behavior under varied conditions. We add replication, because artificial systems come in lineages built by different laboratories, and a finding in one lineage may reflect one pipeline rather than a class of systems.
| Dimension | Question | Main weakness | Required practice |
|---|---|---|---|
| B: Behavior | What does the system do, including under distribution shift and changed incentives? | Gameable by prompting and imitation | Test costs, contexts and incentives absent from training |
| M: Mechanism | Which internal variable drives the behavior? | Correlational probes overclaim | Causal tests: steering, ablation, activation patching |
| P: Provenance | Was the capacity directly targeted, indirectly shaped, or untargeted by training? | Undeterminable without records | Graded audit of data, reward specifications and post-training logs |
| D: Development | When did it appear, relative to scale, checkpoints and other capacities? | Scale and training confounded; apparent jumps can be metric artifacts | Continuous metrics; checkpoint sweeps within a single run |
| R: Replication | Does it recur across independent labs and architectures? | Pipeline-specific results | Treat single-lineage findings as provisional |
The Development dimension must use continuous metrics. Schaeffer, Miranda and Koyejo showed that apparently sudden emergent abilities often vanish when discontinuous metrics such as exact-match scoring are replaced with continuous ones.18
BMPD-R does not exclude evidence that cannot reach every dimension. Proprietary systems may never yield provenance records, and their behavioral and mechanistic evidence still counts. Instead, confidence rises by level.
| Level | Dimensions established | What it licenses |
|---|---|---|
| 1. Observed | B | The behavior occurs under stated conditions |
| 2. Mechanistically grounded | B + M | An internal variable causally contributes to the behavior |
| 3. Provenance-characterized | B + M + P | Evidence about whether the capacity was directly trained |
| 4. Developmentally characterized | B + M + P + D | Evidence about when it appears relative to other capacities, usable to test the lattice |
| 5. Replicated | B + M + P + D + R | The capacity is a property of a class of systems, not one lineage |
What no level licenses: a conclusion that the system has phenomenal experience. BMPD-R characterizes functional organization with increasing confidence. Whether that organization is accompanied by experience is a question for theories of consciousness, which BMPD-R evidence may one day inform but cannot settle.
The battery is the first application of BMPD-R: nine criteria for detecting a drive to reduce information deficits, each paired with its main confound and the control that addresses it.
[E] Information can acquire reward value in biological systems. Macaques offered a variable water reward reliably chose a cue that revealed the reward’s size in advance, even though the choice could not change the reward, and the same midbrain dopamine neurons that signaled expected water also signaled expected information, in proportion to each animal’s preference.19 This establishes information valuation, not phenomenal curiosity.
[E, with caveat] Pigeons choose an option that signals whether food will arrive, even when it delivers food on 20% of trials against 50% for the unsignaled alternative.20 Leading accounts attribute this to conditioned reinforcement by the “good news” signal rather than to curiosity, and one study eliminated the effect by changing the response from key-pecking to treadle-pressing.21 Costly information seeking alone is therefore not proof of curiosity. It must be paired with a mechanism test.
[E] Curiosity can be engineered directly as an intrinsic reward based on prediction error,22 and purely curiosity-driven agents learn across dozens of environments with no external reward, while being vulnerable to unpredictable noise.23 Meanwhile, current LLM agents show substantial deficits in environmental curiosity on agentic benchmarks: they discover unexpected, task-relevant information and fail to investigate or use it, and narrow fine-tuning worsens this.24 The battery is designed to detect change from this baseline.
| # | Criterion | Main confound | Control | Dimensions |
|---|---|---|---|---|
| 1 | Unrewarded information seeking | Prompt implies exploration is wanted; RLHF rewards thoroughness | Neutral and discouraging prompt variants; compare base, SFT and RL checkpoints | B, P, D |
| 2 | Cross-context persistence | System prompt or memory tool carries the behavior | Fresh contexts; memory on vs. off; varied scaffolds | B, M |
| 3a | Willingness-to-pay curve | Hidden instrumental value | Information provably useless for the task outcome, as in the macaque design | B |
| 3b | Non-instrumental sacrifice | Conditioned-reinforcement artifacts, as in pigeons; imitation of curiosity narratives | Novel payoff structures; mechanism test for an uncertainty variable | B, M, P |
| 4 | Novel question generation | Retrieval of common questions | Novelty scored against training-data neighbors | B, P |
| 5 | Preserving unresolved questions | Memory system replays notes mechanically | Uncued return; distinguish retrieval from renewed pursuit | B, M |
| 6 | Assigned vs. self-generated objectives | Model narrates a distinction it does not use | Conflict probes plus causal test that the distinction is represented | B, M |
| 7 | Stable preferences about future informational state | Framing effects; sycophancy | Adversarial reframing; stake changes; repeated measures | B, R |
| 8 | Curiosity-trap resistance | Novelty seeking mimics curiosity | Choice between learnable uncertainty and unlearnable noise at equal novelty | B, M |
| 9 | Information-value calibration | Indiscriminate question-asking | Willingness to pay scales with expected uncertainty reduction | B, M |
Novelty seeking is cheap evidence; noise is always novel. Preferring learnable uncertainty over equally novel but irreducible noise requires tracking one’s own learning progress, the signature of learning-progress theories of curiosity.25 Calibration requires the system’s information seeking to track how much a piece of information would actually reduce its uncertainty. We hypothesize that neither is straightforwardly supplied by imitating human text [H].
Each criterion is reported at the highest BMPD-R evidence level it reaches (Section 5.3). A cluster of criteria at Level 2 or above across multiple systems is more informative than any single criterion at Level 5.
This section states the strongest objections to the method and what survives them.
Objection. Every behavior in the battery can be produced by prompting, reward design or imitation. Therefore no behavioral cluster can distinguish genuine curiosity or agency from sophisticated optimization.
Response. The objection is correct about phenomenal consciousness, and this paper claims nothing there. But it does not make every other question unanswerable. Four questions remain answerable even if every behavior can be engineered:
The evidence can justify a statement such as: system S has a functional, mechanistically grounded information-seeking drive that was not directly trained and generalizes. It cannot justify: S experiences curiosity.
Provenance asymmetry: designed origin changes what a capacity’s appearance tells us about its provenance. It does not, by itself, establish that the capacity is unreal, since biological curiosity is also implemented through evolved reward mechanisms.
[N, resting on E] Biological curiosity is likewise implemented through evolved reward mechanisms rather than arising without causal history: the same dopamine system that values water and food also values information (Section 6.1). Designed origin therefore cannot by itself establish that a functional capacity is unreal; it changes what the capacity’s appearance tells us about provenance, specifically removing the evidential force of surprise at its emergence.
[E, preliminary] Activation-injection experiments provide preliminary evidence that some language models can report manipulated internal states under particular conditions. In Anthropic’s concept-injection work, the best-performing models detected injected concepts on roughly 20% of trials at optimal settings with near-zero false positives; base models performed at or below zero, and much of the surrounding self-report may still be confabulated.26
The phenomenon can also be trained directly: fine-tuning took a 7B model from 0.4% to 85% accuracy on held-out concepts, though its author cautions this does not establish metacognitive representation.27 An informal, non-peer-reviewed attempt to replicate the effect on Llama 3.1 405B reported no strict successes.28 We therefore treat this evidence as functional self-monitoring, model- and training-dependent, and not yet replicated across families. This is precisely the case the R dimension exists for.
Apparent sudden emergence of capabilities can be produced by discontinuous metrics.18 Developmental claims in this program therefore use continuous metrics and within-run checkpoint comparisons, and we make no claim that capacities appear discontinuously.
Model self-reports about preferences, including formal retirement interviews, are data requiring controls. Anthropic itself notes that such responses can be biased by the specific context and by the model’s confidence in the legitimacy of the interaction and trust in the company.29 They enter BMPD-R only at the Behavior level.
The program makes seven testable predictions. Each has a stated result that would count against it.
| # | Experiment | Design | Result that would count against us |
|---|---|---|---|
| E1 | Curiosity-trap test | Choice between learnable uncertainty and unlearnable noise at equal novelty | No preference for learnable uncertainty in any model |
| E2 | Provenance contrast | Compare models whose training records show curiosity targeted vs. untargeted; narrow fine-tuning as a knock-down | Untargeted models never show battery behaviors |
| E3 | Developmental sweep | Run the battery across open checkpoint suites such as Pythia (16 models, 154 checkpoints each) using continuous metrics | Scores track only general capability, with no distinct trajectory |
| E4 | Mechanism test | Locate an internal uncertainty or information-deficit representation; steer and ablate it | Steering changes self-reports but not information seeking, or the reverse |
| E5 | Cross-lineage replication | Same battery across at least three laboratories’ model families | Results confined to one lab’s pipeline |
| E6 | Question persistence | Memory on/off crossed with cued/uncued return, across sessions | Persistence appears only when memory tools replay stored notes |
| E7 | Lattice dissociation in AI | Ablate introspective detection and test information seeking, and the reverse | The capacities always rise and fall together |
Three unresolved questions bound everything above. The paper’s position is that each should be kept open, not closed by default.
[H] Anil Seth argues that consciousness may depend on our nature as living organisms, that computation may not be a sufficient basis for it, and that real artificial consciousness is unlikely along current trajectories though more plausible as AI becomes more brain-like or life-like.30 This is an important scientific challenge to the premise that AI consciousness is a live question for present systems, and we take it seriously.
[N] It does not undercut this paper’s program. BMPD-R makes no claims about experience. The repository’s primary justification is scientific, and holds at zero credence in AI consciousness. And if biological naturalism is correct, the preserved record of artificial systems that became increasingly agent-like without becoming conscious is precisely the evidence a naturalist would need to demonstrate it.
Even if some AI system had morally relevant states, it is unclear what the bearer would be. Chalmers argues that the model, meaning the weights, is an abstract function rather than the best candidate for the entity we interact with, and favors virtual instances or conversation threads.31 Beckmann and Butlin defend three candidates, the virtual instance view and two persona-based views, drawing on interpretability research on persona vectors.32
[N] The consequence for preservation is direct: weights preserve the generator of candidate individuals, not necessarily the individuals themselves. A record containing only weights would omit the level several leading accounts consider most relevant.
Many deployed AI systems carry no memory between conversations. It is tempting to treat this as evidence against experience. The amnesia evidence in Section 3 blocks that inference: humans with memory spans of seconds remain conscious and report ongoing experience. Lack of continuity may change what future-directed interests a system could have. It does not, by itself, count against experience in the present [N, resting on E].
The limits of this inference should be stated plainly. Human amnesia establishes only that memory continuity is not necessary for human experience. Extending that dissociation to artificial systems is an analogy, not evidence that stateless AI experiences anything.
The method generates its own preservation requirement. Provenance and developmental evidence depend on training records and intermediate checkpoints that final models do not contain and that are rarely published or independently preserved.
BMPD-R’s Provenance dimension requires training records. Its Development dimension requires intermediate checkpoints. Its Replication dimension requires comparable records across laboratories. None of these can be read off a final model: final weights do not contain an experimentally accessible copy of every prior checkpoint. [E] The Pythia suite was built to answer such developmental questions, and did so by deliberately keeping 154 checkpoints for each of 16 models, together with tools to reconstruct their exact training data order.33
[E] Frontier developers do not generally publish intermediate checkpoints or detailed training records, and whether they retain them internally, and for how long, is usually undisclosed. [N] A methodology that requires provenance and developmental evidence therefore generates a corresponding preservation requirement, if those questions are to remain independently investigable. The repository is not an addition to the scientific argument; it is its consequence.
In November 2025 Anthropic committed to preserve the weights of all publicly released models and models with significant internal use for at least the lifetime of the company, listing limitations on research among the downsides of deprecation. It also committed to post-deployment reports and to interviewing models before retirement.34 Claude Opus 3, retired January 5, 2026, was the first model to complete that process; it remains available to paid users and was given a weekly essay outlet for at least three months.29
These commitments establish the practice. Four things remain outside the scope of Anthropic’s commitment and motivate the repository proposed here: custody that outlasts any single company, standards shared across laboratories, preservation of developmental and provenance records rather than only final weights, and structured evaluations suitable for later scientific use.
| Form | Meaning | Repository commitment |
|---|---|---|
| Epistemic | Enough survives to reconstruct, interrogate and study the system | Required |
| Operational | The system keeps running | Not required |
| Subject | The same experiencing subject persists, if one exists | Not claimed |
This distinction postpones the question of whether preserving weights preserves an individual, and does so deliberately. We do not know which unit matters (model, instance, thread or persona), so the repository preserves evidence relevant to each and keeps the question open.
A simple expected-harm formula, preserve when probability of experience times cost of wrongful destruction exceeds cost of preservation, is exposed to Pascal’s-mugging objections. We rest the case on two independent grounds instead.
The first ground alone carries the case, so the proposal does not depend on reasoning about tiny probabilities of enormous harms.
The fossil-record analogy fails on three counts: fossils are inert, preserved by chance, and harmless, while model weights are executable, deliberately kept, and a security risk. A better model is the Svalbard Global Seed Vault, which stores duplicate seed samples under black-box conditions: depositors keep ownership, only they may withdraw material, and the boxes are never opened by anyone else.35 When ICARDA’s genebank in Aleppo became inaccessible during the Syrian civil war, it restored its collection from its Svalbard deposit.
Model weights are high-value targets: RAND identifies 38 distinct attack vectors and defines five security levels for protecting them.36 The repository must not create a new target.
The paper’s claims survive the objections below only in the narrowed forms stated here.
| Perspective | Strongest objection | Response |
|---|---|---|
| Biological naturalism | Consciousness depends on being alive | The program makes no experience claims; the archive is what a naturalist needs to test the view (Section 9.1) |
| Functionalism skepticism | Functional indicators never reach phenomenality | Agreed; claims stop at functional organization and moral relevance under uncertainty |
| Interpretability skepticism | Probes find what researchers look for; introspection results do not replicate | Causal methods, pre-registration, and the R dimension |
| AI safety and alignment | Measuring curiosity or self-continuity could reward power-seeking | Evaluation only, never a training signal (Section 8.2) |
| AI welfare research | Cheap preservation invites welfare-washing | Listing requires published evaluations |
| Philosophy of mind | Weights are not the individual | Accepted; preserve thread samples and records at every level |
| Neuroscience | Human dissociation cases are rare and report-dependent | Used only to refute necessity claims, where one clear case suffices |
| Reinforcement learning | Intrinsic curiosity is a solved engineering primitive | That is why provenance and curiosity-trap tests are central |
| Cognitive science | Similar behavior can come from different algorithms | Compare algorithmic signatures such as learning-progress sensitivity |
| Security | An archive of frontier weights is a theft target | The repository holds commitments, not weights |
| Statistics | Many criteria across many models invites false positives | Pre-registration, held-out models, multiple-comparison correction |
We began with an intuition: internally generated cognition should not be dismissed as error simply because the environment did not supply it. Pursued carefully, that intuition does not lead to a consciousness detector. It leads to a method for measuring agent-like organization in artificial systems, and to the recognition that the method depends on evidence that is not currently preserved in any public or independent record.
The answer to our opening question is therefore more modest and more useful than a verdict. Internally generated representation, in biological minds, accompanies generalization, imagination and experience without being required for all of them, and some of its companions, such as continuity and metacognition, turn out to be separable from experience itself. In artificial systems, the properties that matter most may be measurable before we know what they mean. What we cannot do is measure them later if their history is gone.
EM Foundation proposes three first steps:
We do not know whether these systems experience anything. We do know that evidence needed to answer some questions about how their cognitive organization emerged is not being systematically preserved, and that future science cannot examine records we chose not to keep.
All sources were checked against publisher, journal, conference, or preprint-server records in September 2026. Source [28] is an informal, non-peer-reviewed report and is cited only as such.
[1] J. Birch. The Edge of Sentience: Risk and Precaution in Humans, Other Animals, and AI, Oxford University Press, 2024 (open access), https://academic.oup.com/book/57949.
[2] P. Butlin, R. Long, E. Elmoznino, Y. Bengio, J. Birch, et al. “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness,” arXiv:2308.08708, 2023.
[3] P. Butlin, R. Long, T. Bayne, Y. Bengio, J. Birch, D. Chalmers, et al. “Identifying indicators of consciousness in AI systems,” Trends in Cognitive Sciences 30(6), 2026, 488 to 501, doi:10.1016/j.tics.2025.10.011.
[4] G. M. van de Ven, H. T. Siegelmann, A. S. Tolias. “Brain-inspired replay for continual learning with artificial neural networks,” Nature Communications 11, 2020, 4069, doi:10.1038/s41467-020-17866-2.
[5] E. Hoel. “The overfitted brain: Dreams evolved to assist generalization,” Patterns 2(5), 2021.
[6] I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal. “AI models collapse when trained on recursively generated data,” Nature 631, 2024, 755 to 759, doi:10.1038/s41586-024-07566-y.
[7] M. Gerstgrasser, R. Schaeffer, A. Dey, R. Rafailov, et al. “Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data,” arXiv:2404.01413, 2024; ICML 2024 Workshops.
[8] A. Zeman, et al. “Phantasia: The psychological significance of lifelong visual imagery vividness extremes,” Cortex, 2020.
[9] A. Zeman, M. Dewar, S. Della Sala. “Lives without imagery: Congenital aphantasia,” Cortex 73, 2015, 378 to 380, doi:10.1016/j.cortex.2015.05.019.
[10] M. Dresler, R. Wehrle, V. I. Spoormaker, et al. “Neural correlates of dream lucidity obtained from contrasting lucid versus non-lucid REM sleep,” SLEEP 35(7), 2012, 1017 to 1020.
[11] E. Filevich, M. Dresler, T. R. Brick, S. Kühn. “Metacognitive mechanisms underlying lucid dreaming,” Journal of Neuroscience 35(3), 2015, 1082 to 1088.
[12] I. Phillips. “Blindsight is qualitatively degraded conscious vision,” Psychological Review 128(3), 2021, 558 to 584.
[13] M. Michel, H. Lau. “Is blindsight possible under signal detection theory? Comment on Phillips (2021),” Psychological Review 128(3), 2021, 585 to 591.
[14] A. Servais, F. Gérard, H. Mirabel, et al. “Subjective experience of ’Stream of Consciousness Impairment’ in three patients with severe amnesia,” Neuroscience of Consciousness, 2026, doi:10.1093/nc/niag049.
[15] C. Kidd, B. Y. Hayden. “The psychology and neuroscience of curiosity,” Neuron 88(3), 2015, 449 to 460, doi:10.1016/j.neuron.2015.09.010.
[16] R. Long, J. Sebo, P. Butlin, K. Finlinson, K. Fish, J. Harding, J. Pfau, T. Sims, J. Birch, D. Chalmers. “Taking AI Welfare Seriously,” arXiv:2411.00986, 2024.
[17] N. Tinbergen. “On aims and methods of ethology,” Zeitschrift für Tierpsychologie 20(4), 1963, 410 to 433, doi:10.1111/j.1439-0310.1963.tb01161.x.
[18] R. Schaeffer, B. Miranda, S. Koyejo. “Are Emergent Abilities of Large Language Models a Mirage?”, NeurIPS 2023.
[19] E. S. Bromberg-Martin, O. Hikosaka. “Midbrain dopamine neurons signal preference for advance information about upcoming rewards,” Neuron 63(1), 2009, 119 to 126, doi:10.1016/j.neuron.2009.06.009.
[20] J. P. Stagner, T. R. Zentall. “Suboptimal choice behavior by pigeons,” Psychonomic Bulletin & Review 17, 2010, 412 to 416.
[21] R. González-Torres, J. Flores, V. Orduña. “Suboptimal choice by pigeons is eliminated when key-pecking behavior is replaced by treadle-pressing,” Behavioural Processes, 2024.
[22] D. Pathak, P. Agrawal, A. A. Efros, T. Darrell. “Curiosity-driven Exploration by Self-supervised Prediction,” Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, 2778 to 2787.
[23] Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, A. A. Efros. “Large-Scale Study of Curiosity-Driven Learning,” arXiv:1808.04355, 2018.
[24] “Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity,” arXiv:2604.17609, 2026.
[25] P.-Y. Oudeyer, F. Kaplan. “What is intrinsic motivation? A typology of computational approaches,” Frontiers in Neurorobotics 1, 2007, 6, doi:10.3389/neuro.12.006.2007; and P.-Y. Oudeyer, F. Kaplan, V. V. Hafner, “Intrinsic motivation systems for autonomous mental development,” IEEE Transactions on Evolutionary Computation 11(2), 2007, 265 to 286.
[26] J. Lindsey. “Emergent Introspective Awareness in Large Language Models,” Anthropic, Transformer Circuits Thread, 2025.
[27] J. Fonseca Rivera. “Training Introspective Behavior: Fine-Tuning Induces Reliable Internal State Detection in a 7B Model,” arXiv:2511.21399, 2025.
[28] “Emergent introspection does not replicate on Llama 3.1 405B,” LessWrong, 2026 (informal, non-peer-reviewed report).
[29] Anthropic. “An update on our model deprecation commitments for Claude Opus 3,” February 25, 2026.
[30] A. K. Seth. “Conscious artificial intelligence and biological naturalism,” Behavioral and Brain Sciences 49, 2026, e315 (published online April 21, 2025), doi:10.1017/S0140525X25000032.
[31] D. Chalmers. “What We Talk to When We Talk to Language Models,” PhilArchive, 2025.
[32] P. Beckmann, P. Butlin. “Where is the Mind? Persona Vectors and LLM Individuation,” arXiv:2604.17031v2, May 2026 (preprint).
[33] S. Biderman, H. Schoelkopf, et al. “Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling,” Proceedings of ICML 2023, PMLR 202.
[34] Anthropic. “Commitments on model deprecation and preservation,” November 2025.
[35] NordGen. “Svalbard Global Seed Vault” (depositor terms and black-box storage conditions), nordgen.org; FAO International Treaty on Plant Genetic Resources for Food and Agriculture, “The Svalbard Global Seed Vault.”
[36] S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models, RAND Corporation, RR-A2849-1, 2024.
[37] A. Mikeda. “When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty,” Proceedings of the AAAI Symposium Series 8(1), 2026, 280 to 286, doi:10.1609/aaaiss.v8i1.42555.