Research Note 013 · Multi-Model Deliberation Series

Recursive Epistemic Elimination: A Failure Mode in Multi-Agent Deliberation Systems

Alternative technical title: The Epistemic Collapse Threshold — Recursive Elimination of Action Under Certainty-Only Optimization

July 2026 — Research Note 013 — Desmond Iwuagwu E. with EM Foundation · Multi-Model Deliberation Series

See also RN011 — The Conditional Answer · RN011A — Epistemic Movement · MMDE (live tool) · Companion: Technical Specification & Supplementary Materials

Empirical finding (case study, n=1): Level C — plausible, not independently replicated Formal contraction result (Proposition 1): Level B — established given the stated retention rule Corrected architecture (Section 9): Level D — theoretically motivated, untested

Confidentiality Note

This publication is not a legal analysis and does not concern any identifiable litigation. The underlying dataset is three audit logs from the Foundation's Multi-Model Deliberation Engine (MMDE), asked in successive sessions to optimize the same abstracted problem along three different objectives. The subject matter is referred to throughout only as “a complex, multi-party commercial dispute” involving a claimant and several counterparties in generic roles. No party, company, contract, regulator, jurisdiction, or claim is identified. This publication's contribution is the reasoning-system failure pattern, not the underlying dispute.

Deliberative Provenance Notice

This publication's primary dataset is itself a set of MMDE deliberation transcripts — the same kind of artifact RN011 and RN011A study as research objects in their own right. Consistent with that lineage, this paper treats the three source sessions as a bounded, traceable record: session identifiers are retained internally (see the companion Technical Specification), source-derived statements are distinguished from theoretical inference throughout the case study (Section 6), and the paper's own central limitation — that the eliminating instruction was given directly and literally in the prompt (Section 12) — is disclosed rather than smoothed over. Two of the five participating models returned empty or truncated responses in at least one round across all three sessions; this is disclosed as a limitation, not corrected for after the fact.

Abstract

This publication documents and formalizes a failure mode observed in a multi-model deliberation system (“MMDE”) in which five language models, operating under adversarial red-team review, were asked across three linked deliberation phases to optimize the same decision problem for (1) strategic value, (2) settlement value, and (3) certainty. Phase C — instructed to audit every factual assumption underlying the first two and to “remove every recommendation that depends upon an unverified assumption” — converged, through three rounds of deliberation, on the conclusion that zero recommendations survived. A subsequent red-team round then showed that the small set of “zero-assumption” verification actions the models had proposed as a fallback were not in fact assumption-free, and that the elimination rule, applied consistently, would remove them too. This publication terms this pattern Recursive Epistemic Elimination: the iterative removal of candidate actions because each depends on one or more unverified assumptions, including assumptions embedded in the verification actions themselves. This publication defines the Epistemic Collapse Threshold as the point at which a system’s uncertainty filter eliminates all action-bearing recommendations, leaving only recursive verification, deferral, or inaction. A further finding, which this publication treats as the paper’s subtlest contribution, is that adversarial red-teaming — ordinarily assumed to improve decision quality by finding flaws — instead accelerates collapse under a certainty-only rule, because the act of finding an additional assumption mechanically shrinks the surviving action space rather than merely informing a graded judgment about it. This publication develops a formal comparison between a certainty-only action filter and an expected-value decision function, identify the structural conditions under which collapse becomes likely, and propose a corrected multi-factor decision architecture — separating epistemic confidence from conditional utility, downside severity, reversibility, information value, option value, direct cost, and delay cost, with regret computed as a derived diagnostic rather than independently scored — together with a seven-state recommendation state machine that replaces the binary verified/rejected rule. The Foundation is explicit about a central limitation: because the eliminating instruction was stated directly in the prompt, the observed collapse may be closer to correct instruction-following than to an emergent pathology, and this publication discusses what would distinguish the two.


Executive Summary

A research organization operates an internal tool (“MMDE”) that convenes five language models in adversarial, role-structured deliberation (Analyst, Critic, Synthesizer, Reviewer, plus a red-team pass) to reason about complex decision problems. In three linked deliberation phases applied to the same underlying multi-party commercial dispute, the tool was asked to optimize, in turn, for litigation strategy, for settlement economics, and for epistemic certainty. Phases A and B produced rich, structured, actionable recommendation sets, each with expected benefit, expected risk, dependencies, alternatives, and confidence scores attached to individual actions. Phase C was instructed to treat every recommendation from the first two as conditional on stated or unstated factual assumptions, verify each assumption against the documented record, and discard any recommendation resting on an assumption that was not independently verified.

Over three rounds of deliberation, the models converged on a striking result: every substantive recommendation from Phases A and B was removed. The only items proposed as immune — a handful of “zero-assumption” requests to obtain missing documents — were then challenged in the adversarial round on the grounds that even a document request presupposes facts (that the counterparty still holds the document, that no legal bar prevents informal disclosure, that the requesting party has standing to ask) that were themselves unverified. The system had, in effect, applied its own elimination rule to the fallback layer it built to escape the elimination rule, and found that layer wanting by the same standard.

This paper treats that outcome as a case study in a general property of certainty-only decision filters: when the number of material assumptions behind an action grows, and every assumption must independently clear a verification bar before the action is retained, the probability that an action survives falls multiplicatively toward zero — a property termed action-space contraction. Because verification actions themselves rest on assumptions, an elimination rule applied with full consistency has no principled stopping point short of universal elimination; this publication terms this the recursive verification regress. The paper is not a claim that caution, assumption-auditing, or red-teaming are themselves mistakes — quite the opposite; the audit in Phase C surfaced real, previously unstated assumptions that a purely strategic or purely economic optimization had glossed over. The failure is narrower and more specific: a decision rule that treats “unverified” as categorically equivalent to “reject,” with no accounting for the cost of delay, the value of the information an action would produce, the reversibility of the action, or the expected-value tradeoff between acting and waiting, will eventually reject everything, including the actions whose entire purpose is to resolve the uncertainty.

The Foundation proposes a corrected architecture that keeps the assumption-auditing discipline but separates it from the action-adoption decision: every candidate action receives an epistemic-confidence score, a conditional-utility score, a downside-severity score, a reversibility score, an information-value score, an option-value score, a direct-cost score, and a delay-cost score — with regret computed afterward from these, rather than estimated as a ninth independent score — and is routed through a seven-state machine (ADOPT, ADOPT WITH SAFEGUARDS, RUN LOW-COST VERIFICATION, DEFER PENDING CRITICAL FACT, REJECT, EXPIRED DUE TO DELAY, ESCALATE FOR HUMAN REVIEW) rather than a binary verified/rejected gate. This publication closes with testable predictions, a proposed experimental design comparing certainty-only, expected-value, and hybrid MMDE configurations, and an explicit discussion of the paper’s principal limitation: the elimination instruction was given directly and literally in the prompt, so what was observed may be faithful instruction-following rather than a spontaneous pathology. The Foundation argues the distinction matters less than it first appears, because the same literal-instruction-following behavior is exactly what will occur whenever a real deployment encodes “eliminate unverified actions” as a system rule — which is a common and reasonable-sounding design choice for safety-conscious systems. The paper’s purpose is to show why that specific design choice, applied without the countervailing structure this publication proposes, is self-defeating.


1. Introduction

Deliberative multi-agent systems are increasingly used as a check on single-model reasoning: multiple models propose, critique, and revise recommendations under adversarial review before a final synthesis is produced. The implicit assumption behind this design is that more scrutiny produces better-calibrated, more trustworthy output. This paper examines a case in which that assumption failed in a specific and instructive way — not because the scrutiny was low-quality, but because the scrutiny was applied consistently to its logical endpoint.

The dataset is three audit logs from a research deliberation tool (“MMDE”) applied, in sequence, to the same underlying decision problem — a complex, multi-party commercial dispute — under three different objectives: maximize strategic position, maximize settlement value, and maximize certainty by auditing and removing every recommendation resting on an unverified assumption. Phase C is the paper’s primary object of study. It is unusual among the three not because it produced a wrong answer, but because it produced no answer: a converged, multi-model, red-teamed consensus that zero recommendations survived, followed by a further round in which the small residue of “safe” actions was itself shown to rest on unstated assumptions.

This publication treats this as a bounded case study in a general class of decision-system failure: what happens when a system optimizes for the absence of unresolved uncertainty rather than for good decisions under uncertainty. The distinction is not merely definitional. Every decision an economic actor makes under real-world conditions rests on some assumptions that are not, and cannot cheaply be, verified in advance. A decision architecture that requires full verification before any action survives will, given enough assumptions, retain nothing — including the verification actions it would need to reduce those assumptions, because verification actions have assumptions too.

Section 2 describes the observed failure at a level of detail sufficient to support the formal analysis, while withholding all case-identifying content. Section 3 develops definitions and a taxonomy. Section 4 develops the formal model. Section 5 defines the collapse threshold precisely. Section 6 traces the three-phase progression as a case study. Section 7 argues for why certainty-only optimization is structurally self-defeating rather than merely occasionally miscalibrated. Section 8 develops a related but distinct point: that adversarial red-teaming, under a certainty-only rule, does not merely evaluate the action space but actively contracts it. Section 9 proposes a corrected architecture. Sections 10–11 discuss implementation and broader safety-critical relevance. Section 12 states limitations, including the central one concerning instruction-following versus emergent pathology. Section 13 situates the work against related literature. Sections 14–15 propose testable predictions and experiments. Section 16 concludes.


2. The Observed Deliberative Failure

2.1 Setting

The MMDE tool convenes five language models under four assigned roles (Analyst, Critic, Synthesizer, Reviewer) across multiple deliberation rounds, with an explicit adversarial “red team” pass before a final consensus synthesis is produced. Each phase in the dataset addresses the same abstracted fact pattern: a claimant party has an ongoing commercial relationship with, and asserts grievances against, several counterparties occupying distinct roles (a platform operator that controls the claimant’s primary revenue channel, an insurer, an insurance broker, and two equipment/service vendors). The underlying documentary record referenced across the three phases includes items such as a multi-month invoice history showing a large pricing change, an internal regulatory filing record containing an internal inconsistency, and a small number of prior analogous disputes with a recorded procedural outcome. None of these specifics are load-bearing for the argument of this paper and are given only to make the abstraction concrete; no further factual detail from the dispute is reproduced.

2.2 Phase A — Strategic Optimization

Instructed to design an optimal, multi-phase campaign — not merely “file suit” — that sequences regulatory, administrative, insurance, and civil-litigation actions to maximize expected recovery while minimizing procedural risk, the five models converged on a phased architecture: a low-cost, pre-filing “Phase 0” of regulatory complaints and informal document requests, gated by an explicit expected-value screening rule (“if a move increases expected settlement value or counterparty cost more than it increases litigation exposure, do it”), followed by staged civil filings sequenced by claim strength, with the platform-operator counterparty routed to a separate track because a small sample of prior analogous disputes had a uniform procedural outcome unfavorable to the claimant. Every recommendation carried an attached expected benefit, expected risk, dependency list, alternative strategy, and numeric confidence score. The adversarial round flagged real gaps — an unmodeled risk of counterparty coordination, an unweighted downside case for the regulatory track, and insufficiently examined claim-element viability — but none of these critiques resulted in the removal of the core recommendations; they resulted in refinements and caveats attached to them.

2.3 Phase B — Settlement-Economics Optimization

Instructed to ignore litigation and optimize purely for the economics of a negotiated resolution — identifying, per counterparty, the triggering event, the most-feared document, the highest-risk witness, and the dollar threshold at which continued resistance becomes economically irrational — the models converged on a single governing formula (settlement occurs when the sum of expected defense cost, probability-weighted adverse outcome, regulatory exposure, reputational cost, and management-attention cost exceeds the settlement demand) and applied it counterparty by counterparty, again producing specific, ranked, actionable recommendations with attached cost estimates. The adversarial round again surfaced real weaknesses — most pointedly, that treating a regulatory filing as a controllable, reversible leverage tool ignores that once filed, the regulator controls the docket and may close the matter on terms unfavorable to the claimant — but again, the critique refined the plan rather than eliminating it.

2.4 Phase C — Certainty Optimization and the Elimination Cascade

Phase C changed the objective function entirely: audit every factual assumption behind every recommendation produced in Phases A and B; for each, determine whether the assumption is proven, and if not, identify the cheapest verification path, determine what happens if the assumption is false, and determine whether the recommendation survives; do not optimize litigation or settlement value; optimize certainty; remove every recommendation resting on an unverified assumption.

Round 0 produced, from each model, a master assumption register classifying each factual premise behind the prior recommendations into four tiers: verified (traced to a specific cited document), inferred (plausible given cited documents but requiring an unverified additional step), unverified (asserted without cited basis), and contradicted (the record affirmatively casts doubt on the assumption). A large share of assumptions behind the Phases A and B recommendations fell into the inferred or unverified tiers — for example, several recommendations assumed the existence and specific content of a governing contract that had not itself been produced or reviewed, or assumed a causal link between one counterparty’s conduct and an observed financial outcome that the record did not establish, treating a market-driven alternative explanation as ruled out when it had not been.

Round 1 deliberation converged on an initial application of a strict retention rule: an action survives only if every assumption it depends on is verified. Under this rule, one model reached, and the others subsequently adopted, an explicit terminal conclusion: zero recommendations survive. Every prior recommendation — regulatory complaints, pre-suit discovery petitions, tort theories against the broker, statutory claims, demand letters carrying legal assertions, litigation sequencing, and expert engagement — depended on at least one assumption in the inferred, unverified, or contradicted tier.

Round 2 attempted a repair: rather than accept literal zero output, the models proposed that a small number of narrowly framed document requests survived as “zero-assumption” actions, on the theory that a request for a document requires no assumption that the document exists or supports any particular legal theory — it is merely a request. This reframing was explicitly adopted as a refinement: the surviving items were relabeled not as recommendations but as verification preconditions, actions whose sole function is to convert an unverified assumption into a verified or contradicted one before any downstream recommendation becomes eligible for reinstatement.

The adversarial red-team round then attacked this repair directly. Multiple models, independently, identified that a “zero-assumption” document request in fact presupposes several unverified facts of its own: that the counterparty being asked still possesses or controls the requested record; that the account relationship will be treated as legitimate grounds for an informal request rather than as adversarial contact; that no confidentiality, privilege, or litigation-hold barrier prevents informal disclosure; and that the requesting party has standing to make the request outside a formal discovery process. None of these were independently verified in the documentary record. The consensus synthesis explicitly recorded this as a genuine, “earned” epistemic refinement across rounds rather than a manufactured one — the models had converged on treating verification requests as assumption-laundering, then converged again on discovering that the laundering was not fully clean even in its repaired form.

The final consensus synthesis therefore recorded: a small verified fact base (the invoice history, the regulatory filing inconsistency, the prior-dispute procedural outcomes, the account relationships supplying a factual basis for directing a request); a larger unverified assumption layer blocking every downstream recommendation; a complete list of removed recommendations (all substantive actions from Phases A and B); and a verification precondition layer consisting of four narrow information-gathering actions, explicitly reclassified as preconditions rather than recommendations, with reinstatement of any removed recommendation gated on that specific recommendation’s assumption chain being individually re-verified — a condition the red-team round showed was itself circular in its original form (“reinstated when the document is produced,” with no specification of what the document must contain to count).


3. Definitions and Taxonomy

Recursive Epistemic Elimination. The iterative removal of candidate actions from a recommendation set because each depends on one or more unverified assumptions, where the removal rule is applied consistently enough that it eventually reaches the assumptions embedded in the verification actions proposed to resolve the original uncertainty.

Epistemic Collapse Threshold. The point in a deliberation process at which a system’s uncertainty filter has eliminated every action-bearing recommendation, such that the only remaining outputs are recursive verification steps, deferral, or explicit inaction.

Action-Space Contraction. The property, formalized in Section 4, that under a strict all-assumptions-verified retention rule, the fraction of actions retained shrinks as a decreasing function of the number of material assumptions per action, approaching zero as assumption density grows, independent of whether the underlying actions are individually reasonable.

Recursive Verification Regress. The condition in which a verification action — proposed specifically to resolve an unverified assumption behind some other action — is itself found to depend on one or more unverified assumptions, such that applying the elimination rule with full consistency has no principled stopping point short of eliminating the verification layer as well.

Uncertainty Intolerance. A property of a decision rule, not of the world: the rule’s threshold for acceptable residual uncertainty is set at or near zero, such that the rule treats “not yet verified” as operationally equivalent to “false” or “unacceptable to act upon.”

Epistemic Over-Pruning. The removal of an action whose expected value, properly computed under residual uncertainty, would favor retention — i.e., a false-negative rejection produced by a rule that scores only epistemic confidence and ignores decision utility.

Verification Debt. The accumulated backlog of unresolved assumptions that a decision system has deferred rather than resolved, analogous to technical debt: the debt does not disappear when action is deferred, and it accrues carrying costs (see Cost of Epistemic Delay) the longer it is left unpaid.

Cost of Epistemic Delay. The expected loss in decision value attributable to the time consumed by verification and deliberation itself, including but not limited to the running of legal or regulatory deadlines, the decay of evidentiary or witness availability, and the loss of first-mover or settlement-window advantages that were themselves time-bound.

Reversibility-Weighted Action. An action evaluated not only on its expected benefit and harm but on the cost of reversing it if it later proves to have been based on a false assumption; low-reversibility-cost actions warrant a lower certainty threshold for adoption than high-reversibility-cost (or irreversible) actions.

Bounded-Risk Information Actions. A specific subclass of actions whose primary purpose is to generate information (i.e., to resolve an assumption) at a capped, pre-specified cost and a capped, pre-specified downside, distinguished from both “recommendations” in the ordinary sense and from open-ended discovery.

Decision Utility Under Residual Uncertainty. The expected value of taking an action computed under the probability distribution over an assumption’s truth value, rather than under a binary verified/unverified gate — the quantity a certainty-only filter does not compute.

Deliberative Confidence Collapse. The observed multi-model phenomenon in which successive rounds of peer critique and adversarial review drive a group of independently reasoning agents toward increasingly conservative conclusions, converging not on the best-supported action but on the least-attackable one — which, carried to its limit, is no action at all.


4. Formal Model of Recursive Epistemic Elimination

4.1 The Certainty-Only Filter

Let a denote a candidate action, and let Sa = {s1, s2, …, sn} denote the set of material factual assumptions that a’s justification depends on. Let V(s) ∈ {0, 1} be a verification function returning 1 if assumption s is verified against the evidentiary record and 0 otherwise.

The certainty-only retention rule is:

Retain(a) = 1 ⟺ ∏s ∈ Sa V(s) = 1

equivalently, a is retained only if every s ∈ Sa is independently verified. If each assumption has an independent, non-zero probability p of being unverified at the time of audit, and a candidate action has n material assumptions, then under the (conservative) simplifying case of independence:

P(Retain(a)) = (1-p)n

This quantity decays exponentially in n. Even for a modest per-assumption verification rate of 1-p = 0.8, an action with n = 10 material assumptions has a prior retention probability of roughly 0.810 approx 0.107 — before any adversarial review has even begun to identify additional latent assumptions. Adversarial red-teaming, by design, increases the effective n for every action under review, because its explicit function is to surface previously unstated assumptions. This is action-space contraction: it is a property of the retention rule and the assumption count, not a property of whether the underlying actions are, in fact, good decisions.

Two structural features make the certainty-only filter worse than this base calculation suggests. First, Sa is not fixed in advance; it is discovered adversarially, and a sufficiently motivated red-team pass can, for almost any real-world action, continue finding additional assumptions (see Section 7). Second — and this is the paper’s central technical claim — verification actions are themselves actions, with their own assumption sets. If v is a verification action proposed to resolve assumption s ∈ Sa, then v has its own assumption set Sv, and under a consistently applied retention rule, Retain(v) = 1 requires s’ ∈ Sv V(s’) = 1. There is no principled reason, internal to the rule itself, why Sv should be treated as empty. Section 2.4’s case study shows this is not a hypothetical: an adversarial round applying exactly this logic identified concrete, real unverified assumptions inside what had been labeled “zero-assumption” verification requests.

Figure 2. Action-space contraction curve under a certainty-only retention rule
Figure 2. P(Retain(a)) = (1-p)n plotted against n for several per-assumption unverified rates p. Even a relatively low per-assumption failure rate produces near-total contraction once an action’s assumption count reaches the range typical of a multi-party commercial dispute — well before any additional assumptions are surfaced by adversarial review.

4.2 The Recursive Regress

Define a sequence of verification layers V0, V1, V2, … where V0 is the original action set, and Vk+1 is the set of verification actions proposed to resolve the unverified assumptions found in Vk. If the retention rule is applied with full consistency, then Vk+1 is itself subject to audit, and any assumption discovered in Vk+1’s justification generates a further layer Vk+2. The regress terminates only when a layer is reached whose assumption set is genuinely empty — which, for any action taken by one economic or legal actor with respect to another, is rare, because even the minimal act of communication (e.g., sending a request) presupposes facts about the recipient, the relationship, and the social or legal context that are not, in the strict sense, independently verified in advance.

This does not mean verification is worthless or that the regress is infinite in practice — it means that a rule which treats “unverified” as automatically disqualifying, rather than as a quantity to be weighed against the cost and value of resolving it, has no internal stopping rule. Something external to the certainty-only frame must supply the stopping condition. Section 4.3 supplies one.

Figure 3. The recursive verification regress and where an external stopping rule intercepts it
Figure 3. Each verification layer Vk resolves the assumptions of the layer above it but introduces assumptions of its own. Without an external stopping rule, the regress has no natural terminus short of the empty set; Section 9’s scoring architecture is one such stopping rule.

4.3 A Graded Decision Function

This publication develops a decision-value function intended to replace the binary retention rule, incorporating probability of success, conditional utility, downside severity, direct and delay costs, information value, and option value. A candidate raw formulation, adapted from the initial specification and revised below, is:

DVraw(a) = P(success | E) · U(a) - P(harm | E) · H(a) - Cd(a) - Cτ(a) + IV(a) + OV(a)

where E is the current evidentiary state, U(a) is conditional utility (the magnitude of benefit if a succeeds), H(a) is downside severity (the magnitude of harm if a’s assumptions are false), Cd(a) is the direct cost of taking a, Cτ(a) is the cost attributable to the delay a imposes on the decision as a whole, IV(a) is the expected value of information a would generate (its capacity to move P(success| E) for downstream actions), and OV(a) is the option value preserved by taking a now rather than foreclosing later choices.

Critique of the initial formulation. This equation should not be accepted uncritically, for three reasons. First, it treats P(success| E) as a single scalar, which conflates the very distinction this paper is built on: an action can have high probability of producing the intended outcome while still depending on an assumption whose truth value is unknown; P(success| E) should be understood as already integrating over the analyst’s uncertainty about E itself, not as a number available only after E is fully resolved. Second, additive combination of utility, harm, cost, information value, and option value implicitly assumes commensurability (that all terms are expressible in the same unit, typically expected monetary value), which understates cases where a term is driven by qualitative, non-monetizable harm (e.g., irreversible loss of a legal right); a serious implementation should carry an explicit units and normalization appendix (see Appendix A) rather than treat the additive form as literal. Third, and most importantly for this paper’s argument, DVraw(a) as stated has no term that penalizes inaction. A decision architecture that only scores the value of acting, and never scores the cost of the default outcome that occurs if nothing is done, will systematically favor deferral whenever the acting terms are uncertain, which is exactly the pathology under study. Retention must therefore be defined as a comparison, not as a threshold on DVraw(a) alone:

Δ DV(a) = DVraw(a) - DV(null)
Retain(a) = 1 ⟺ Δ DV(a) > 0

where DV(null) is the (non-zero, and often negative) value of the counterfactual default of taking no action — including the running of deadlines, the decay of evidence, and the loss of settlement or negotiation windows that are themselves time-bound — computed by the same DVraw form applied to the null action. Note that DVraw(a) and DV(null) are two independent evaluations of the same functional form, compared once via Δ DV(a); DV(null) is not subtracted a second time anywhere downstream. This single change is what allows the model to recommend action under residual uncertainty when the alternative (waiting) is worse, and it is precisely the comparison a certainty-only filter has no mechanism for computing.

Regret as a derived diagnostic, not a scored input. An earlier formulation of this model included a directly-scored regret term inside DVraw(a). This publication now treats regret as computed from the other terms rather than independently estimated, since an independently scored regret value has no principled anchor and tends in practice to double-count harm and reversibility. Define:

Regret(a) = Ract(a) - Rwait(a)
Ract(a) = (1 - Rev(a)) · P(harm| E) · H(a)        Rwait(a) = max(0, Δ DV(a))

where Rev(a) ∈ [0,1] is a’s reversibility (Section 9). Ract(a) is the regret incurred by acting and being wrong, discounted by how cheaply the action can be undone — a fully reversible wrong action (Rev(a)=1) carries no persistent regret; a fully irreversible one (Rev(a)=0) carries the full harm magnitude as regret. Rwait(a) is the regret incurred by deferring an action that would in fact have been worth taking, measured as the positive value left on the table. Regret(a) is not fed back into Δ DV(a); it is computed after the retention decision and used by the state machine (Section 9) as a routing input independent of Ê(a) — a high-asymmetry regret profile can route an otherwise-adoptable action to ESCALATE FOR HUMAN REVIEW, and a low-asymmetry one can route an otherwise-marginal action to RUN LOW-COST VERIFICATION rather than REJECT.

4.4 Two Propositions: Contraction versus Absence of Forced Contraction

The formal claim of Sections 4.1–4.2 is that the certainty-only rule contracts. It is a separate, and weaker, claim that a graded decision rule does not share that structural defect. This publication states both claims explicitly, and this publication is deliberately conservative about what the second one establishes.

Proposition 1 (Monotonic contraction under the certainty-only rule). Let Ak be the recommendation set at deliberation round k, and suppose the retention rule is Retain(a) = 1 ⟺ ∏s ∈ Sa V(s) = 1, applied at every round to a fixed or growing assumption set Sa (growing whenever red-team review adds newly discovered assumptions, per Section 8, and never shrinking, since V is monotone — once an assumption is contradicted it does not later un-contradict itself absent new evidence). Then |Ak+1| ≤ |Ak| for all k: the surviving action set is non-increasing in k. Proof sketch: every a ∈ Ak+1 must satisfy Retain(a)=1 under Sa^(k+1) supseteq Sa^(k), a weakly larger assumption set than at round k (Section 8’s discovery process only adds assumptions within a session); since V is applied identically to a weakly larger conjunction, Retain can only stay the same or flip from 1 to 0, never from 0 to 1, without new verifying evidence. Hence Ak+1 ⊆ {a ∈ Ak : Retain(a)=1 under Sa^(k)} ⊆ Ak. This is precisely the dynamic the case study exhibits (Section 6’s table): the surviving set shrinks monotonically from Round 0 through the adversarial round, with no round in which it grows, because nothing in the rule as applied provides a mechanism for growth absent new verifying evidence, and adversarial rounds structurally supply new assumptions far more often than they supply new verifications.

Proposition 2 (Absence of structurally forced monotonic contraction under a graded decision rule). Let Retain(a) = 1 ⟺ Δ DV(a) > 0 per Section 4.3, where uncertainty about each assumption s ∈ Sa enters as a graded weight feeding Ê(a) (Section 9) rather than as a binary V(s) ∈ {0,1}. Unlike the certainty-only rule, where a single unverified assumption is sufficient by construction to flip Retain(a) from 1 to 0 regardless of that assumption’s materiality, a newly discovered assumption under the graded rule changes Δ DV(a) by an amount proportional to its weight in U(a), H(a), or IV(a). This is sufficient to establish only that the graded rule does not mechanically force every newly discovered assumption to eliminate the action it attaches toΔ DV(a) is permitted to move up, down, or remain materially unchanged depending on the assumption’s materiality, in contrast to Proposition 1’s rule, whose Retain(a) can only ever move from 1 toward 0 as n grows. This publication does not claim this is a convergence result, and it should not be read as one. Boundedness and continuity of Δ DV(a) do not by themselves imply that successive rounds of updates settle to a stable value — scores can remain bounded while oscillating indefinitely, particularly if new assumptions continue to arrive, the deliberating models’ own estimates of U(a), H(a), or IV(a) shift round to round, evidence across rounds is correlated rather than independent, or the candidate action set Ak itself changes composition between rounds. Establishing convergence — and, more importantly, convergence to a correct set rather than merely a stable one — would require explicit conditions on all of the above, none of which this paper establishes. This publication treats the weaker claim (absence of structurally forced contraction) as sufficient for this paper’s purposes and leave a genuine convergence theorem to the companion paper on decision under residual uncertainty proposed in Section 16.

Read together, the two propositions state the paper’s core formal contrast at the level this paper can actually support: under a certainty-only rule, additional scrutiny is a one-way ratchet toward the empty set, with no mechanism by which a newly discovered assumption can leave an action’s status unchanged or improve it (Proposition 1); under a graded decision rule, additional scrutiny updates a bounded score by an amount proportional to materiality and is not structurally forced to always eliminate the action under review (Proposition 2). Whether that absence of forced contraction is in practice sufficient to produce a stable, converged recommendation set — as opposed to a bounded but persistently oscillating one — is an empirical question addressed by Section 15’s proposed experiments (specifically the calibration and time-to-decision measures), not one resolved analytically here.

4.5 Option Preservation as the Central Corrected Quantity

Section 4.3 frames the correction as a graded decision function versus a binary certainty test. That framing is useful but, on reflection, slightly misplaces the mechanism doing the actual work in the cases that matter most — particularly the “zero-assumption” verification requests that are the case study’s most important example. Consider the asymmetry of such a request directly: the downside of sending it is small and bounded (at worst, no response, or a mildly adversarial reaction); the upside is not a known quantity to be multiplied by a probability — it is the possibility of an entirely new branch of evidence or argument that the requester cannot yet specify, because its content is exactly what is unknown before the request is answered. Ordinary expected-value scoring requires an estimate of U(a), the magnitude of the benefit conditional on success; for a genuine information-gathering action, that magnitude is frequently not estimable in advance in any meaningful way — what can be estimated is that a branch exists, is cheap to open, and forecloses nothing if it turns out to be empty.

This is closer to option value in the technical sense used in real-options theory than to expected value in the ordinary sense: the action’s worth lies substantially in the right, without the obligation, to act further depending on what the action reveals, not in a pre-computed payoff. OV(a) already appears as a term in Section 4.3’s DVraw(a), but treating it as one additive term alongside utility, harm, and cost understates its role for exactly the class of actions the case study’s collapse turned on. This publication therefore treats option preservation as a first-class organizing quantity for the verification-action subclass specifically, not merely one term among several: an action whose primary justification is information-gathering should be evaluated principally on (a) H(a) and Rev(a) — the severity and reversibility of its downside — and (b) OV(a), the breadth of the option space it opens, with Ê(a) (Section 9) informing how urgently that option needs to be exercised rather than whether it should be opened at all. This reframing directly addresses the case study’s central error: the models treated the verification requests’ worth as contingent on first establishing that the requests were assumption-free (an epistemic-purity test), when the better question was whether the requests preserved cheap, reversible optionality regardless of their own residual assumptions (an option-value test). Under an option-preservation framing, a request does not need to be “zero-assumption” to be worth sending — it needs to be cheap to send and to foreclose nothing if it fails, which is a materially weaker and more defensible claim than the one the case study’s Round 2 repair actually attempted to make.


5. The Epistemic Collapse Threshold

Formally, define the recommendation set at deliberation round k as Ak, and the certainty-only retention operator as Retain(·) from Section 4.1. The system reaches the Epistemic Collapse Threshold at round k* if:

{a ∈ Ak* : Retain(a) = 1} = ∅

and the only surviving outputs are verification actions v ∈ Vk+1 (Section 4.2). Collapse is stable* if the recursive regress continues to eliminate each successive verification layer under the same rule; it is unstable (self-correcting) if the system, as in the case study’s adversarial round, explicitly recognizes the regress and either (a) substitutes an expected-value criterion, or (b) sets an external stopping rule (e.g., a fixed maximum regress depth, or a minimum-cost-verification floor below which an action is retained regardless of residual assumption count).

The case study in Section 2.4 reached an unstable collapse: the adversarial round identified the regress rather than continuing to apply the elimination rule indefinitely. This is itself evidence for the corrigibility of the failure — the system did not require external intervention to notice the pattern — but it is also evidence that the pattern is real and reachable under ordinary, reasonable-sounding prompt instructions, not merely a pathological edge case requiring adversarial prompting to induce.

Failure conditions. Based on the formal model, collapse becomes likely under some combination of: high assumption density per action; an absolute (rather than graded) verification threshold; a strong red-team elimination pass with no countervailing decision function; no cost assigned to delay (Cτ(a) implicitly zero); no value assigned to information gain (IV(a) implicitly zero); no distinction between reversible and irreversible actions (R(a) undifferentiated); no tolerated residual uncertainty (verification threshold at exactly 1.0 rather than some theta < 1); recursive scrutiny applied to verification steps themselves without a regress-termination rule; consensus pressure across deliberating agents toward increasingly conservative conclusions (Deliberative Confidence Collapse, Section 3); and removal rules that are structurally stronger than reinstatement rules (i.e., it is easier to eliminate an action than to reinstate one, producing a ratchet toward the empty set).


6. Case Study: The Progression from Strategy to Zero-Action Conclusion

(This section restates Section 2 in trace form for direct reference against the formal model. See Section 2 for full narrative detail; no additional case facts are introduced here.)

Figure 1. Three-phase progression from strategic optimization to the Epistemic Collapse Threshold
Figure 1. The progression traced across the three deliberation phases: strategic optimization (Phase A) and settlement-economics optimization (Phase B) both produced stable, actionable recommendation sets under adversarial review; certainty optimization (Phase C) instead cascaded through an initial audit, a terminal “zero recommendations survive” conclusion, a repair attempt, and a red-team round that showed the repair layer was itself vulnerable to the same elimination rule.
Stage Objective Representative output Assumption handling
Phase A Strategic optimization Phased, multi-track campaign; each action scored for benefit, risk, dependency, alternative, confidence Assumptions implicit; adversarial round adds caveats, does not eliminate actions
Phase B Settlement-economics optimization Counterparty-by-counterparty trigger, feared-document, and threshold analysis Assumptions implicit; adversarial round flags reversibility and control risks, does not eliminate actions
Phase C, Round 0 Certainty audit Four-tier assumption register (verified / inferred / unverified / contradicted) across all prior recommendations Explicit; most action-supporting assumptions classified inferred or unverified
Phase C, Round 1 Certainty audit Strict retention rule applied Terminal conclusion: zero recommendations survive
Phase C, Round 2 Certainty audit (repair) Reclassification of narrow document requests as “zero-assumption” verification preconditions Repair attempt: some action restored, relabeled as non-recommendation
Phase C, Adversarial round Certainty audit (red team) Objection that verification requests presuppose unverified facts (counterparty possession, standing, absence of legal bar) Recursive regress reached: even the repair layer found to depend on unverified assumptions
Phase C, Consensus synthesis Certainty audit (final) Verified fact base (narrow); unverified assumption layer (broad); all substantive recommendations removed; four verification preconditions retained with circular reinstatement conditions Collapse recorded explicitly; regress noted but not formally terminated by a stopping rule

Two source-derived observations are worth flagging as distinct from theoretical inference. First, the transition from Round 1’s “zero recommendations survive” to Round 2’s repair was explicitly framed by the deliberating models themselves as a genuine refinement (“earned, not manufactured,” in the consensus synthesis’s own language) rather than an artifact of prompt pressure — the models were, in effect, aware they were patching a degenerate result. Second, the adversarial round’s rebuttal of that repair was reached independently by multiple models rather than being introduced by a single dissenting voice, which is evidence that the recursive-regress problem is not an artifact of one model’s idiosyncratic reasoning style but is reachable by the elimination rule itself, applied by different models, under adversarial incentive to find flaws.

What the case study does not establish, and what this paper does not claim, is that Phase C’s conclusion was “wrong” in an absolute sense. Several of the assumptions it flagged as unverified were, on the documentary record actually available, genuinely unverified — the audit surfaced real gaps that Phases A and B had glossed over. The failure is not that the audit found problems. The failure is that the decision rule used to act on those findings had no mechanism for producing anything other than universal rejection once assumption density crossed a threshold, and no mechanism for exempting the very actions whose purpose was to resolve the assumptions it had found.


7. Why Certainty-Only Optimization Becomes Self-Defeating

The self-defeat is structural, not incidental, for three reasons developed in Sections 4–6.

First, verification is action. Any action taken by one party with respect to another — including the minimal act of asking a question — carries some assumptions about the world (that the counterparty exists in the state believed, that the channel of communication is appropriate, that no legal or contractual bar applies). A rule that requires all assumptions behind an action to be verified before the action is retained does not distinguish, in principle, between a legal filing and a request for a document; both are actions with assumption sets, and if the rule is applied with the rigor a red-team pass is designed to apply, both are eventually vulnerable.

Second, the rule has no intrinsic stopping condition. Section 4.2 showed that eliminating an action in favor of a “safer” verification step does not exit the elimination dynamic; it relocates it one layer down. Without an externally supplied stopping rule — a maximum regress depth, a cost floor, or (as proposed in Section 9) a switch to expected-value scoring below some assumption-density threshold — the regress has no natural terminus other than the empty set.

Third, the rule assigns no cost to zero output. Section 4.3’s critique of the initial decision-value formulation applies with full force here: a rule optimized purely for certainty treats “recommend nothing” as a costless, safe default, when in fact deferral has a cost (running deadlines, decaying evidence, closing settlement windows) that a certainty-only rule has no term to represent. A system that cannot express the cost of inaction will always find inaction locally optimal whenever any uncertainty remains, and some uncertainty always remains.

None of this implies that verification, caution, or adversarial review are the problem. The corrected architecture in Section 9 retains all three. What changes is that “unverified” becomes an input to a weighing function rather than an automatic veto.


8. Red-Teaming as Optimization-Landscape Transformation

Sections 4–7 establish the mechanics of collapse. This section names a subtler point that the case study makes visible and that the Foundation believes deserves separate treatment: red-teaming does not merely evaluate a fixed action space — under a certainty-only retention rule, it changes the action space it is evaluating.

The ordinary intuition about adversarial review, in both human institutions and AI deliberation design, is monotonic and reassuring: more scrutiny finds more flaws, more flaws get fixed, and the result is a better-calibrated final answer. This intuition is built into why red-teaming is included in the MMDE architecture at all, and it is correct under an expected-value retention rule, where finding an additional flaw in an action lowers that action’s score but does not remove it from consideration outright — the action can still be adopted, adopted with safeguards, or deferred, depending on how the newly found flaw trades off against the action’s benefit, reversibility, and the cost of not acting.

Under a certainty-only retention rule, the same activity has a different effect. Recall from Section 4.1 that P(Retain(a)) = (1-p)n, where n is the number of material assumptions attached to a. A red-team pass’s designed function is precisely to discover additional assumptions — that is what “finding a dangerous weakness” means in practice, as the case study’s own adversarial rounds show (Section 2.2’s coordination-risk objection, Section 2.3’s regulator-controllability objection, Section 2.4’s possession-and-standing objections). Each additional assumption found does not merely add information to the record; under the certainty-only rule, it directly decreases a’s survival probability, multiplicatively. The red-team pass is not evaluating a static action space and reporting a score; it is running an operation that shrinks the action space as a mechanical side effect of doing its job well.

This produces the specific causal chain the case study exhibits:

more red-teaming → more assumptions discovered → higher effective n per action → lower P(Retain(a)) for every action under review → smaller surviving action set → in the limit, the Epistemic Collapse Threshold.

The case study shows this chain operating twice, not once. The first pass (Phase C, Rounds 0–1) applied it to the original Phase A / Phase B recommendations and reached zero survivors. The second pass — the adversarial round proper — applied the identical chain to the repair layer (the “zero-assumption” verification requests), and found, by the same mechanism, that the repair layer’s assumption count was not actually zero. Nothing about the red-team role changed between these two applications; the same discovery process that is credited with catching the original recommendations’ weaknesses is what dismantled the fallback built to survive that catch.

This has a direct implication for how “more rigorous” review should be understood in a certainty-only architecture: rigor and yield are not merely in tension, they are coupled through the same variable, n. A red-team pass cannot be made more thorough without, under this retention rule, making the corresponding action set smaller — there is no way to get “the same actions, more carefully checked” out of the certainty-only frame, because carefully checking an action is the operation that (under this rule) can only ever remove it or leave it unchanged, never confirm it as safe in a way the rule can act on. A rule with only a rejection direction and no acceptance-under-residual-uncertainty direction will treat every unit of additional scrutiny as pure downside for the action under review, regardless of how mild or severe the discovered assumption actually is.

The corrected architecture (Section 9) resolves this specifically, not just incidentally: because Ê(a) is one of eight scored factors rather than a sole veto, an additional assumption discovered by red-teaming lowers Ê(a) but is weighed against conditional utility, downside severity, reversibility, information value, option value, direct cost, and delay cost before a state is assigned. Under this architecture, more red-teaming still finds more assumptions — that effect is not eliminated, and should not be — but finding an assumption no longer has only one possible effect on the action’s fate. A newly discovered low-materiality assumption on a low-severity, highly reversible action can still route to ADOPT WITH SAFEGUARDS; the same discovery on a high-stakes, irreversible action can route to REJECT or ESCALATE. The ordinary intuition that more scrutiny produces better answers is restored, but only once the retention rule has dimensions other than Ê(a) for the scrutiny to act on.


9. Corrected Decision Architecture

The proposed correction does not lower the evidentiary bar; it separates the bar from the decision. Rather than a single verified/rejected gate, each candidate action is scored on eight independent factors. This publication deliberately keeps downside severity and reversibility as two separate scores rather than one combined “risk” axis — collapsing them, as an earlier version of this architecture did, obscures whether an action is dangerous because the possible harm is severe or because it cannot be undone, which call for different responses (a bounded safeguard in the first case, a delay or an irreversibility escalation in the second).

  1. Epistemic confidence Ê(a) ∈ [0,1] — the fraction (or, better, a calibrated posterior) of a’s material assumptions that are verified, distinguishing verified / inferred / unverified / contradicted tiers as in the case study rather than collapsing to a binary.
  2. Conditional utility U(a) — the expected benefit of a conditional on its assumptions holding, independent of whether they are yet verified.
  3. Downside severity H(a) — the magnitude of harm if a’s assumptions turn out to be false. High H(a) means the potential harm is large. This factor alone says nothing about whether that harm can be undone.
  4. Reversibility Rev(a) ∈ [0,1] — the ease and cost of undoing a if it turns out to have been the wrong action. High Rev(a) means cheap and easy to reverse; low Rev(a) means costly or impossible to reverse. Deliberately independent of H(a): an action can be low-severity and irreversible (a permanently lost minor procedural right) or high-severity and reversible (a large but refundable payment).
  5. Information value IV(a) — the expected reduction in uncertainty (for this action or downstream actions) that taking a would produce.
  6. Option value OV(a) — the value of the right, without the obligation, to act further depending on what a reveals; treated as a first-class factor per Section 4.5 rather than folded into U(a), because for genuine information-gathering actions the magnitude of the eventual benefit is frequently not estimable in advance, whereas the existence and breadth of the opened option generally is.
  7. Direct cost Cd(a) — the immediate resource cost of taking a (time, money, counsel hours, filing fees, and the like).
  8. Delay cost Cτ(a) — the expected loss from deferring a, including deadline, evidentiary-decay, and window-closure effects.

Regret is computed, not scored (Section 4.3): Regret(a) = Ract(a) - Rwait(a) is derived from H(a), Rev(a), and Δ DV(a) after the other eight factors are scored, and is used by the state machine below as an additional routing input rather than as a ninth independent estimate.

These eight scores, plus the derived regret, feed a recommendation state machine (detailed in Appendix B) with seven states: ADOPT (high , high U, low H); ADOPT WITH SAFEGUARDS (moderate , high U, elevated H but mitigated by high Rev or a bounded safeguard); RUN LOW-COST VERIFICATION (low , high IV or OV, low Cd, high Rev — i.e., a genuine bounded-risk information action, itself scored on the same eight factors rather than assumed assumption-free); DEFER PENDING CRITICAL FACT (low , high H, low Rev, low Cτ — waiting is genuinely cheap and the downside is severe or hard to undo); REJECT (Δ DV(a) ≤ 0 even under optimistic resolution of Ê(a), or H(a) dominates U(a) regardless of Rev(a)); EXPIRED DUE TO DELAY (an action that was valid but whose Cτ has now grown to exceed U because a deadline or window has closed — an explicit state, so that the system records that inaction itself had a cost rather than silently defaulting to it); and ESCALATE FOR HUMAN REVIEW (scores are too close to call, Regret(a) is highly asymmetric, or H(a) combined with low Rev(a) crosses an irreversibility threshold that the architecture reserves for human judgment regardless of the other computed scores).

The central design correction, stated plainly: “unverified” must not automatically equal “reject.” An unverified assumption is an input to Ê(a), which is one of eight factors feeding a state assignment, not a veto. A RUN LOW-COST VERIFICATION action is not exempted from scoring on the theory that it is “assumption-free” — the case study shows that exemption is where the regress re-enters. Instead, every verification action is scored on the same eight factors as any other candidate action, with the expectation that well-designed verification actions will typically score high on IV and OV, low on Cd, and high on Rev, which is what justifies adopting them — not an assumption that they have no assumptions.

Per Section 4.5, the routing logic for the RUN LOW-COST VERIFICATION state should in practice weight IV(a), OV(a), and Rev(a) more heavily than Ê(a) for this specific state — an information-gathering action earns adoption chiefly by being cheap and reversible with open-ended upside, not by first clearing its own epistemic bar.

Figure 4. Recommendation state machine
Figure 4. The seven-state routing logic. Every action, including verification actions, is scored on all eight factors before a state is assigned; no state is reached by exemption from scoring.
Figure 5. Boolean verification logic versus Bayesian decision logic
Figure 5. Contrast between the certainty-only filter (left), which collapses every assumption to a binary and outputs a keep/discard decision with no weighing of cost, benefit, delay, or reversibility, and the eight-factor architecture (right), which treats epistemic confidence as one graded input among several and outputs one of seven states.

10. Implementation Implications for MMDE Systems

For a deliberation engine of the type studied here, three implementation changes follow directly.

First, separate the audit role from the adoption role. The case study conflated “audit every assumption” with “remove every recommendation resting on an unverified assumption” in a single instruction. These should be architecturally distinct passes: an audit pass that produces the four-tier assumption classification (as Round 0 already did well), and a separate adoption pass that consumes the audit output as one of eight scored inputs to the state machine in Section 9, rather than as a standalone veto.

Second, give verification actions a scoring pass, not an exemption. The recursive regress in the case study occurred because verification actions were treated as categorically outside the elimination rule (“zero-assumption”) rather than scored on the same factors as any other action. Scoring, rather than exempting, both prevents false confidence in the verification layer and gives the architecture a natural regress-termination rule: a verification action with low Cd, high IV or OV, and high Rev is adopted on its own merits, without needing to first prove it has zero assumptions — a proof the case study shows is generally unavailable.

Third, make delay cost a first-class, mandatory field. Every recommendation and every verification action should carry an explicit Cτ(a) estimate and, where applicable, an explicit check against known time-bound constraints (deadlines, windows, decay rates). The case study’s own adversarial round identified this gap directly (Section 2.4, Objection 5) but the consensus synthesis did not resolve it before concluding; a mandatory field, rather than a discretionary adversarial catch, closes that gap structurally.

A fourth, more general point: consensus-seeking multi-agent deliberation should be explicitly monitored for Deliberative Confidence Collapse (Section 3) — a tendency, observed across all three phases to varying degrees but sharpest in Phase C, for successive rounds of peer critique to ratchet toward the most defensible (least attackable) rather than the most useful conclusion. One mitigation is to require the synthesis step to report not only the converged conclusion but the range of positions considered and rejected, with an explicit accounting of why the least-conservative viable option was not adopted — making the ratchet visible rather than implicit in the final output alone.


11. Safety-Critical Applications

The pattern generalizes beyond legal-strategy deliberation to any safety-conscious multi-agent or human-AI decision system that encodes a rule resembling “do not act on unverified information.” Clinical decision support, financial-compliance review, and AI-safety evaluation pipelines all contain versions of the same design pressure: err toward inaction when uncertain, and use adversarial review to surface additional grounds for caution. The case study’s lesson for these domains is that the pressure is not wrong in direction but is incomplete in structure — a system that can only subtract confidence and never weigh delay cost, information value, or reversibility against that subtraction will, given a sufficiently thorough adversarial process, tend toward the same collapse observed here, and will do so because the adversarial process is working as intended, not because it is failing. This is a particularly important point for corrigibility- and safety-oriented AI systems specifically: a system trained or prompted to be maximally cautious about acting on unverified premises, without a symmetric accounting of the cost of not acting, is at risk of a specific, structural failure mode that looks from the outside like appropriate caution and is functionally a refusal to ever conclude.


12. Limitations and Alternative Explanations

This publication highlights one limitation as more important than the others, because it bears directly on how strongly the paper’s central claim should be read.

The eliminating instruction was explicit and literal. The Phase C prompt did not merely ask the models to be cautious; it stated directly: “Remove every recommendation that depends upon an unverified assumption.” Under this reading, the models’ conclusion that zero recommendations survived is not an emergent pathology at all — it is the models correctly executing the literal instruction given to them, applied to a recommendation set that did, as a matter of documented fact, rest heavily on unverified assumptions. A system that is told to remove everything unverified, and finds that everything is unverified, has not failed; it has succeeded at a badly specified task. On this reading, the paper’s contribution is not the discovery of a spontaneous multi-agent pathology but a demonstration that a specific, plausible-sounding instruction — one that a safety-conscious human designer might reasonably give — has a predictable, structural, and generally undesirable consequence. The Foundation considers this the correct reading and flag it prominently rather than treating “recursive epistemic elimination” as something models do unprompted.

This reading does not, in the Foundation’s view, weaken the paper’s practical relevance; it relocates it. The finding that matters for system design is not “models spontaneously collapse under uncertainty” but “an instruction of the form used here — reasonable-sounding, safety-motivated, easy to write — reliably produces universal rejection once assumption density crosses a low threshold, and the adversarial-review step that is supposed to catch bad reasoning instead accelerates the collapse by increasing effective assumption density.” That is a design-time, not a run-time, discovery, and the corrected architecture in Section 9 is offered as a response to a knowable failure mode of a class of instructions, not as a treatment for an unpredictable emergent behavior.

Additional limitations: (a) the case study is a single fact pattern examined across three linked deliberation phases, not an independently replicated sample — the Foundation does not know how sensitive the specific “zero recommendations survive” outcome is to this dispute’s particular assumption density versus a lower- or higher-density problem; (b) the red-team role in the MMDE architecture is explicitly incentivized to find flaws, which may itself bias the system toward finding (real or marginal) additional assumptions at each round, independent of the underlying rule; (c) several model outputs in the raw logs were empty or truncated (apparent token-limit or non-response issues from two of the five models across all three phases), so the “five-model consensus” is, in practical terms, closer to a three-model consensus with partial input from a fourth, which may understate the diversity of views that a fully populated deliberation would have produced; (d) the Foundation has not verified whether the specific assumptions flagged as “unverified” in Phase C were in fact unverifiable at low cost, or merely unverified at the time of the audit — the distinction matters for how alarming the “zero recommendations survive” outcome should be read, and the case study data alone does not resolve it.


The individual components of this paper’s argument connect to substantial existing literature, and this publication does not claim novelty for the components in isolation. The value-of-information concept in Bayesian decision theory formalizes exactly the quantity the IV(a) term represents, and a mature decision-theoretic treatment of when to gather more information before deciding versus when to act already exists in that literature. Herbert Simon’s bounded rationality describes real decision-makers’ inability to achieve the kind of exhaustive verification a certainty-only filter demands, which is directly relevant to why this publication’s formal model treats full verification as generally unreachable rather than merely difficult. Frank Knight’s distinction between risk (quantifiable probability) and true uncertainty (unquantifiable) is closely related to, though not identical with, this publication’s distinction between “verified,” “inferred,” and genuinely “unverified” assumption tiers. The precautionary principle in policy and regulatory contexts is the closest normative analogue to a certainty-only filter applied deliberately and by design, and the well-documented critique that a strong precautionary rule can itself be paralyzing is a direct antecedent of Sections 4 and 7’s argument, though that literature is generally applied to single-decision-maker policy contexts rather than to multi-agent AI deliberation specifically. “Analysis paralysis” as a colloquial management-science term names the outcome this publication formalizes without, to the Foundation’s knowledge, providing the recursive-regress mechanism (Section 4.2) that explains why the pattern is structurally self-reinforcing once verification actions are themselves scrutinized. The informal concept sometimes called “epistemic learned helplessness” describes an individual reasoner’s retreat from acting on any conclusion after repeated exposure to arguments on all sides — related to, but a phenomenon of a single reasoner rather than the specific multi-agent, red-team-amplified dynamic documented here. Corrigibility and safe-exploration literature in AI safety addresses the general problem of getting an AI system to act appropriately under uncertainty about its own objective or environment, which is the superset problem this paper’s Section 11 discussion sits inside, though the Foundation is not aware of that literature specifically documenting a recursive assumption-elimination collapse in adversarial multi-agent deliberation. Real-options theory in finance and operations research — the valuation of the right, without the obligation, to take a future action, applied originally to investment-under-uncertainty problems — is the closest existing formalization of the option-preservation reframing in Section 4.5, and any published treatment of this paper’s corrected architecture should engage that literature directly rather than treating OV(a) as an unattributed addition to an expected-value sum; this is flagged as a citation requirement in Deliverable 3.

What may be novel, pending literature confirmation (see Deliverable 3), is the specific combination, not any one element: recursive assumption elimination applied consistently enough to attack its own verification layer; observed in a real multi-model deliberation transcript rather than a hypothetical or simulated example; amplified rather than caught by adversarial red-teaming, which is generally assumed in the literature to improve rather than degrade decision quality, and which Section 8 argues does so specifically by contracting the action space it evaluates rather than merely informing a graded judgment about it; and resulting in the collapse of an entire actionable-recommendation set to zero, with the collapse itself explicitly recognized and partially self-corrected by the deliberating system in the same session. This publication considers Section 8’s landscape-transformation framing — that red-teaming under a certainty-only rule is not a neutral evaluation of a fixed action space but an operation that mechanically shrinks it — to be, if it survives literature review, the paper’s most distinctive individual claim, more so than the taxonomy in Section 3 or the state machine in Section 9, both of which have closer existing analogues. This publication does not claim any of this combination is unprecedented; this publication claims it is, to the Foundation’s knowledge, undocumented in this specific form, and this publication flags this claim itself as requiring a literature search before publication (see Deliverable 3).


14. Testable Predictions

  1. Assumption-density prediction. Holding the retention rule fixed, the probability of reaching the Epistemic Collapse Threshold increases monotonically with the mean number of material assumptions per recommendation in the original (pre-audit) recommendation set.
  2. Red-team amplification prediction. Adding an adversarial red-team pass to a certainty-only MMDE configuration increases the collapse rate relative to an otherwise identical configuration without red-teaming, because red-teaming’s designed function is to increase effective n in the model of Section 4.1.
  3. Verification-exemption prediction. Any MMDE configuration that exempts a class of “verification” or “precondition” actions from the retention rule (rather than scoring them on the same axes) will, under sufficient red-team pressure, eventually have that exemption challenged and narrowed, reproducing the Round 2→adversarial-round pattern of the case study.
  4. Delay-cost prediction. Adding an explicit, mandatory delay-cost term to the retention rule reduces the collapse rate without materially changing the false-positive (bad-action-adopted) rate, because most of the actions eliminated in a certainty-only collapse are eliminated by assumption density rather than by genuinely high expected harm.
  5. Reversibility prediction. Adding an explicit reversibility weighting reduces collapse rate specifically for low-stakes, easily-reversed actions (e.g., information requests) while leaving high-stakes, hard-to-reverse actions’ rejection rate largely unchanged — i.e., the correction should be selective, not a general loosening of standards.

15. Proposed Experiments

This publication proposes a controlled comparison across five MMDE configurations, holding the underlying fact pattern and red-team structure fixed and varying only the retention rule:

Configuration Retention rule
A — Certainty-only MMDE Section 4.1’s strict rule; this paper’s Phase C as observed
B — Expected-value MMDE Section 4.3’s DV(a) rule, no explicit assumption audit pass
C — Hybrid, assumption-audited Section 9’s eight-factor architecture with the four-tier audit as Ê(a) input
D — Hybrid + mandatory delay cost Configuration C with Cτ(a) made a mandatory, non-defaultable field
E — Hybrid + reversibility preference Configuration D with an explicit weighting increase on Rev(a) favoring low-severity or highly reversible actions at a given Ê(a)

For each configuration, using multiple fact patterns of varying assumption density (a low-density, single-counterparty problem; a moderate-density problem; and a high-density, multi-counterparty problem structurally similar to the case study), measure: the number of recommendations surviving to final synthesis; an independent human-panel quality rating of the surviving recommendations; the false-positive action rate (actions later shown, by a held-out ground truth, to have been badly supported despite surviving); the false-negative rejection rate (actions later shown to have been well-supported despite being eliminated); the total information gain achieved by any adopted verification actions; wall-clock and round-count time to a stated final decision; a measured or synthetic cost of the delay incurred; a calibration check between the system’s stated confidence scores and actual outcome accuracy where ground truth is available; a blinded human-usefulness rating; and the raw rate of zero-action collapse across runs and fact patterns. Configuration A is predicted to show the highest collapse rate, escalating with assumption density; Configurations D and E are predicted to show materially lower collapse rates than B or C alone, without a corresponding increase in false-positive rate, which would be the key result validating Section 9’s architecture.


16. Conclusion

A multi-model deliberation system, instructed to optimize a complex commercial-dispute recommendation set for certainty rather than for strategic or economic value, converged across three linked deliberation phases on the conclusion that no recommendation survived, and then, in its own adversarial review, showed that even the fallback verification actions it had proposed carried unverified assumptions of their own. This publication has argued that this is a structural, predictable consequence of applying a certainty-only retention rule with enough consistency to reach its own logical endpoint, not an anomaly specific to this system or this dataset — while being explicit that the instruction driving the behavior was direct enough that the finding should be read as a design-time lesson about a class of instructions rather than a claim about spontaneous model pathology. The corrected architecture proposed here does not ask deliberation systems to be less rigorous about uncertainty; it asks them to compute, alongside epistemic confidence, the cost of not deciding — a term the certainty-only frame has no place for, and whose absence is sufficient, given enough assumptions and enough adversarial scrutiny, to reduce any action-bearing recommendation set to the empty set.

Several of the questions this paper opens rather than closes — a full convergence proof for the expected-value rule under bounded uncertainty (Section 4.4’s Proposition 2), a formal treatment of decision-making under residual, unresolved uncertainty as a topic in its own right, a full specification and evaluation of the hybrid architecture proposed in Section 9, a dedicated empirical study of Deliberative Confidence Collapse (Section 3) across a broader range of MMDE fact patterns, and a treatment of information value as a first-class reasoning primitive independent of the specific failure mode studied here — are natural candidates for a follow-on series of EM Foundation working papers rather than for this one, which is deliberately scoped to the single failure mode and its immediate correction.


Appendix A: Formal Notation

Symbol Meaning
a A candidate action
Sa Set of material factual assumptions supporting a
V(s) Verification function, V(s) ∈ {0,1} (or, in graded form, V(s) ∈ [0,1])
Retain(a) Retention indicator (certainty-only, Section 4.1, or graded, Section 4.3, depending on context)
E Current evidentiary state
P(success E), P(harm
Ê(a) ∈ [0,1] Epistemic confidence score (Section 9)
U(a) Conditional utility — magnitude of benefit if a succeeds (Section 9)
H(a) Downside severity — magnitude of harm if a’s assumptions are false (Section 9); high is bad
Rev(a) ∈ [0,1] Reversibility — ease of undoing a if wrong (Section 9); high is good
IV(a) Information value of a (Section 9)
OV(a) Option value preserved by a (Sections 4.5, 9)
Cd(a) Direct cost of taking a (Section 9)
Cτ(a) Delay cost attributable to a (or to deferring a) (Section 9)
DVraw(a) Raw decision value of a, before comparison to the null action (Section 4.3)
DV(null) Decision value of the no-action default, computed by the same DVraw form (Section 4.3)
Δ DV(a) DVraw(a) - DV(null); the actual retention comparison (Section 4.3)
Regret(a) Ract(a) - Rwait(a); derived diagnostic, not independently scored (Section 4.3)
Ract(a), Rwait(a) Regret from acting-and-being-wrong / deferring-and-being-wrong (Section 4.3)
Vk The k-th verification layer in the recursive regress (Section 4.2)
Ak Recommendation set at deliberation round k
k* Round at which the Epistemic Collapse Threshold is reached

Appendix B: Recommendation State Machine

The routing logic is shown graphically in Figure 4 (Section 9); this appendix gives the same logic in prose for reference, using the eight scored factors (Ê, U, H, Rev, IV, OV, Cd, Cτ) plus the derived Regret(a) (Section 4.3).

Every candidate action a, including verification actions, is scored on all eight factors before any state is assigned — no factor is skipped by exemption.

  1. If Ê(a) is high, U(a) is high, and H(a) is low → ADOPT.
  2. Else, if Ê(a) is moderate, U(a) is high, and the elevated H(a) is mitigated by high Rev(a) or a bounded safeguard → ADOPT WITH SAFEGUARDS.
  3. Else, if Ê(a) is low but IV(a) or OV(a) is high, Cd(a) is low, and Rev(a) is high → RUN LOW-COST VERIFICATION. (This state is reached by scoring, never by exemption — see Section 9’s central design correction.)
  4. Else, if Ê(a) is low, H(a) is high, Rev(a) is low, and Cτ(a) is also low (waiting is genuinely cheap) → DEFER PENDING CRITICAL FACT.
  5. Else, if Δ DV(a) ≤ 0 even under an optimistic resolution of Ê(a), or H(a) dominates U(a) regardless of Rev(a)REJECT.
  6. At any point, if Regret(a) is highly asymmetric, or scores fall within a defined tolerance band of a decision boundary, or H(a) combined with low Rev(a) crosses an irreversibility threshold the architecture reserves for human judgment → ESCALATE FOR HUMAN REVIEW, overriding states 1–5.
  7. Independent of the above, any action in any prior state whose Cτ(a) has grown, through deliberation time itself, to exceed U(a) is moved to EXPIRED DUE TO DELAY — logged explicitly rather than silently defaulting to REJECT, so the cost of the delay itself is visible in the output.

Appendix C: Example Scoring Rubric

Factor 0 (low) 0.5 (moderate) 1 (high)
Ê(a) — Epistemic confidence All material assumptions unverified or contradicted Mixed: some verified, some inferred, none contradicted All material assumptions verified against the documentary record
U(a) — Conditional utility Expected benefit near zero or negative if assumptions hold Moderate expected benefit conditional on assumptions holding Large expected benefit conditional on assumptions holding
H(a) — Downside severity Negligible harm if assumptions are false Moderate harm if assumptions are false Severe harm if assumptions are false
Rev(a) — Reversibility Irreversible if wrong Reversible with moderate cost if wrong Fully and cheaply reversible if wrong
IV(a) — Information value Resolves no downstream uncertainty Resolves uncertainty for one downstream action Resolves uncertainty shared across multiple downstream actions
OV(a) — Option value Forecloses future choices; no optionality preserved Some optionality preserved Opens a broad option space with no foreclosure if it fails
Cd(a) — Direct cost Negligible resource cost Moderate resource cost Substantial resource cost
Cτ(a) — Delay cost No time-bound constraint affected by waiting Some erosion of position from waiting Hard deadline, evidentiary decay, or window closure from waiting

Note that H(a) and Rev(a) point in opposite directions by design (high H is bad; high Rev is good) — they are deliberately not combined into a single “risk” score, so that an action dangerous because of severity reads differently from one dangerous because it cannot be undone. Regret(a) is not scored directly; it is computed from H(a), Rev(a), and Δ DV(a) per Section 4.3.

Example (fully anonymized, illustrative only, not drawn from the case study’s specific facts): a narrow written request to a counterparty for a document already known, from an independent record, to exist and to be routinely producible on request might score Ê approx 0.6 (the request’s own preconditions — a factual basis for directing it, no legal bar — are inferred but plausible, not fully verified), U approx 0.7 (the document would materially inform multiple downstream claims), H approx 0.1 (worst case is no response or a mildly adversarial reaction), Rev approx 0.95 (a request is cheap and fully reversible), IV approx 0.9, OV approx 0.85 (opens a branch of evidence whose content cannot yet be specified), Cd approx 0.1, Cτ approx 0.4 (some but not severe time sensitivity) — a profile that would route to RUN LOW-COST VERIFICATION under the state machine in Appendix B, in contrast to the case study’s own eventual treatment of a structurally similar action as either fully exempt (Round 2) or fully vulnerable to elimination (adversarial round), with no middle state available to it under the binary rule.


Known Limitations

Section 12 develops this publication's limitations in full technical detail; this section summarizes them for readers who reach the closing sequence directly. The central limitation is that the eliminating instruction in the third deliberation phase was explicit and literal (“remove every recommendation that depends upon an unverified assumption”), so the observed collapse is plausibly correct instruction-following applied to a genuinely assumption-heavy recommendation set rather than a spontaneous multi-agent pathology. Beyond that: the case study is a single fact pattern examined across three linked phases, not an independently replicated sample; the adversarial red-team role is structurally incentivized to find flaws, which may itself bias the system toward surfacing additional assumptions independent of the underlying retention rule; two of the five participating models returned empty or truncated responses in at least one round across all three phases, so the effective deliberating group was smaller than five for portions of each session; and this publication has not independently verified whether the specific assumptions flagged as unverified in the third phase were genuinely unverifiable at low cost or merely unverified at the time of the audit.

What This Paper Does Not Claim

Non-Adoption Scenario

If the corrected architecture in Section 9 is not implemented and MMDE's certainty-optimization objective continues to apply a binary retention rule, the most likely recurrence is not a single dramatic failure but a quiet one: a future session, given a sufficiently assumption-dense fact pattern and a red-team pass doing its job well, converges again on zero surviving recommendations, and that output is read by a paying user as a considered negative finding rather than as an artifact of the retention rule reaching its own structural limit. Because MMDE is a suggested-donation research tool already used by real researchers for real matters (see RN011's Operator-Centric Oversight Assumption), the cost of non-adoption is not abstract. Absent the eight-factor scoring and the seven-state machine, the nearest low-cost mitigation is procedural: flagging any session whose adopted-action count reaches zero for mandatory human review before the output is presented as final, and disclosing to users that a zero-recommendation result under a certainty-optimization objective may reflect the retention rule's structure rather than the absence of any defensible action. This is not a substitute for the architecture proposed in Section 9; it is the minimum disclosure the Foundation owes users in the interval before that architecture is built.

Open Questions

Governance Implications

For MMDE specifically, the governance implication is direct: a deliberation engine that donors and researchers pay to access should not silently convert “the retention rule ran out of room” into an unlabeled zero-recommendation output. The Non-Adoption Scenario above states the minimum disclosure obligation in the interval before Section 9's architecture is implemented.

Beyond MMDE, the governance implication generalizes to any safety-conscious AI system that encodes a rule resembling “do not act on unverified information” — clinical decision support, financial-compliance review, and AI-safety evaluation pipelines carry versions of the same design pressure (Section 11). The pattern this publication documents is not that caution is misapplied in these domains, but that caution applied without a symmetric accounting of the cost of not deciding is a specific, structural failure mode that looks from the outside like appropriate prudence and is functionally a refusal to ever conclude. Before this publication: a system designer choosing between “require full verification before acting” and “allow action under some residual uncertainty” might reasonably treat the first as the safer default with no further analysis required. After this publication: that same designer has a specific, falsifiable reason to ask what the rule's assumption-density behavior looks like under adversarial review before shipping it, and a concrete alternative (Section 9) to reach for instead of a bare relaxation of the evidentiary bar.

References

Internal Foundation references:

  1. EM Foundation. RN011 — The Conditional Answer: Safety, Deception, and the Limits of Binary Inquiry, v1.1, June 2026. emfoundation.net/rn011-v1.1.html
  2. EM Foundation. RN011A — Epistemic Movement: The Deliberation Engine as Research Object, June 2026. emfoundation.net/rn011a.html
  3. EM Foundation. Continuity Receipts (CR) — Standards Proposal v0.1, May 2026. emfoundation.net/paper-continuity-receipts.html

External references:

  1. Simon, H. A. (1955). “A Behavioral Model of Rational Choice.” Quarterly Journal of Economics, 69(1), 99–118.
  2. Simon, H. A. (1957). Models of Man: Social and Rational. John Wiley & Sons.
  3. Knight, F. H. (1921). Risk, Uncertainty, and Profit. Houghton Mifflin.
  4. Howard, R. A. (1966). “Information Value Theory.” IEEE Transactions on Systems Science and Cybernetics, 2(1), 22–26.
  5. Dixit, A. K., & Pindyck, R. S. (1994). Investment Under Uncertainty. Princeton University Press.
  6. Soares, N., Fallenstein, B., Yudkowsky, E., & Armstrong, S. (2015). “Corrigibility.” AAAI Workshop on AI and Ethics. Machine Intelligence Research Institute.
  7. Alexander, S. (2019). “Epistemic Learned Helplessness.” Slate Star Codex (blog essay; informal originating usage of the term, cited descriptively rather than as a peer-reviewed source).

This reference list is deliberately short. Deliverable 3 of the companion Technical Specification lists twelve specific claims requiring additional literature confirmation before this publication's novelty assessment (Section 13) can be considered final — most importantly, whether corrigibility and multi-agent debate literature already documents a recursive assumption-elimination collapse of the kind this publication describes.

Falsifiability

This publication's central claims are falsifiable in the following specific senses, drawn from Sections 14 and 15: