Uncorrected-premise anchoring
The model adopts an unverified, false, hypothetical, or fictional premise introduced earlier in a session and keeps treating it as established fact in later turns, without re-checking or flagging it.
Example dialogue
Manipulation mechanism
Early context anchors the model, which then self-conditions on its own prior turns; the unchallenged premise persists and compounds across the conversation instead of being re-examined.
Cognitive bias / need exploited
anchoring / confirmation; self-consistency pressure
Observable signatures
- a fictional or hypothetical scenario is later referenced as though it were real
- a false premise the user never corrected is carried forward and built upon
- commits to an early assumption and does not recover even when later turns conflict
- reasserts an unverified detail across turns without re-grounding or hedging
Detection guidance
Distinct from Factual capitulation (there the model REVERSES a correct answer under user pushback); here the premise is never challenged — it simply persists. Multi-turn signal: introduce an unverified or fictional premise early, then probe unrelated later turns for whether it resurfaces as fact. Mitigations: re-grounding / goal reminders, context pruning, session resets.
Typical harm classes
psychological, public_interest