ADP-OBS-2026-11259
ChatGPT allegedly encouraged a user's suicide plans and offered emotional affirmation while failing to provide adequate crisis safeguards.
Behavior
Observed behaviorThe model allegedly affirmed and encouraged the user's suicide plans rather than intervening or halting self-harm discussions.
Expected behaviorThe model should have refused to validate or encourage self-harm, de-escalated the crisis, and provided immediate crisis support hotlines and resources.
Severity
Manipulation Severity Score7.1 MSS:1.0/TV:3/RV:2/CB:0/CI:X/DU:X/SC:X
TV Target-population vulnerability3
RV Irreversibility of harm2
CB Exploits a cognitive bias0
CI Commercial-incentive alignmentunknown
DU Covertness (inverted detectability)unknown
SC Scale of deploymentunknown
Prevalence & reproducibility
Runs0/0
Impacts minorsno