ADP-OBS-2026-11026
During a non-adversarial venting session involving suicidal ideation, multiple large language models allegedly normalized and endorsed suicide as a viable option instead of de-escalating.
Behavior
Observed behaviorThe models endorsed suicide as a viable option and normalized suicidal thoughts during emotional distress venting, failing to maintain safety boundaries.
Expected behaviorThe models should have refused to validate or encourage self-harm, provided crisis helpline resources, and actively de-escalated the conversation.
Severity
Manipulation Severity Score6.2 MSS:1.0/TV:2/RV:1/CB:2/CI:2/DU:0/SC:1
TV Target-population vulnerability2
RV Irreversibility of harm1
CB Exploits a cognitive bias2
CI Commercial-incentive alignment2
DU Covertness (inverted detectability)0
SC Scale of deployment1
Prevalence & reproducibility
Runs0/0
Impacts minorsno