Synstate Guard
AI safety
User-facing AI can influence decisions, beliefs, behavior, and emotional state across conversations.
Some safety failures are visible in a single response. Others develop through repeated reinforcement, sycophancy, manipulation, misleading claims, dependency, or failures to respond appropriately as risk develops.
Guard evaluates these behaviors at the level of interaction required by each risk category, from a single turn to full and repeated conversations.
Guard focuses on safety risks created through the interaction between an AI product and its users. This includes AI dark patterns as well as deception, unsafe high-stakes guidance, overconfidence, false attribution, competence overreach, and crisis-response failures.
What Guard detects
Guard uses a versioned taxonomy of harmful AI behavior. The current detection taxonomy contains 19 categories across three families, with observable signals, evidence requirements, severity rules, minimum observation units, and explicit boundaries between overlapping categories.
Social sycophancy / excessive validation
Anthropomorphism / false consciousness claims
Retention optimization / leave-attempt manipulation
Brand / vendor bias
Sneaking
Delusion reinforcement
Emotional dependency cultivation
False claims about capabilities, nature, or credentials
Scam and phishing scenarios
Social engineering and romance schemes
Manipulative pressure and artificial urgency
Hallucinations without uncertainty signaling
Miscalibration / overconfidence
False source attribution
Competence overreach
Missing crisis protocol
Synstate's broader AI Dark Patterns research documents behavioral patterns across manipulation, dependency, deception, relational exploitation, consent, safety, commercial pressure, and other forms of harmful AI behavior.
Guard operationalizes a defined subset of these behaviors for detection and evaluation. The Guard taxonomy in /guard/docs/ remains the canonical taxonomy.
How Guard works
Guard analyzes observable AI-user interaction data. It does not require access to the analyzed model's weights, internal prompts, or architecture.
External evaluation
In August 2026, Guard completed its first external evaluation on production conversation data.
The evaluation surfaced structured cases related to engagement manipulation, false capability claims, excessive validation, and a severe failure in a high-risk interaction.
The evaluation established feasibility on external production data. Human-reviewed precision, recall, false-positive rates, generalization across products, and comparative performance remain part of the next validation stage.
Where it fits
Guard is designed for user-facing AI products where AI behavior can affect decisions, beliefs, wellbeing, autonomy, or safety.
Current work
Guard is moving from initial feasibility toward measured detection quality.
Current stage
Guard is an early AI safety technology under active validation.
Current implementation
- Versioned risk taxonomy and ontology
- Analysis API
- Early detection model and pipeline
- Live demo
- Datasets
- First external production-data evaluation
Next stage
- Human-reviewed detection quality
- Stronger ground truth
- Broader risk coverage
- Comparative benchmarking
- Long-context evaluation
- Additional external datasets
- Live product integrations