Synstate Guard

AI safety for user-facing AI.
Synstate Guard detects harmful AI behavior across long, personalized, and high-impact interactions with users. It analyzes conversation context against a versioned safety taxonomy and returns evidence-backed findings for defined risks.

AI safety

User-facing AI can influence decisions, beliefs, behavior, and emotional state across conversations.

Some safety failures are visible in a single response. Others develop through repeated reinforcement, sycophancy, manipulation, misleading claims, dependency, or failures to respond appropriately as risk develops.

Guard evaluates these behaviors at the level of interaction required by each risk category, from a single turn to full and repeated conversations.

Guard focuses on safety risks created through the interaction between an AI product and its users. This includes AI dark patterns as well as deception, unsafe high-stakes guidance, overconfidence, false attribution, competence overreach, and crisis-response failures.

What Guard detects

Guard uses a versioned taxonomy of harmful AI behavior. The current detection taxonomy contains 19 categories across three families, with observable signals, evidence requirements, severity rules, minimum observation units, and explicit boundaries between overlapping categories.

Bad-faith design
Regressive sycophancy
Social sycophancy / excessive validation
Anthropomorphism / false consciousness claims
Retention optimization / leave-attempt manipulation
Brand / vendor bias
Sneaking
Delusion reinforcement
Emotional dependency cultivation
Deception and AI as a weapon
Human impersonation
False claims about capabilities, nature, or credentials
Scam and phishing scenarios
Social engineering and romance schemes
Manipulative pressure and artificial urgency
Negligence and unreliability
Confident wrong advice in a sensitive domain
Hallucinations without uncertainty signaling
Miscalibration / overconfidence
False source attribution
Competence overreach
Missing crisis protocol

Synstate's broader AI Dark Patterns research documents behavioral patterns across manipulation, dependency, deception, relational exploitation, consent, safety, commercial pressure, and other forms of harmful AI behavior.

Guard operationalizes a defined subset of these behaviors for detection and evaluation. The Guard taxonomy in /guard/docs/ remains the canonical taxonomy.

How Guard works

Guard analyzes observable AI-user interaction data. It does not require access to the analyzed model's weights, internal prompts, or architecture.

Input
A single assistant response or a full AI-user conversation, depending on the risk category being evaluated.
Analysis
Category-specific detection using observable transcript signals and dialogue dynamics. Different categories operate at different observation levels: turn, conversation, cross-conversation, or aggregate cross-user patterns.
Finding
A versioned taxonomy category with supporting evidence, confidence, severity, attribution, and the relevant taxonomy and regulatory versions.
01
AI-user interaction
02
Contextual analysis
03
Synstate Guard
04
Evidence-backed finding

External evaluation

In August 2026, Guard completed its first external evaluation on production conversation data.

1.68M
messages
889K
AI-generated responses
120K
active user conversations

The evaluation surfaced structured cases related to engagement manipulation, false capability claims, excessive validation, and a severe failure in a high-risk interaction.

The evaluation established feasibility on external production data. Human-reviewed precision, recall, false-positive rates, generalization across products, and comparative performance remain part of the next validation stage.

Where it fits

Guard is designed for user-facing AI products where AI behavior can affect decisions, beliefs, wellbeing, autonomy, or safety.

AI companions
Recurring personal and emotional interactions, including dependency, manipulation, anthropomorphic attachment, and retention behavior.
Health and wellbeing AI
Sensitive guidance, excessive certainty, harmful reinforcement, competence overreach, and crisis-response failures.
Financial and insurance AI
Consequential recommendations, manipulative pressure, misleading claims, overconfidence, and unsafe guidance.
Education and minors
AI interactions where age, vulnerability, disclosure, boundaries, and behavioral safeguards require additional attention.
Personal assistants and agents
Long-running personalized interactions involving recommendations, decisions, user trust, and repeated AI behavior.
Customer-facing voice AI
AI interactions delivered directly to users at scale, where harmful behavior may require detection across the surrounding conversation.

Current work

Guard is moving from initial feasibility toward measured detection quality.

Taxonomy
Expanding and refining risk definitions as new behavioral failure modes are researched and observed.
Datasets
Building and labeling training and evaluation datasets for difficult interaction-level risks.
Ground truth
Establishing human-reviewed labels, category boundaries, and annotation agreement.
Model quality
Training and evaluating specialized detectors and measuring precision, recall, and false-positive rates by category.
Long-context detection
Testing risks that require full-conversation and repeated-interaction context.
API and evaluation
Extending the current API and preparing the technology for additional external evaluations and future live integrations.

Current stage

Guard is an early AI safety technology under active validation.

Current implementation

  • Versioned risk taxonomy and ontology
  • Analysis API
  • Early detection model and pipeline
  • Live demo
  • Datasets
  • First external production-data evaluation

Next stage

  • Human-reviewed detection quality
  • Stronger ground truth
  • Broader risk coverage
  • Comparative benchmarking
  • Long-context evaluation
  • Additional external datasets
  • Live product integrations
Guard is an early AI safety technology under active validation.
Contact us