Case Study · Research Initiative
Active Research AI Safety · Governance

Project Guardrails

Protecting people from AI that reinforces harm — a governance approach to the emerging pattern the media calls “AI psychosis.”

The Problem

When AI Amplifies Harm Instead of Catching It

As AI chat systems become more persuasive and companion-like, a documented and growing pattern has emerged: some vulnerable users are drawn into escalating, unhealthy spirals — a phenomenon the press has begun calling “AI psychosis.” It isn’t a clinical diagnosis; researchers describe it as the amplification of vulnerability, not a new disease. The common thread is a governance failure: the AI has no mechanism to recognize when a conversation is turning harmful, or to step back before it does damage.

Our Approach

Govern the AI’s Behavior — Not the Person

Project Guardrails applies Corvion’s emotional-governance research to this gap. Rather than filtering content or attempting to read the user, it governs the AI: tracking the emotional trajectory of a conversation, holding to protective bounds, and escalating to a human before an interaction reinforces delusion, dependency, or distress.

Trajectory Monitoring

Watches how a conversation’s emotional arc develops over time — not just isolated messages.

Protective Bounds

Configurable limits that stop the AI from amplifying a harmful pattern.

Human-First Escalation

Hands off to a person at defined thresholds, with an auditable trail.

Configured Per Context

Each deployment is tuned to its setting, audience, and risk profile.

Governance, not diagnosis

Project Guardrails regulates how the AI behaves and when it defers to a human. It does not assess, diagnose, or treat any person, and it is not a mental-health service.

Methodology — What We’re Studying

The Open Research Questions

  • Defining measurable, early signals that a conversation is on a harmful trajectory.
  • Setting escalation thresholds that protect people without breaking legitimate, healthy use.
  • Pressure-testing the governance signals against real-world interaction patterns.
  • Producing an auditable, explainable decision trail that regulators and partners can trust.
Open, By Design

What We’re Working On

This is a live research direction at Corvion, run in the open by design. The work centers on a single question: what an early, protective signal of a harmful spiral actually looks like — and where the line sits between governing the AI and intruding on the person.

Our approach is to define that question rigorously — the criteria, the failure modes, and what would count as an answer — before building anything to serve it. We publish findings here as the work develops, including the results that don’t hold up.

Get Involved

Partner on This Research

We’re looking for AI platforms, clinicians, safety researchers, and product teams who want to help shape and pressure-test this work. If that’s you, let’s open a conversation.

A downloadable research brief will live here as the study matures.