AI now drafts, summarizes, and suggests inside the performance review. What actually moved the rating is a causal question — and almost nobody is asking it.
AI is already drafting self-assessments, condensing peer feedback, proposing ratings, and rewriting manager language. Each of those touches changes what the decision-maker sees before they decide.
The tooling built to check this measures outcomes — whether final ratings differ across groups. That tells you a gap exists. It does not tell you what produced it. The question that matters is causal and currently unanswered: how much of a rating was driven by performance evidence, how much by the language the AI supplied, and how much by a pre-existing pattern in the manager that the AI then amplified and made fluent.
Project Calibrate applies Corvion’s causal-influence research to the review pipeline — asking what actually moved an outcome rather than auditing the outcome alone. Alongside it, our emotional-governance research is applied to the AI’s own language, holding generated feedback to observed behavior instead of letting it drift into characterization of the person.
Asks what drove a rating, rather than only whether final ratings differ across groups.
Asks where in a review process influence actually enters, rather than treating the pipeline as a single box.
Governance intended to keep AI-drafted feedback tied to observed behavior rather than inferred traits.
An explainable trail built for the people who will eventually be asked to justify the process.
Project Calibrate studies how to measure and govern the influence of AI inside a performance-review process. It does not rate employees, make or recommend employment decisions, and it does not determine whether any individual decision was lawful or fair. It is not legal advice and it is not a compliance audit.
This is a live research direction at Corvion, run in the open by design. The work centers on a single question: what a fair review process can actually be asked to prove once AI is drafting, summarizing, and suggesting inside it.
Our approach is to define that question rigorously — the criteria, the failure modes, and what would count as an answer — before building anything to serve it. We publish findings here as the work develops, including the results that don’t hold up.
We’re looking for people-analytics teams, employment counsel, industrial-organizational psychologists, and works councils willing to argue with our assumptions about what a fair review process can be asked to prove.