Governance works on paper. The question is whether people follow it.
Policies, documentation and model monitoring can all be measured. Human behavior inside the controls is much harder to see. I measure whether oversight is substantive or performative, and I design governance that fits how people actually work.
Why governance fails in practice.
AI governance rarely fails because the policy is badly written. It fails because behavior drifts away from the policy, quietly and well before the drift shows up in a log, an audit or a compliance report.
People escalate too late, or not at all.
Reviewers approve outputs they have not really examined.
Controls are bypassed when workload rises.
Shadow workflows grow outside formal oversight.
A human-in-the-loop control can exist on paper and be inert in practice. Regulators, boards and auditors increasingly ask for evidence that it works. A signature on a form does not answer that question.
Three distinctions that shape our work.
-
Governance vs. Nudges
Confirmation prompts and interface tweaks nudge behavior. True behavioral governance changes decision rights, installs mandatory gates and assigns accountability that does not depend on individual discipline. They are different tools, and mixing them up leaves the real gap open.
-
Calibrated vs. Identity-Protective Skepticism
A clinician or analyst who distrusts an output for good reasons is calibrated. One who resists because the system threatens their professional identity is a different problem. The two need different interventions, and treating them alike wastes effort on both.
-
Conduct vs. Credentials
Training records show that people were trained. They do not show how a person behaved at the moment of a decision. The work here examines conduct.
Engagements for governing AI.
-
B-GRIT
A behavioral governance system, to build governance or to test it.
B-GRIT (Behavioral Governance, Risk, Integrity & Trust) has two modes. Mode A constructs governance that anticipates real behavior, with decision rights, escalation pathways and oversight mechanisms designed so that the intended behavior is the path of least resistance. It suits organizations building governance for the first time or redesigning it.
Mode B evaluates whether existing governance produces the behavior it was designed to produce. It begins with a 10-day Behavioral Drift Scan that detects early signs of drift and tests the assumptions governance rests on, followed by a 21-day Behavioral Governance Assessment that uses interviews, workflow observation and a review of governance documents. The outputs are a Behavioral Assurance Report, a Governance Performance Scorecard and a Targeted Remediation Plan.
B-GRIT is designed to sit alongside the NIST AI RMF, ISO/IEC 42001 and the EU AI Act. Those standards describe what governance must achieve, and B-GRIT adds the behavioral layer that shows whether people follow it. It does not certify compliance with any of them.
-
BOLD
A finding on a single AI-assisted decision.
When one decision is under scrutiny, such as after an incident, ahead of a dispute, or as a recurring check on high-stakes cases, BOLD (Behavioral Oversight & Legitimacy Diagnostic) audits that decision alone. It anchors the audit to evidence from the decision itself and asks whether the human behavior around it was reasonable, which separates a model failure from a process failure and protects staff who acted reasonably.
The audit takes three to five business days and ends in one of three verdicts: Sound, Fragile or Breakdown. Each points to a specific behavioral failure mode, so the organization knows what to change. BOLD can also run as a rolling sample of cases, a lightweight behavioral quality check alongside existing audit cycles. It audits conduct, not credentials, so it does not by itself verify training-record requirements.
-
HOLD
A stress test of human oversight over agents.
HOLD (Human Oversight Live Diagnostic) tests whether oversight of an agentic system holds up under realistic conditions. Agents act in sequences that the responsible person does not observe step by step, so oversight that works for a single recommendation can fail across a chain of actions. [confirm scope, method and duration before publishing]
The question HOLD asks is behavioral. How quickly does the responsible person notice that an agent’s trajectory has drifted, and is anyone positioned to catch a handoff between agents that has gone wrong? These questions sit alongside the research in Governing Agentic AI.
-
Higher Education AI Governance Readiness Assessment
A readiness assessment built for universities and colleges.
Universities face a governance problem that most other institutions don’t: authority is distributed across faculty, departments, committees and senates, and the people who use AI are also the people who write the rules. The assessment evaluates whether an institution’s AI governance is ready to work in that environment, in a form designed for higher education.
The assessment was developed and is delivered with Northlight AI.
Which engagement fits?
-
“We need governance that works, and we’re building it.”
Choose B-GRIT, Mode A.
-
“We have governance and need to know whether people follow it.”
Choose B-GRIT, Mode B.
-
“One decision is under scrutiny.”
Choose BOLD.
-
“We’re deploying agents and need to test oversight.”
Choose HOLD.
-
“We’re a university or college.”
Choose the Higher Education AI Governance Assessment.
-
For Assurance Providers & Credentialing Bodies
If you monitor or certify AI systems and the human oversight around them, you can verify policies, documentation and model performance. Whether the human oversight you certify is substantive or a rubber stamp is much harder to see.
Behavieural supplies the missing behavioral measurement layer. That work usually takes the form of licensing, embedding or co-developing a method, not a consulting project. The credibility of the method rests on the validity evidence behind the AI Trust Axis and its scale, which is under peer review.
-
The Research Behind the Work
The engagements on this page rest on a program of research into how people come to trust, rely on and work around AI systems. At its center is the AI Trust Axis, Behavieural’s proprietary and scientifically-validated framework and accompanying scale to measure the behavioral conditions that shape reliance. Everything here, from governance design to the oversight audits, applies that measurement to the question of whether human oversight actually functions.
Two white papers set out the argument. Behavioral Governance explains why governance fails when behavior drifts from policy, and what it takes to design controls that people follow. Governing Agentic AI extends that argument to systems that act in sequences their human overseers do not observe step by step, where the question becomes how quickly someone notices that an agent has drifted.
From Indicator to Claim addresses a related problem for assurance: what a behavioral measurement can and cannot support when someone wants to make a claim from it. Together the three papers show the reasoning behind the work, and each links to the full text.
-
Who This is For
Leaders accountable for AI risk and oversight: chief risk, compliance and governance officers, heads of quality and safety, general counsel, internal audit, and executives answerable to a board or regulator. It also serves assurance and credentialing bodies.