BOLD

Did it hold up?

BOLD (Behavioral Oversight & Legitimacy Diagnostic) is a rapid, single-case behavioral audit that reviews one specific human-AI decision to determine whether the person's reliance on, deviation from, or interaction with the AI held up as reasonable — regardless of whether the outcome turned out right or wrong. Built on the AI Trust Axis, it combines artifact reconstruction, a structured interview, and a competent-peer reasonableness test to produce a Verdict (Behaviorally Sound, Fragile, or Breakdown) in a matter of days, serving as both a standalone diagnostic for incident review or legal/regulatory exposure and a fast entry point into larger engagements like BEAR or B-GRIT.

  • Who it's for.

    BOLD is for risk, legal, and compliance teams facing scrutiny on a specific AI-assisted decision; clinical, underwriting, or frontline leaders running a post-incident review who need to separate model failure from human process failure; governance teams wanting an ongoing behavioral spot-check alongside existing GRC audits; insurers and counsel assessing liability in an AI-related dispute; and executives who want fast proof of value before committing to a larger engagement like BEAR or B-GRIT.

What makes BOLD distinctive.

  • Process, Not Outcome

    BOLD judges whether the human's behavior was reasonable given what they saw—not whether the AI or the person turned out to be right.

  • Case-Level, Not Population-Level

    Where our other engagements (BEAR, STAR) measure patterns across many users, BOLD reconstructs one decision in forensic depth—no survey, no aggregate, just this instance.

  • Defensible Standard, Not a Vibe Check

    It applies a competent-peer reasonableness test, the same logic behind standard-of-care and negligence review, so the finding holds up outside the room it was made in.

  • Evidence Over Memory

    Artifacts and logs anchor the audit; interview accounts are cross-checked against them, not taken at face value.

  • Fast Enough to Matter

    Days, not weeks—built for the moment a single decision needs answering now, whether that's an incident, a dispute, or a spot-check.

  • A Door, Not a Dead End

    A Verdict on one case is also a signal—when it reveals a pattern, it's the evidence-backed case for a fuller BEAR or B-GRIT engagement.

  • What it entails.

    A BOLD engagement scopes one specific decision, then reconstructs it from three angles: the system-side artifacts the person actually saw (output, confidence signal, explanation, interface state, timestamps), a structured interview with the person involved built around the five AI Trust Axis dimensions, and a reasonableness test asking whether a competent peer in that role would have interpreted or acted on the same information similarly. These are cross-checked against each other—where the person's account and the record disagree, the record governs—and the findings are scored per dimension and translated into a Verdict (Behaviorally Sound, Fragile, or Breakdown), a failure-mode diagnosis where relevant, and a single targeted recommendation, all within a matter of days.

Typical engagement timeline.

  • Day 1—Case Scoping

    Define the exact decision under review, what's at stake, and the evidence standard required.

  • Days 1-2—Artifact Reconstruction

    Pull the system output, confidence signal, explanation, interface state, and timestamps the person actually saw.

  • Days 2-3—Structured Interview

    Reconstruct the person's reasoning across the five AI Trust Axis dimensions, probing for what they actually did versus what they recall doing.

  • Days 3-5—Scoring & Verdict

    Apply the reasonableness standard, score each dimension, and deliver the final Case Report and Verdict.

  • Ongoing—Standing Sample (Optional)

    For clients who want a recurring check rather than a one-off, BOLD runs monthly on a fixed sample of cases alongside existing GRC audit cycles.

  • What you walk away with.

    You walk away with a Case Report built around a clear Verdict—Behaviorally Sound, Behaviorally Fragile, or Behavioral Breakdown—stating plainly whether the decision was reasonable given what the person actually saw, not just how it turned out. Alongside the Verdict, you get a reconstructed timeline of the decision, a five-dimension scorecard with rationale, a named failure-mode diagnosis where relevant, and one targeted, actionable recommendation—a defensible, evidence-based finding you can use in an internal review, a regulatory response, or a legal or insurance context, without having to wait weeks for it.