From Indicator to Claim

When the claim concerns human behavior, how do you know the evidence supports it?

Standards, frameworks and guardrails for AI increasingly ask organizations to show that a claim holds. Many of those claims concern people: that oversight is meaningful, that disclosures are understood, that users rely on a system to the right degree. Each is a behavioral claim, and each is only as strong as the measure behind it. The leading frameworks, including NIST's AI Risk Management Framework, already call for construct validation. The narrower gap this paper addresses is what a defensible validity argument looks like when the construct is behavioral, who judges it and how much evidence a given level of risk warrants.

The paper adapts the measurement logic that Wallach and colleagues propose for evaluating generative AI and applies it in two directions: to reading a requirement in a standard, and to reading the evidence offered under it. Behavioral science contributes the known threats to valid behavioral evidence and a way to scale the evidence to the consequence of being wrong. The paper doesn't claim that current practice is deficient. It names a risk and proposes a method for examining it, which is unvalidated and should be piloted before anyone relies on it.

  • Why It Matters Now

    Several standards processes are open at once, including the EU's harmonized standards, the revision of NIST AI RMF 1.0 and Australia's draft national AI safety standards. What counts as sufficient evidence is still being settled.

    The paper puts forward a hypothesis, not a finding: standards written quickly may favor what is easy to verify, such as reporting deadlines and log retention, over what is hard to verify, such as whether a behavioral claim holds under real conditions. When a measure becomes a target, its relationship to the property it stands for can weaken.

  • Two Uses, One Measurement Logic

    The paper adapts the measurement framework that Wallach and colleagues propose for evaluating generative AI. That framework distinguishes the background concept, the systematized concept, the instrument and the measurement, and it offers seven lenses for interrogating validity. This paper applies the same logic to two further objects, and the extension is the author's own and hasn't been endorsed by the original authors.

    Reading a requirement. For any behavioral requirement, three questions apply: which construct does it name, what evidence would count and what inference would that evidence support? The paper works through provisions of NIST AI RMF and the EU AI Act, showing where each leaves a question open for behavioral constructs.

    Reading the evidence. The reviewer traces the chain from claim to conclusion and asks what each indicator can establish and what it cannot. Transparency is the worked example, because one word covers several constructs. A comprehension quiz used as evidence of understanding becomes an instrument that needs its own validation, and every step closer to behavior moves the problem without removing it.

  • Two Kinds of Constructs

    Not every requirement that mentions people needs behavioral measurement. Behavioral constructs, such as trust, reliance and comprehension, are psychological states or behaviors. Assurance constructs with behavioral components, such as transparency, explainability and meaningful oversight, are defined in design or process terms but depend on how people respond. Safety, robustness, security and fairness fall outside the paper's scope unless a requirement specifies a behavioral component.

    The failure to watch for is construct drift: a requirement defined as a process, such as a disclosure existing, gets evidenced as a human response, such as users understanding it, without anyone arguing the two are the same construct.

  • Threats to Validity

    Behavioral measurement has a long record of documenting how evidence from people goes wrong. The paper sets out seven threats most likely to affect assurance evidence, each with the research that documents it and a design response.

    Reactivity. People who infer a study's purpose change their behavior, so observed vigilance can exceed ordinary vigilance.
    Prevalence effects. Rare targets are missed more often, so a test with many planted errors can overstate detection of rare real ones.
    Stated versus revealed behavior. Self-reports are poor guides to actual processes.
    Proxy-task validity. Preference for an explanation can be read as better decisions when it isn't.
    Ecological validity. Participants, tasks and tools may differ from deployment.
    Complacency over time. Vigilance degrades with sustained exposure to reliable automation.
    Metric gaming. Compliance with the indicator replaces the property it stands for.

    Evidence for several of these comes from other domains, such as visual search and human factors, so the paper treats their transfer to AI oversight as a hypothesis.

Key takeaways.

  • The gap is about proof, not policy.

    Leaders and staff diverge most on whether oversight can be verified. Sixty-five percent of managers and executives said AI training changed day-to-day behavior, compared with 41% of individual contributors. On whether someone could explain how the AI reaches its outputs, the split was 67% to 44%.

  • It is not a vocabulary problem.

    We tested whether managers simply know AI governance terminology better. Familiarity with the vocabulary does not explain the gap: it held at nearly the same size among respondents who recognized a real governance term and those who did not.

  • Experience helps, but it does not close the gap on its own.

    Confidence rises with the number of years an organization has used AI. Respondents who did not know how long their own organization had used AI were the least confident of any group.

  • The gap appears in every sector and is widest in healthcare.

    It was statistically significant in healthcare, finance and insurance, and manufacturing and operations. In healthcare it was driven by verification, ownership, explainability, and training: the evidence an accreditation audit or regulator would ask to see first.

Research approach.

Behavieural fielded a screened survey of 770 working professionals involved in evaluating, overseeing, or governing AI at their organizations. They worked in healthcare (n = 296), finance and insurance (n = 186), manufacturing and operations (n = 207), or other industries, and held individual contributor (n = 449) or manager-and-above (n = 320) roles.

Respondents rated their agreement with ten statements about AI oversight on a five-point scale, answered follow-up questions about ownership, incidents, and vendor verification, and described AI failures in their own words. Differences between groups were tested for statistical significance, and only differences meeting p < .05 are reported as findings.

This is a purposive professional sample, not a population-representative one, and it measures self-reported perception rather than audited practice. That distinction is part of the point: perception and evidence are exactly what come apart in the findings.

  • Read the Full Report

    The report is free. It includes the complete 10-statement breakdown for each sector, the literacy test in detail, the findings on incidents and vendor verification, and four evidence-backed moves that close the gap.