Case Studies

Fixing the Human Side of the Model You Can't Touch: A Buyer-Side Trust Calibration Engagement in Commercial Lending

The Problem That Brought Them In

The client (a mid-market commercial lender; business loan underwriting division, ~120 underwriters) had licensed a widely-used, off-the-shelf AI credit risk scoring platform two years earlier—a large, multi-tenant vendor serving hundreds of institutions, with a standard contract that offered no custom development, no access to model internals, and no interface changes for any single mid-market customer. Whatever this engagement produced had to work entirely on the buyer's side of the relationship, because the platform itself was fixed.

The trigger was a fairness complaint pattern, not a performance metric. Over four months, the underwriting team had logged a rising number of internal escalations from relationship managers arguing that the tool's declines on small-business applicants "didn't make sense"—not because the decisions were provably wrong, but because underwriters themselves couldn't articulate a consistent reason for them beyond "the score said no." The platform's factor list was generic by design, built to apply uniformly across every institution licensing it, and had never been translated into language this bank's own staff could use with relationship managers or applicants. Two long-tenured underwriters had begun manually re-scoring every declined application by hand, effectively running a shadow underwriting process in parallel to the tool the bank was paying for—with nowhere official in the vendor's platform, or the bank's own policy, to record that this was happening.

Diagnostic Approach

Using the AI Trust Axis framework, the diagnostic focused on where trust breakdown was originating on the buyer's side, since the tool itself was fixed and unchangeable:

  • Cognitive Alignment: Did underwriters understand what the score did and didn't account for?

  • Autonomy Safety: Did underwriters feel authorized to exercise judgment within the bank's own policy, independent of the vendor's tool?

  • Fairness Comprehension: Could underwriters explain a decline to a relationship manager or applicant in the bank's own terms?

  • Interaction Effort: How much extra work was the shadow re-scoring process actually creating?

  • Failure Recovery Intelligence: What happened, internally, when a decline was later shown to be wrong on appeal?

Methodology combined the AI Trust Axis Scale, interviews with underwriters and relationship managers, observation of the escalation and appeal process, and review of 9 months of decline-and-appeal outcomes.

Key Findings

  • The bank had adopted the vendor's score as a decision, not an input. Internal policy had quietly evolved to treat the score as final, even though the bank's own credit policy technically permitted underwriter discretion—nobody had documented what "using discretion" was supposed to look like in practice, and the platform offered no structured way to record an override when one occurred.

  • Underwriters had no internal vocabulary for explaining a decline. The vendor's score shipped with one standard factor list, uniform across every client on the platform; the bank had never built its own translation from that generic list into language its underwriters could use confidently, so "the score said no" became the default explanation because no better one existed anywhere in the workflow.

  • The shadow re-scoring process was informal and invisible. The two underwriters running parallel manual reviews weren't following any approved bank protocol, and nothing in the platform surfaced that this workaround was happening—it wasn't visible to bank management, let alone to the vendor, until this engagement uncovered it.

  • There was no closed loop on overturned declines. When an appeal succeeded, the reversal wasn't fed back to the underwriting team, or captured anywhere the platform could use it, in any structured way—so lessons from past misses never accumulated into shared judgment on either side of the relationship.

The Intervention

Every element below was implemented entirely within the bank's own policies, training, and workflow, since the platform itself offered no configuration options and no vendor engagement was available at this account tier:

  • A documented discretion protocol, formally authorizing and defining when and how underwriters could depart from the vendor's score within existing bank credit policy—effectively building, on paper, the structured override record the platform itself had no field for.

  • An internal "translation layer:" a bank-authored reference guide mapping the vendor's generic score factors to the bank's own underwriting language, so staff had a consistent, defensible way to explain declines despite the platform's one-size-fits-all output.

  • A structured appeal-review loop: every overturned decline now feeds a quarterly case review shared across the underwriting team—a manual substitute for the appeal-outcome feedback mechanism the platform didn't provide.

  • Manager-led calibration sessions, run monthly, where underwriters discuss recent borderline cases together, building shared judgment about when the score was reliable and when it wasn't, entirely from the bank's own case history rather than any model-level insight the vendor could have supplied.

Outcomes

Measured over the 6 months following implementation, against the 6 months prior:

  • Informal shadow re-scoring dropped from an estimated 15% of declines to near zero, replaced by the documented discretion protocol.

  • Relationship-manager escalations citing "unexplainable" declines fell 47%.

  • Appeal turnaround time improved 22%, driven by underwriters having a consistent internal explanation ready rather than needing to reconstruct reasoning after the fact.

  • The proportion of appeals resulting in reversal fell from 34% to 19%—read by the bank as a sign that first-pass declines were becoming more defensible, not that fewer applicants were appealing.

  • Post-training AI Trust Axis Trust Integrity Score rose from 2.1 to 3.4 (out of 5).

Every gain here came from work the bank had to invent for itself—a discretion protocol standing in for a feature the platform didn't have, a translation guide compensating for a factor list built for hundreds of institutions rather than one, a manual review cycle replacing a feedback loop the platform never closed. None of it changed what the vendor's tool actually does; all of it worked around the fact that the tool wasn't built to be adapted, explained, or corrected for any single client's own vocabulary and judgment.

"We kept waiting for the platform to give us a better way to explain a decline. Once we accepted that wasn't coming, the real fix turned out to be entirely ours to build—we just hadn't realized how much of the problem was our own policy vacuum, not their model.”
— VP, Commercial Credit