The Maze: AI builders have spent two years selling autonomy. The firms that insure a claim, approve a loan or explain a diagnosis are now asking a less glamorous question: who owns the call when the model is wrong? A BCG analysis of 5,802 earnings calls shows “human in the loop” accelerating fastest in insurance from Q1 to Q2 2026. The phrase is becoming a liability-management signal, not merely an AI-product feature.
Insurance moved from 0.5% to 1.9% of calls mentioning the phrase. That is a +270.7% increase in one quarter, the biggest jump in the four-sector comparison. The absolute level still looks small. That is the point. A phrase only needs to enter a few executive conversations before it starts appearing in procurement, controls and operating-model work. For an insurer, the hard cases are not a long tail of customer-service tickets. They can be pricing exceptions, claims outcomes, fraud flags and decisions that a regulator or customer may later want explained.
The other regulated sectors are moving too, but from small bases. Healthcare rose from 0.2% to 0.5% (+127.4%), while financial institutions rose from 0.5% to 0.9% (+104.5%). Those growth rates should not be read as adoption rates. They do show where management teams are beginning to connect AI capability with a named person who can intervene. A review step is not enough. The reviewer needs the information, time and authority to challenge the output before the decision becomes an expensive fact.
TMT has the opposite profile: high starting awareness, little fresh urgency. It began at 1.4% and reached 1.6%, a +9.8% increase. Technology, media and telecom groups have talked about human review longer because they build, sell or operate the systems. The sharper shift is happening among companies that inherit the decision risk rather than write the code. That matters for commerce too: the retailer, lender, insurer or marketplace ultimately owns the customer outcome even when its AI stack comes from someone else.
“Human in the loop” is two different jobs. One is retrospective: audit a decision later, review a sample, update a policy. The harder one is in-the-moment: give a qualified person authority to stop or alter a consequential decision before it lands. The second model is expensive because the cases that escape automation are usually ambiguous, time-sensitive and senior enough to need real judgment. A queue alone does not solve it. A system can route the cases it finds uncertain; it cannot reliably identify the cases that are confidently wrong and catastrophically costly.
Why it matters: The next AI bottleneck may not be model performance. It may be exception capacity. If a system handles 90% of routine work, the remaining 10% can still hold most of the downside. Operators need escalation rules based on irreversibility, customer impact and decision authority—not a queue built from the model’s own confidence score. That requires a service level, a named owner and an override fast enough to affect the decision. It is management plumbing, but that plumbing decides whether AI can safely scale. It is unglamorous work, and it determines whether automation can scale safely.
Sources: Andreas Horn’s LinkedIn post and BCG CEO Data Point exhibit (Q1–Q2 2026; BCG Analysis, July 2026; n=5,802).
Images: Cover AI-generated


