A growing number of financial institutions are exploring the integration of AI into their workflows with uneven outcomes. A widely cited 2025 MIT study found that roughly 95% of corporate AI pilots fail to deliver a measurable financial return.1 For boards and executives who have watched budgets flow into chatbots and copilots with little to show for it, the figure invites an uncomfortable question: Is AI a dud?
The evidence suggests the opposite. The high failure rates of these initiatives reflect how institutions deploy AI, not whether the underlying technology works. The AI capability curve is steepening, and the gap between institutions seeing real returns and those stuck in perpetual pilots comes down to a handful of choices that have little to do with which model a vendor sells.
Does AI Actually Work?
When considering the efficacy of an AI pilot, it helps to separate two things that are often conflated: the maturity of the technology versus the maturity of its deployment.
On the technology side, progress is accelerating. Independent benchmarking by METR, which measures the length of tasks AI systems can complete reliably, shows frontier models advancing from short, novelty-grade outputs a few years ago to sustained, day-long task execution today — work that, in human terms, runs beyond a full nine-to-five.2
On the deployment side, the picture is messier, and that is where the disappointment often stems from. Industry analysts place agentic AI somewhere on the early downslope of the classic hype cycle — past the peak of inflated expectations and working through the trough of disillusionment before the plateau of productivity.3 Historically, that journey takes two to five years. While adoption appears to be moving faster than in prior technology waves, being early nonetheless offers no protection against poorly orchestrated AI deployments.
Why Outcomes Vary: Purpose and Conviction
Institutions that have experienced positive outcomes following AI deployments often share two traits. First, they deploy with purpose. A useful illustration or measure of this would be whether an institution approaches the deployment goal with concrete KPIs. For instance, “increase NPS by three points” or “cut delinquency-management costs by six percent.” Those numbers carry a concrete purpose. On the other hand, launching “an AI chatbot for our website” lacks that directionality and clear goal. The data shows that roughly 78% of companies use AI, with typical impacts of below 10% in cost savings and under 5% in revenue uplift, and that only about 1% of U.S. firms have truly scaled AI across the enterprise.4
The differentiator is workflow redesign. An estimated 90% of the value captured by successful firms comes from reshaping and inventing workflows, not from sprinkling AI on top of legacy processes.5 High performers embed AI where decisions are actually made, across entire processes rather than via isolated pilots, and move deliberately from “humans using tools” to AI-supported workflows that humans orchestrate. Those same firms are roughly five times more likely to do strategic workforce planning for AI.
Second, they deploy with conviction. Successful AI is run as a CEO- and board-sponsored, multiyear program with a funded roadmap rather than via a series of hedged experiments. High performers are about three times as likely as their peers to expect transformative business change from AI within three years, and roughly one-third devote more than 20% of their digital budget to it while systematically tracking KPIs.6
For an institution considering AI implementation, these three practices can be most useful in driving meaningful impact down the line:
- Set a baseline metric, a target and a named owner for each AI deployment from day one.
- Redesign the underlying workflow end-to-end rather than automating a single step. For every process, label each task explicitly as AI-led, human-in-the-loop or human-only, with documented escalation rules.
- Pair that with a multiyear value thesis, for instance, “+3 points of NPS from AI in loan servicing,” and stage gates that define what evidence is required to move from pilot to scale.
What ‘AI’ Actually Means Today
Part of the confusion on this topic is linguistic. “AI” is more than just large language models; large language models are only one of roughly five families of machine-learning models, each suited for different jobs. Tree-based models work like auto-generated if-this-then-that rules and excel at fraud detection and delinquency-risk prediction. Linear and logistic models capture straightforward relationships and support risk scoring and segmentation. Neural networks recognize complex patterns and underpin default prediction and fraud detection. Unsupervised models find hidden structure in data, grouping borrowers by behavior or financial stress without being told the right answer. Large language models, the technology behind tools such as ChatGPT, predict the most likely next word in a sequence and, on their own, have no goals or understanding.
The more consequential shift for financial institutions is the arrival of AI agents. An agent is best understood as a workflow orchestrator that can think, decide and act, not merely respond. Agents break goals into steps, call systems, and enterprise data and adapt as they go. In servicing and collections, this translates into concrete roles: QA agents who review interactions against custom scorecards, loan-servicing agents who handle payment plans and hardship options, delinquency-management agents who tailor outreach, and back-office agents who process documents and complete routine tasks.
The New Risk Categories
Because agents take action across multiple steps within a process rather than simply generate text, they introduce risks that traditional model governance was not designed to catch. The most counterintuitive is compounding error. A 98% per-step success rate sounds excellent until five steps are chained together, at which point overall accuracy falls to roughly 90%. That means one in ten workflows would fail. In a regulated collections process, that is not an acceptable margin.
Three further failure modes deserve board-level attention. Agents can hallucinate or misrepresent policy, inventing non-compliant financial advice, referencing competitors the institution does not endorse, or making insensitive assumptions about a borrower’s circumstances. They can miss edge cases, failing to recognize manipulative language or legal traps; a borrower who says “I have no money to pay you” may be invoking a script that makes further contact a potential FDCPA or RFDCPA violation. And they can be manipulated directly, through prompt-injection attempts that try to hijack the agent into saying something off-brand or non-compliant.
Mitigations for these cases should be non-negotiable. Hard-coded guardrails and specially trained oversight models that check other models’ work can catch hijacking attempts and block poor responses before they reach a borrower. Rigorous evaluation in simulated environments also quantifies error rates across categories such as jailbreaking, hallucinations, content moderation, competitor references and policy-sensitive requests, such as bankruptcy or cease-and-desist notices. Finally, human-in-the-loop review remains essential for high-uncertainty cases. These measures build confidence in risk management by grounding error rates in concrete data and expressing them as a number that can be insured against.
What Regulators Expect
Supervisory expectations from the CFPB, FDIC, OCC and FFIEC are converging around four themes, and institutions should map any AI deployment against all of them:
- Explainability: The ability to explain, in plain language, how AI is used in a product or process; specific and accurate reasons for credit denials and other adverse actions, with no “black box” excuses; and documented model purpose, assumptions and limitations.
- Borrower Impact: Demonstrated testing for disparate impact and unfair bias across protected classes; ongoing monitoring of approval rates, pricing and servicing outcomes; and clear disclosures, opt-outs where applicable and easy access to a human review or complaint path.
- Model Governance: Treating AI as “models” under SR 11-7-style guidance, with an inventory, named owners and use-case documentation; independent validation before go-live and regular back-testing for drift and performance; and board and senior-management oversight.
- Data Privacy: Lawful data use under FCRA, GLBA and UDAAP, with limits on surveillance and non-traditional data; strong vendor controls, including rights to audit, explain, remediate and shut down AI systems; and cybersecurity, access control and logging around training data, prompts and outputs.
Where the Opportunities Are
When deployed with clarity and discipline, AI can drive material gains across an institution. Use cases most beneficial for banks include customer service, where omnichannel agents deliver consistent, compliant responses around the clock; fraud and risk monitoring, with real-time anomaly detection and lighter manual-review loads; loan servicing, with instant, accurate answers and automated follow-ups that lower handle times; delinquency management, where outreach adapts to borrower intent and personalized payment plans improve cure rates; and back-office automation, covering data entry, document processing and QA scoring.
In continuous QA, every voice or chat interaction is scored by an evaluation agent against criteria tied to regulatory and internal policy, flagging non-compliance, surfacing trends on team and topic dashboards, and routing feedback to supervisors. In automated servicing, an agent captures caller intent in real time, retrieves the next-best response from approved knowledge bases and servicing guidelines, and writes updated information, such as payment status, promises-to-pay and address changes, back to core systems so the next interaction starts with the current context. In both cases, the value comes from redesigning the whole loop, not affixing a chatbot onto the front of it.
The Board’s Checklist
AI is neither a silver bullet nor a dud; it is a capability that rewards purposeful, well-governed deployment and punishes the opposite. Before approving an AI program, boards should confirm that management has four things in place: a documented AI strategy tied explicitly to borrower value and operational efficiency; a governance framework aligned with regulatory expectations; robust vendor controls, testing protocols and model-lifecycle oversight; and clear policies for privacy, fairness, explainability and incident escalation. The institutions that get this right will not be the ones that adopt AI earliest or spend the most. They will be the ones who deploy it with clear objectives, redesigned workflows and humans firmly in the loop.
Michael Goh is the cofounder and CEO of Krew. Prior to Krew, he led speech benchmarking at Artificial Analysis, an Andrew Ng-backed company, where he worked alongside OpenAI and Amazon to conduct performance and quality evaluations for their multimodal LLMs, focusing on speech inputs. Michael earned his MS in computer science from the University of Chicago, where he was a Global Fellow, and his BA in economics and management from the University of Oxford, where he was a Fung Scholar.
- MIT, “State of AI in Business 2025” (Challapally et al., 2025).
- METR, “Measuring AI Ability to Complete Long Tasks” (2025).
- Gartner, “Hype Cycle for Artificial Intelligence” (July 2025).
- The Wall Street Journal, “Companies Are Struggling to Drive a Return on AI. It Doesn’t Have to Be That Way” (2025).
- Boston Consulting Group, “The Widening AI Value Gap” (2025).
- McKinsey & Company, “The State of AI in 2025: Agents, Innovation and Transformation” (2025).



