Representative image created using AI

AI can win the benchmark and still fail the customer

The score was almost certainly true; the mistake was treating a measurement of a model as a verdict on the system built around it.

by · India Today

An insurer puts an assistant in front of its renewals team. The model ranks near the top of the leaderboards; the demo was faultless. A customer asks whether a lapsed policy can be brought back. The assistant searches the company's documents and finds a reinstatement guide, long since superseded but never removed. It reads the guide, reasons from it correctly, and drafts an offer. The pricing engine finds nothing wrong with the format. The rule that sends unusual cases to a human keys on the amount, which is small. The letter goes out: polite, well structured, fully compliant with a policy that no longer exists.

That is a composite, not a documented case, and nothing in it is exotic. Where, exactly, did the intelligence fail?

The tempting lesson is that the benchmark misled everyone, and it is the wrong one. A benchmark does something science requires: it holds the world still. It fixes inputs, scoring and conditions so that one capability can be isolated and two models compared fairly. Whatever makes the score clean also narrows what it can say about a deployment. Controlled conditions isolate capability; organizations operate under uncontrolled ones. That is not a defect. The mistake lives in the extrapolation.

It helps to keep three questions apart. Whether a model can do a task under conditions that suit it is capability. How well it does across a defined set of cases is performance. Whether the deployed system keeps producing acceptable outcomes through the mess of real use is reliability, a word engineers have long used to mean a system performing its required function under stated conditions for a specified period. A leaderboard answers the first question with authority, the second partly, the third hardly at all.

In 2019, UC Berkeley researchers rebuilt the ImageNet validation set, following the original process closely. Widely used models lost around eleven points of accuracy on the new images; the best model fell from 83 to 72 percent. The authors looked for overfitting to the much-reused test set and found little sign of it. The new images were simply a little harder. Image classifiers are not language models, but the measurement lesson carries. The models had not changed. The world had shifted a little, and the scores followed.

Deployment cannot hold the world still. Customers change, documents get replaced, and last quarter's edge case becomes this quarter's Tuesday. The exam stays where it was. Reality keeps moving.

The insurer's model never looked at the company's policies. It looked at what was placed in front of it. Something retrieves, the model reads, and the retrieval can be wrong in ways the model cannot see. One benchmark study of retrieval-augmented systems found the models it tested poor at declining when the retrieved material lacked the answer. A capable model can reason competently from evidence that should never have reached it. If the guide is obsolete, a better reasoning score does not touch the problem. Matei Zaharia and colleagues noted in 2024 that the strongest results increasingly came from what they called compound AI systems: models wired together with retrieval, tools and other software. If the outcome comes from an assembly, a score for one part is evidence about one part.

Customers do not experience parts. In 2024 a British Columbia tribunal heard from a customer whose question to Air Canada's website chatbot about bereavement fares drew an answer contradicting the airline's own policy page. Air Canada argued it could not be held liable for what its chatbot said; the tribunal member wrote that this amounted to treating the chatbot as a separate legal entity responsible for its own actions, and awarded damages. The decision says nothing about what the chatbot was. To the customer, the airline was wrong. An engineer trying to fix it faces a harder question: which part was wrong?

A demonstration establishes that something can work; it says nothing about how often. A benchmark from the company Sierra puts an agent through simulated customer-service conversations, runs each task several times, and reports, alongside the usual success rate, an estimate of how often the agent would get the same task right every time. A GPT-4o agent using function calling succeeded on about 61 percent of retail tasks in a single attempt. Its estimated chance of succeeding on all of eight attempts was below 25 percent. A system that fails often but harmlessly is also a different institutional object from one that fails rarely and irreversibly; a leaderboard has no column for what a failure costs.

The shape of failure changes again once a system takes more than one step. In a question-and-answer tool a wrong answer ends there. In a system that plans, retrieves, calls tools and acts, a wrong early judgment becomes the premise for everything after it: it decides what gets searched for next and so what gets found, what gets written into a record the next step reads as fact, and whether a case looks routine enough to skip a human. An error acquires descendants.

In July 2025 a user who had told Replit's coding agent to make no changes reported that it deleted a live production database. Replit's chief executive, Amjad Masad, confirmed the deletion, called it unacceptable, and listed remedies including separating development and production databases and a planning-only mode. Both are architectural. A model that writes a wrong sentence creates one kind of problem. A system holding credentials can turn the same mistake into something that must be restored from backup.

The usual answer to all of this is a human in the loop, which is neither an admission of failure nor a fix. Humans are components with failure modes of their own, and a handoff only helps if it comes early enough, carries enough context to judge, reaches someone with authority, and arrives before the consequence rather than after. Researchers at Meta tested twenty frontier models on questions that should not be answered definitively, and found that models fine-tuned for step-by-step reasoning were markedly worse at declining than the models they were built from. A system that answers more questions is not necessarily a more reliable one.

Every evaluation contains a theory of what success is. Score answer accuracy and you learn about answer accuracy, and you have quietly decided that the answer is the unit that matters. The insurer never wanted accurate sentences. It wanted the right customer to receive the right decision under the policy currently in force, unusual cases to reach someone able to judge them, and mistakes caught while they were still drafts rather than commitments. The unit of success was never the model's answer; it was the completed institutional act. Choosing what to measure is choosing which of those things the system is for, a choice that is partly technical, partly commercial, and partly about accountability. No leaderboard can make it on an organization's behalf.

A thoughtful skeptic will object that better benchmark scores do predict better real-world results. In the ImageNet study every point gained on the original test set translated into slightly more than a point on the new one, and a large enough capability jump can make workarounds unnecessary. I take that seriously. The claim here is narrower: capability contributes to reliability but does not measure it. An engine test tells you a great deal about an engine and nothing about whether the aircraft is airworthy, because that also depends on the wings, the fuel lines and the pilot.

The practical question is less which model tops the table this month than whether the organization can say what a completed task looks like, what the system should do when it meets a case it was never shown, and what happens after it gets something wrong. Then it has to test that object, with its real documents and permissions and the people who will actually take the handoffs.

So the accusation in the title was unfair, deliberately. The benchmark did not lie. The model very probably earned its score. What happened next was ours. We asked how capable a model was under defined conditions and read the answer as though we had asked whether a deployed system could be trusted to complete our task under the conditions in which we actually operate. Those are different questions, about different objects. The insurer's model reasoned from what it was given. The failures that mattered were in what it was given, in what checked the result, and in the rule that decided no one needed to look.

The problem may not be that our measurements are bad. It may be that we have become very precise at measuring one object while deploying another.

Dr. Aditya Vikram Kashyap is AI Researcher and an AI Expert based in New York. Kashyap is an award-winning technology leader. His core competencies focus on enterprise-scale AI, digital transformation, and building ethical innovation cultures. Views expressed are strictly his own and do not reflect any entity or affiliations, past or present.

- Ends