article image

Your AI Is Not Guessing. It Is Reading Bad Data Very Confidently.

Two measured failures, one bad assumption, and the four conditions that have to be true before AI gets an operational question right

09:41, and a customer is on the phone to the sales desk.

Customer: Any chance this one could go out today?

Salesperson: Let me check for you.

The salesperson types the question into the AI assistant the business switched on last month. Back it comes in under a second. Clear English. A number in it. The calm tone of a colleague who has just gone and looked.

The AI assistant: Yes. Forty units on hand, and a vehicle on that route this afternoon.

Salesperson: Yes, no trouble at all.

Nobody went and looked.

At 16:20 the order is still sitting on the warehouse floor. By then four different people have each spent twenty minutes finding out that it was never going anywhere.

Nobody did anything wrong. That is the part worth staying for.

What was that yes actually standing on?

"Can this go out today" sounds like one question. It is at least five.

Is the stock physically there. Is it free to sell, or already promised to somebody else. Is there a vehicle going the right way with room on it. Has anybody planned that route yet. And will the truck start and the driver turn up.

Five answers. They live in five different places, and every one of them is perfectly capable of being wrong on its own.

The AI did not get five attempts at this. It gave one answer, built out of five separate pieces. And it handed that answer over with the same confidence it would have used if all five had been sound.

One yes at the sales desk commits six teams to a day that none of them have agreed to yet.

Here is the uncomfortable part. A report that is wrong at least shows you its workings. The dates are there. The filters are there. The column that does not add up is sitting in front of you, not adding up.

A sentence gives you none of that. It arrives finished and tidy, with nothing to squint at and nothing to disagree with. Sounding certain is not the same as being right. It is just a very convincing impression of it.

We have made this argument before. Scattered data is why AI keeps disappointing people, and we set that out in you don't need more software, you need one version of the truth. What follows is the evidence underneath it. It turns out to be worse than we said.

Two measured failures, and they multiply

The first failure is that AI is not especially good at finding the right answer inside a real company's records.

Researchers ran a test where AI systems had to answer plain-English questions using real business records. Not tidy practice data, where the column marked customer name reliably contains a customer name. Real records, with values entered inconsistently, abbreviations nobody wrote down, and three spellings of the same city.

A person who knew the data scored 92.96%. The best AI available at the time scored 40.08%.

That was 2023, and this field does not sit still. The systems have improved a great deal since, and the published scores have climbed a long way. They still have not caught the human.

And the score was never the interesting finding anyway. What made the questions hard was not the English, and it was not the arithmetic. It was working out what the entries in the records actually meant.

The second failure is that the number it finds is usually not true in the first place.

A 2025 study of grocery retail compared stock records against what was genuinely sitting on the shelf. Just 35.3% of records were accurate. A quarter of them claimed more stock than was actually there.

A second study the same year found between 50 and 70% of individual product lines wrong at any given moment. Worth reading twice, that one. It means the normal state of a stock figure is roughly.

Now put the two failures together. Something that is imperfect at finding the number, pointed at a number that was never right, gives you a confident sentence with nothing visibly wrong with it.

We take the first half apart elsewhere in this series, and say what to do instead. The second half is a warehouse discipline problem rather than an AI one. It does not improve on its own.

And the plan was optimistic to begin with

Both of those failures assume the only problem is whether a record is accurate. There is a bigger assumption underneath, and it is this. An operation runs as designed, for ever, so the schedule is a description of what is going to happen.

It is not. A schedule is a forecast of a day that has not happened yet. It was written by somebody who did not know the driver was going to be ill.

The truck will not start. The licence expired on Friday and nobody was watching the date. Ask anybody who has run a dispatch desk what their morning consisted of. Not one of them will describe following a plan.

An AI has no idea it has inherited that assumption. Nobody told it the plan was optimistic. The full version of this argument is the next piece in this series, including why a breakdown stays where it happened instead of reaching the people still selling capacity.

So what actually went wrong at 09:41?

Nothing that anybody in the building could have seen.

The AI answered the question it was asked, using the information it was given. The record said forty. The salesperson trusted a system her company had just bought specifically so that she could trust it. The warehouse found out at half past two.

So "the AI got it wrong" is almost always the wrong diagnosis. And a cleverer AI will not rescue you. A cleverer one reads the same thirty-one items and still calls them forty.

What would have had to be true at 09:41?

Four conditions. None of them are about the AI itself. That is the part people find irritating, because the AI is the part with the budget attached.

One place to look for each number. Say stock lives in the warehouse system. It also lives in the accounting system. And in a spreadsheet somebody keeps out of a well-earned mistrust of the other two. Then there is no such thing as the stock figure. There are three figures and an argument. Ask an AI and it hands you one of them, with no hint that the other two are out there disagreeing.

The number has to be current, and the answer has to say how current, and whether it is a plan or a fact. Most reporting is built on figures pulled together overnight. That is exactly right for spotting a trend, and no use whatsoever for making a promise. A vehicle that was roadworthy last night is not a vehicle that is roadworthy now. A driver who was on the rota on Friday is not a driver standing in the yard this morning.

The answer has to respect who asked. Three people can ask about the same order. A salesperson wants the margin on it. The customer wants to know where it is. An operations manager wants to know what it does to their week. Three different questions, all sitting on top of the same records. Adding a chat box to systems that were never built to know who is allowed to see what does not quietly teach them. That one gets a piece of its own later in this series, because it is the objection every buyer raises and almost nobody answers concretely.

The AI must not be making up the sum. The one that gets skipped, and the one that matters most.

Go back to that test for a moment. What it was scoring is an AI taking your question, writing its own instruction to go and fetch the answer, and running it. That is the thing it is measurably bad at. It is also, almost universally, how "ask your data anything" gets built.

There is a way round it. The business works out in advance how each of its important numbers is calculated, once, properly. Available stock. Free capacity on a route. Deliveries due today. The AI then works out which of those the person was asking for, and what to put into it. It picks from a menu. It does not invent the recipe, and it never arrives at a figure of its own. A later piece in this series makes that case properly, and says who else is making it.

Notice what those four have in common. Not one of them is something you can buy afterwards and switch on. Each one is a consequence of how the systems underneath were put together, and that was settled years before anybody thought about asking questions in plain English.

So the usual order of work is precisely backwards. The AI is not the project. The AI is the easy bit that becomes possible once the project is done. The piece that closes this series turns that into a sequence you can check yourself against.

It is also why the salesperson at 09:41 was not the weak link. She was the last person in a chain of five systems with any chance of catching this. What she was handed was a sentence, not a set of workings.

How this runs on one platform

Illuminate was built down the length of the work rather than across the departments. We made that argument in full in what operations actually means. It turns out to matter here.

Depot WMS, Cargo TMS, Flow OMS and Ryse CRM are not separate systems sending each other updates. They all write to one shared set of records. So the stock figure, the quantity already promised, the space left on the vehicle and the route sit in the same place. One figure each, rather than an overnight negotiation between systems that were never built to agree.

The vehicle is in there too, which is the bit that usually gets left out. Every vehicle sits in the asset registry, Tagz AMS, with its current status, what it can actually carry, and its repair and service history on the same record. So a booked service or an open repair reaches the dispatcher before a load goes on that truck, not after.

The same applies to the people. A driver's record holds their licence and its expiry date, the hours they work, and a daily driving limit. The software that plans the routes reads those limits, instead of them sitting in a folder somebody has to remember to check. And when something goes wrong, the dispatch screen counts it alongside ready, dispatched and en route. Things going wrong is the normal weather of an operating day, not an error.

Every figure also says where it came from and who may see it. Anything about right now is read live from the operation, anything about trends from last night's totals. Finance sees cost. Operations sees how busy the week is. A customer in their own portal sees their own account and nothing else. All of it set up as settings rather than code, because the reporting was built this way long before anything conversational was on the table.

None of that is an AI feature. It is the floor any trustworthy answer has to stand on, whoever ends up building the thing that speaks.

The short version

An AI answering questions about your operation walks into two measured problems and one unexamined assumption. It is imperfect at finding the right number in real business records. Those records are full of numbers that were never true. And the plan it reads was always a forecast, written by somebody who could not see the day coming.

None of that gets solved by a cleverer AI, because none of it is the AI's fault. It gets solved by four conditions. One place to look for each number. Every answer saying how current it is. Permissions built in from the start rather than added later. And an AI that picks from figures the business has already worked out, instead of working them out for itself.

Get those right and the AI becomes the straightforward part. Get them wrong and you have bought a very articulate way of being misinformed. And somebody is still going to have to ring that customer at 16:20.

Frequently asked questions

Why do AI projects fail on operational data?

Usually because of the data rather than the AI. Operational information is commonly spread across several systems, and each one holds its own version of the same figure. It is often wrong even within a single system. Research on retail stock records has found accuracy as low as 35%. An AI reading that information produces confident answers built on incorrect numbers, and gives no signal that anything is wrong.

Can an AI be trusted to answer questions about stock or deliveries?

Only where two conditions hold. Each figure has to come from one agreed place, be up to date, and only be shown to people allowed to see it. And the system itself has to work out the figure, rather than the AI writing its own instruction to go and find it. Where those conditions are not met, how confident an answer sounds tells you nothing about whether it is right.

Why do vehicle and driver availability matter to an AI answer about delivery?

Because a delivery promise depends on a vehicle being fit to drive and a driver being there, and both are commonly recorded outside the systems that make the promise. Service records sit in a workshop log. Licence expiry dates and shift patterns sit in a spreadsheet or a group chat. So a truck off the road is known locally, while what it costs the business stays invisible to everybody upstream. Any system answering "can this go out today" has to hold vehicle and driver availability as part of the same picture, and has to treat things going wrong as normal rather than as an error.

Does better data quality matter more than a better AI?

For operational questions, generally yes. A more capable AI still reads the same records. Where those records are wrong, the only improvement is in how convincingly the wrong answer is expressed. Studies of stock record correction have found sales improvements of around 11% after a full count. That suggests the value sits in the accuracy of the underlying information, not in the sophistication of whatever reads it.

References

The rest of this series

This piece is the overview. Each one below takes a single argument out of it and goes all the way down, and each is published in turn over the coming weeks.

  • Why the plan stops being true by mid-morning - the dispatch desk, the fires, and why things going wrong belong in the system rather than in somebody's head.
  • AI Is Bad at Reading Databases. So Stop Asking It To. - what that test actually measured, why real company records are the hard part, and the approach that sidesteps the whole problem.
  • Who is allowed to know that? - answering operational questions without showing people what they should not see.
  • What to fix before you buy an AI assistant - the order of work, as a sequence you can check yourself against.

Two earlier pieces sit underneath this one. What operations actually means is on why the work runs down through a business while the ERP looks across it. You don't need more software, you need one version of the truth argues the same foundation from the data side.

Talk to us about what your data would have to be



For media inquiries, please contact:
contact@illuminate.ae

Illuminate Software Solutions was founded in Dubai in 2017, when a consulting and solution delivery business that had operated in Canada since 2000 turned its operations and automation experience into products. It operates from the UAE, India and Canada. Its eight products - Cargo, Ryse, Tagz, Flow, Lyst, Depot, Mrkt and Kart - cover delivery logistics, warehousing, order management, CRM, pricing, assets, point of sale and customer portals. Each runs on its own, and none need integrating with each other: they share one data model, so a delivery confirmed in one product updates inventory, invoicing and the customer's order in the others. Anything outside connects to that same layer in real time - ERP systems, carriers, marketplaces, IoT devices - which is what gives AI a connected foundation rather than fragmented data. More than 10 million deliveries and a quarter of a billion dollars in order value have been processed to date, and every product is built in-house by one team with more than 25 years in business operations. illuminate.ae