Posted on Leave a comment

I built an AI that keeps receipts. The mistakes became the useful part

AI is remarkably good at producing answers. It is even better at sounding certain. Ask a difficult question and, within seconds, a system can gather information, connect ideas and return a polished explanation. Yet a harder question arrives later: what happens when reality proves the answer wrong?

I met that problem while building OnTheRice, a Singapore-based AI research publication. Part of its work involves making time-bound directional calls, saving the evidence available at the time, and checking the result later. The experience changed how I think about AI products. Producing an answer was the easy part. Preserving an honest record of it was much harder.

Freeze the answer before reality arrives

A prediction is easy to admire while its outcome is unknown. It is also easy to repair after the fact. A changed sentence, a missing timestamp or a quietly removed failure can turn poor judgement into a convincing success story.

We began locking each call before its result was known: the direction, entry price, source set, publication time and evaluation time. As at 1 September 2026, the public Founder Ledger showed 52 correct and 34 incorrect results across 86 resolved calls, or 60.47 per cent. 10 older records remained visible but excluded because they could not be verified properly.

Those figures are not proof that the system is exceptional. They are useful because they are incomplete, imperfect and inspectable. Once the misses remained on the page, wrong stopped being one category. Sometimes the reasoning failed. Sometimes the information arrived too late. Sometimes the market had already moved. Sometimes an event changed the conditions after publication.

The practical lesson is simple: save the original output, the evidence, the timestamp and the scoring rule before the outcome arrives. Otherwise, learning can become hindsight wearing a lab coat.

More links do not mean more evidence

A second problem appeared when the system began reading large numbers of reports. 10 websites may cover the same event, yet nine might trace back to one wire story. That is not ten witnesses. It is one witness with excellent distribution.

AI systems can confuse information volume with independent confirmation. Counting URLs rewards duplication. It can make a thin claim look strong simply because it travelled far.

Also Read: SEA’s AI boom has a water problem it cannot offset away

We started treating source origin as part of the evidence. Reports were grouped when they repeated the same underlying account, while genuinely independent reporting carried more weight. The practical rule is to trace claims backwards, not merely count how many pages repeat them. Three independent reports can tell us more than 100 copies.

Time belongs inside the evidence

Suppose a report at 8am says oil is falling and another at 5pm says it is rising. Which is wrong? Possibly neither. The world happened between them.

An AI system that ignores time can flatten both statements into one moment. Worse, it can use information published later to explain a decision made earlier. This creates the illusion that the system knew more than it could have known.

Useful records therefore need more than a source link. They need the time the source was published, the time the system found it, the time the output was made and the time it was assessed. This applies beyond markets. A business cannot honestly link a competitor’s price change to falling sales without knowing which came first.

Documentation is not decoration

This approach is not unique to one product. The researchers behind Model Cards for Model Reporting proposed standardised records of a model’s intended uses, evaluation and limitations. The NIST AI Risk Management Framework also treats accountability, transparency, validity and reliability as central parts of trustworthy AI.

The need is growing. Stanford’s 2025 AI Index reported 233 AI-related incidents in 2024, 56.4 per cent more than in 2023. The figure does not mean every incident could have been prevented by better records. It does show why explanations offered only after something goes wrong are not enough.

For an applied AI product, a small receipt can carry the source, timestamp, system version, original output, confidence level, known limits, evaluation rule and eventual outcome. None of this looks as exciting as a smarter model demonstration. It is far more useful during a dispute, audit or failure review.

Also Read: The app worked, the product didn’t: Can we install judgement into AI agents?

Let the system say I don’t know

AI products are designed to answer. Silence looks broken, especially in a demonstration. But forcing a decision when sources conflict or evidence is missing creates artificial certainty.

One of our hardest lessons was to separate wrong from unverifiable. The first means reality contradicted a recorded call. The second means the evidence is too weak to score it honestly. Combining them hides different problems; counting either as a win is worse.

A mature system needs permission to abstain. I don’t know yet should be a valid output when confidence falls below a clear threshold. Teams should record why the system abstained, then test whether the rule was too cautious or appropriately restrained.

Trust needs evidence

The AI industry is understandably focused on better reasoning, larger context windows and stronger models. Yet applied systems face a less glamorous test: can another person inspect what happened?

Can they see the original source? Can they see what the system said, when it said it and what changed afterwards? Can they find the failures as easily as the successes?

Building an AI that keeps receipts taught me that mistakes are not embarrassing leftovers. Properly preserved, they are training data for the product team, evidence for the user and a guard against self-deception.

A system should not earn trust by describing itself as intelligent. Show the work. Keep the misses. Let the record speak.

Editor’s note: e27 aims to foster thought leadership by publishing views from the community. You can also share your perspective by submitting an article, video, podcast, or infographic.

The views expressed in this article are those of the author and do not necessarily reflect the official policy or position of e27.

Join us on WhatsAppInstagramFacebookX, and LinkedIn to stay connected.

The post I built an AI that keeps receipts. The mistakes became the useful part appeared first on e27.

Leave a Reply

Your email address will not be published. Required fields are marked *