For many startup founders, the first attempt to understand the cost of artificial intelligence (AI) begins in the wrong place: the model provider’s pricing page.
They calculate the price of input and output tokens, compare one model against another, and try to forecast usage as if AI were a simple utility meter.
That approach may work for a chatbot answering one question at a time. It breaks down quickly when companies move into autonomous agents — software systems that can plan, retrieve information, run commands, inspect errors, and try again without constant human prompting.
In that world, the model call is often not the expensive part. The real bill sits around it.
Also Read: The app worked, the product didn’t: Can we install judgement into AI agents?
As Erik Perttu, Head of Engineering at Edu2Review, puts it in a recent Agoda report on AI adoption across Southeast Asia and India: “The generation is cheap; the trust-building around it is where the spend actually lives.”
That line captures a growing problem for engineering teams in the region. As AI coding assistants and agentic workflows become part of daily development, costs no longer come only from asking a model to write code. They come from everything needed to make that code usable, safe, and reliable.
The hidden cost of making AI useful
The report points to a case study involving an independent engineering pipeline working on a routine software task: renaming a variable across a dozen interconnected files.
On paper, this is exactly the kind of job AI should handle cheaply. The raw large language model calls needed to generate the code modifications cost about US$0.50. But once the full workflow was measured, the picture changed. Automated context retrieval, prompt construction, syntax parsing, multi-pass test validation, security scanning, and human review accounted for close to 90 per cent of the total financial and computational spend.
In other words, most of the cost was not in generating the answer. It was in proving that the answer could be trusted.
This matters because autonomous agents behave very differently from single-turn assistants. A coding assistant might suggest a function. An agent tasked with resolving a software issue may inspect local files, search dependencies, execute terminal commands, read test failures, re-prompt itself, and produce several corrective patches before stopping.
Each loop can be useful. Each loop also consumes resources.
Without guardrails, these systems can burn through budgets in surprisingly ordinary ways. Teams may dump entire code repositories, database schemas, or raw application logs into a high-context model when only a small slice is needed. Agents may be allowed to retry a failing unit test 15 or 20 times before a human steps in. Static system prompts, API definitions, and architecture notes may be sent repeatedly instead of being cached.
Also Read: Language was never the problem: Inside SEA’s real AI adoption gap
The result is not a dramatic AI failure. It is something more mundane: death by a thousand inefficient calls.
“The real risk isn’t AI being expensive. It is AI being used carelessly,” says M. Ridwan Agustiawan, Head of Engineering at Indonesian media-intelligence firm dataxet.
How dataxet made cost control an engineering habit
For dataxet, part of the Dataxet Group, the issue is not theoretical. The company uses agentic AI in production systems that support infrastructure and data-monitoring reporting. Its agents retrieve metrics from internal and client repositories, detect anomalies, translate raw data trends into executive narratives, and generate prioritised action recommendations for leadership.
That is a higher-stakes use case than a developer asking for help with boilerplate code. The output can influence operational decisions. It also requires the AI system to work across fragmented information sources, a familiar challenge for Southeast Asian companies dealing with multilingual markets, varied data maturity, and uneven legacy infrastructure.
dataxet noticed early that once developers became comfortable with AI tools, invoking an agent became the default response to many routine tasks. Individually, those requests looked harmless. Across an engineering organisation, they added up.
Rather than banning usage or imposing blanket restrictions, dataxet treated AI cost control as an engineering discipline.
The company narrowed the context fed into agents, using tightly scoped and pre-filtered data payloads instead of raw logs. It routed routine data aggregation to deterministic scripts — predictable software that does not need a reasoning model — and reserved high-reasoning large language model calls for anomaly interpretation and executive synthesis. It also tracked token consumption by feature, pipeline, and engineering workflow, making AI usage visible rather than abstract.
That visibility is crucial. Cloud computing went through a similar cycle. In the early days, teams spun up servers freely in the name of speed. The bills came later. The response was FinOps, a set of practices for managing cloud costs without killing innovation. Agustiawan sees a similar shift coming for AI: resource-conscious AI engineering.
Also Read: Why Singapore, Indonesia, and Vietnam are losing the AI race they think they are winning
For Southeast Asian startups, the lesson is particularly relevant. Many operate with lean engineering teams, limited runway, and investor pressure to show productivity gains from AI. The temptation is to either embrace agents everywhere or lock them down as soon as costs rise. Neither approach is sustainable.
Why quotas alone do not solve the problem
The Agoda report also highlights a split between how senior and junior technologists experience AI adoption.
Senior technology leaders (including CTOs, VPs, and architects) are more than twice as likely as junior developers to identify cost as the main barrier to agent adoption, at 32 per cent compared with 15 per cent. Junior developers, meanwhile, are nearly three times as likely to cite lack of skills, at 17 per cent compared with 6 per cent.
That gap helps explain why many organisations reach first for rationing. Across Southeast Asia and India, four in five developers now operate under usage limits, token quotas, or budget restrictions. Among large enterprises with more than 1,000 employees, more than 91 per cent enforce active usage caps.
Caps may be necessary, especially in companies where AI usage has spread faster than governance. But blind quotas can create a false sense of control. If a team is feeding bloated context into every prompt, using the wrong model for simple work, or allowing agents to retry indefinitely, a quota only slows the waste. It does not remove it.
Julius Domingo, Founder and CTO of Yappler, frames the issue more broadly: “The cost of AI not only involves the build, but also the data, process, and infrastructure preparation.”
That is the part many AI return-on-investment calculations still miss. A cheaper model is not always cheaper if it fails more often, requires more retries, or produces output that demands heavier review. A more expensive model may be economical if it completes complex tasks with fewer loops and lower downstream risk.
OpenAI’s Derrick Choi, Head of Codex Applied AI for APAC, makes a similar point in the report, noting that evaluation is shifting from simple token pricing to “how much useful work each dollar of intelligence can deliver.”
For engineering leaders, that means the metric cannot be tokens alone. Oravee Smithiphol, Tech Intelligence and Insights Manager at SCB 10X, the venture arm of Siam Commercial Bank, argues that organisations should assess cost per successful business outcome, weighing spend against code quality, delivery speed, and operational risk.
From AI spend to AI investment
The practical playbook is becoming clearer.
First, companies need to measure AI cost at the pipeline level. Token spend should sit alongside build status, test coverage, and deployment metrics, not in a separate finance spreadsheet reviewed only after the bill arrives.
Second, autonomous loops need termination gates. If an agent cannot fix a test after three attempts, for instance, it should escalate to a human developer rather than continue blindly.
Third, teams should use tiered model selection. Smaller or open-weight models can handle low-risk tasks such as documentation drafts, summarisation, or initial test scaffolding. Frontier models should be reserved for work that requires deeper reasoning, such as architecture decisions, security reviews, or complex debugging.
Also Read: Singapore firms embrace agentic AI, but audit trails remain thin
Finally, governance matters. The report finds that organisations with established AI guidelines show higher production adoption, at 43 per cent compared with 30 per cent, and stronger codebase readiness, at 56 per cent compared with 40 per cent.
The next phase of AI adoption in Southeast Asia will not be defined by who gives developers the biggest token allowance. It will be shaped by who builds the best systems around AI: cleaner context, smarter routing, visible consumption, and clear rules for when machines should stop and humans should step in.
The raw model call may be cheap. Trust is not. For startups hoping to turn AI from an experiment into an operating advantage, that is where the real work begins.
The post The hidden economics of autonomous AI agents appeared first on e27.
