
Artificial intelligence has dramatically changed the way startups are built.
Today, a founder can describe a product idea, open a tool such as Claude or OpenAI, generate hundreds of lines of code, build a prototype and present it as an “AI-powered innovation” within days. What once required a technical team, months of development and significant capital can now be achieved remarkably quickly.
The barrier to building software has never been lower.
But as the ability to build accelerates, another challenge becomes increasingly important: validation.
AI can generate code, analyse information, summarise complex material and produce remarkably convincing answers. But a convincing answer is not necessarily a correct one. And when AI-generated outputs move from a demo environment into industries where mistakes have real consequences, the difference between “working” and “working reliably” becomes critical.
Healthcare is perhaps the clearest example.
For a healthcare AI startup, the question is not simply whether a model can produce an answer. It is whether that answer is clinically reliable, generated from appropriate data, reproducible across relevant populations and settings, understandable to the intended user, and safe enough to inform a clinical decision.
This is where human-in-the-loop verification becomes more than a safety feature. It becomes part of the product itself.
The clinician is not the bottleneck
There is sometimes an assumption that AI automation becomes more valuable as humans are removed from the workflow. In healthcare, that assumption can be dangerous.
AI can process enormous volumes of information far faster than a human. It can identify patterns across patient records, compare information against large knowledge bases and surface potentially relevant findings. But it does not automatically understand the complete clinical context in which those findings will be used.
A clinician does. That is why the most useful healthcare AI may not be the system that attempts to replace clinical judgement, but the one that augments it.
Regulation is increasingly reflecting this distinction. The US Food and Drug Administration’s January 2026 final guidance on Clinical Decision Support Software clarifies the criteria for certain clinical decision-support functions to qualify as non-device software. One important criterion is whether the software enables healthcare professionals to independently review the basis of its recommendations rather than primarily relying on the software’s output.
Singapore’s Health Sciences Authority has similarly refined its framework for Software as a Medical Device (SaMD) and Clinical Decision Support Software (CDSS), as outlined in its update on SaMD risk classification and CDSS qualification guidelines. Its July 2025 revision added, among other changes, a criterion concerning whether CDSS recommendations are based solely on established clinical guidelines when determining whether software qualifies as a non-medical device.
The message for founders is important: automation does not automatically mean removing the professional from the loop.
Depending on its intended purpose, functionality and risk, software may fall within medical-device regulation or qualify for a non-medical-device pathway. Either way, the product needs to be designed around clearly defined accountability.
The clinician interprets the recommendation, considers the patient’s circumstances and decides whether to accept, modify or reject it.
Also Read: Vietnam’s healthtech boom has a talent problem nobody is talking about
“Accurate” is not enough
AI hallucination is often discussed as a technical problem. In healthcare, it is a product and safety problem.
A generative AI system can produce an answer that is fluent, structured and persuasive while being completely wrong. A fabricated reference, incorrect interpretation of a medical record or inappropriate recommendation could have consequences far beyond a poor user experience.
This means healthcare AI cannot be evaluated simply by asking: “How accurate is the model?”
The more useful questions are: Accurate for whom? Under what conditions? Compared with what reference standard? Using which data? And in which real-world population?
Traditional metrics such as sensitivity, specificity, precision, recall and area under the receiver operating characteristic curve (AUC) remain valuable. But a strong metric on a controlled dataset does not automatically translate into reliable performance in clinical practice.
A model can perform exceptionally well in one dataset and behave differently when exposed to another hospital, patient population, imaging device, documentation style or clinical workflow.
This is why validation must extend beyond the model itself.
It includes the quality and representativeness of the data, external validation, clinical workflows, human factors, usability, monitoring and performance after deployment. For regulated software, lifecycle management, verification and validation, change management and post-market considerations are increasingly important parts of the development process. HSA, for example, maintains a lifecycle-oriented framework for software medical devices alongside its SaMD and CDSS classification guidance.
The new startup moat may be trust
For founders, this creates an important strategic shift.
The competitive advantage of an AI startup may no longer be simply how quickly it can build a model.
If thousands of startups can use the same foundation models and increasingly powerful coding tools, the ability to produce a prototype becomes less differentiated.
The harder question becomes: Can you prove that what you built works?
That proof may become one of the strongest forms of competitive advantage.
A startup that combines AI automation with genuine domain expertise, structured validation, transparent outputs and continuous monitoring can build something considerably more defensible than a product that simply places a large language model on top of an existing workflow.
This is particularly relevant for founders entering regulated or high-stakes industries. Domain experts should not be brought in merely to satisfy an advisory requirement after the product has been built. Their expertise should influence the product architecture, validation strategy, workflow design and definition of failure.
Also Read: Healthtech in South and Southeast Asia – Seeing beyond the “obvious”
Human-in-the-loop should therefore not be viewed as a limitation on AI.
It is a mechanism for making AI deployable.
The best systems will know what to automate, when to request human verification and, critically, when not to provide an answer at all.
From “AI versus humans” to “AI plus humans”
The future of AI in healthcare is unlikely to be a simple contest between artificial intelligence and human intelligence.
It is more likely to be a carefully designed partnership.
AI brings scale, speed and the ability to process enormous amounts of information. Humans bring contextual understanding, professional judgement, ethical responsibility and the ability to challenge an output when something does not look right.
The real innovation lies in designing the interface between the two.
As AI lowers the cost and time required to build software, founders will increasingly be judged on something beyond how quickly they can produce a demo.
They will be judged on whether they can demonstrate that their product works, understand where it can fail, and build the mechanisms to detect and manage those failures.
In an age where almost anyone can build software with a prompt, building is becoming easier. Proving is becoming harder.
And for high-stakes AI, that may be where the real startup advantage lies.
—
Editor’s note: e27 aims to foster thought leadership by publishing views from the community. You can also share your perspective by submitting an article, video, podcast, or infographic.
The views expressed in this article are those of the author and do not necessarily reflect the official policy or position of e27.
Join us on WhatsApp, Instagram, Facebook, X, and LinkedIn to stay connected.
The post The missing layer in AI innovation: Human verification appeared first on e27.
