Confidently wrong: the hard half of AI adoption
I build AI into my own financial modelling and forecasting, daily, on a live P&L. Capability was never the constraint. Trust was. An AI that produces a fluent, plausible, precisely formatted answer that is quietly wrong is more dangerous than one that is obviously wrong, because the obviously wrong answer gets caught and the quietly wrong one gets actioned.
What actually made it reliable
The fixes were unglamorous.
Context. Feeding the model the real, specific ground truth, not letting it fill gaps from thin air. Most confident wrongness starts as a gap the model papered over.
Guardrails. Clear limits on what it may assert and where it must defer. In my own practice this has hardened into five standing controls: every output states its assumptions and sources; material analysis is cross-examined by independent models, with a blind judge on the decisions that matter; nothing informs a decision until validated against source data; misses are logged so the workflow improves; and confidential data stays within arrangements its owner has approved. Familiar instincts, new instrument.
A human at the point of consequence. The model drafts; judgement decides. That boundary is not a limitation of the technology. It is the design.
The regulators reached the same answer
What struck me is that this is almost exactly the framework the FDA has started to codify for AI in drug development: a risk-based credibility standard built on transparency, validation and human oversight. Practitioners who learned by doing and regulators who learned by reviewing are converging on the same architecture, independently. When that happens, it usually means the answer is right.
The lesson generalises well beyond pharma. In finance, in drug safety, anywhere the output matters: the value of AI is capped by the quality of your guardrails, not the power of the model. Model capability is now abundant and cheap. Governed reliability is scarce and valuable, and it is built, not bought.
I have written up the full working practice, the five controls and the modelling stack they govern, in the long-form piece on this site.
In brief
Why is "confidently wrong" the real risk? Because fluent, well-formatted wrong answers pass the glance test and get actioned. Obvious errors get caught; quiet ones compound. What stops it? Ground-truth context, explicit guardrails on assertion and deferral, and a human at the point of consequence. In my practice: five standing controls, run daily. Is this compatible with regulation? It is where regulation is going: the FDA's draft credibility framework for AI rests on the same three pillars of transparency, validation and human oversight.
Source: the FDA's 2025 draft guidance on AI in drug development (fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development). Companion piece: What AI actually changes in FP&A (https://jatinderpurewal.com/insights/ai-in-fpa/).
© Jatinder Purewal 2026. All rights reserved.
I take these conversations directly: get in touch.
More Insights · Discuss on LinkedIn: linkedin.com/in/jatinderpurewal