Sounds Right, Points Wrong: What Expert Judgment Is Actually Worth When AI Makes Everything Seem Easy
Image: Depositphotos
Source: Thomas Seoh
Late last year, a founder who was building toward an FDA submission asked veteran regulatory and clinical strategist Shaherah Yancy—genuinely, not as a provocation—what she could tell him that ChatGPT could not. Yancy and a colleague, Heath Naquin, ended up recording a public conversation about exactly this question,[1] because they were both seeing what Naquin called "AI-informed decisions" quietly displacing "lived-experience decisions." Yancy's summary was blunt: building a regulatory strategy from a chatbot, she said, is like trying to build a plane from YouTube videos.
The question underneath that exchange is worth asking directly: what does an expert consultant actually offer, in an age when AI can produce a plausible-sounding regulatory and clinical development strategy on demand?
What AI does well
Skillfully used, AI handles a wide range of tasks in regulatory and clinical development faster, cheaper, and better (e.g., more completely) than humans:
Reviewing published literature and flagging relevant studies
Drafting first-pass protocols, clinical study reports, and SOPs from templates
Checking a submission's formatting and completeness before filing
Summarizing meetings and drafting routine correspondence—even used by FDA[2]
Sorting and coding adverse event reports as they arrive
Monitoring guidance updates and competitor filings
Translating documents for multiregional trials
Generating and checking statistical and data-cleaning code
On well-bounded tasks, with a checkable right answer, AI tools measurably outperform unassisted work.
What AI is not so good at
Some of the weaknesses of AI that have emerged with multiplying experience include:
Defaulting to generic, trend-driven answers on genuine strategic questions ("trendslop")—resistant even to adversarial prompting, and worst precisely on the most novel decisions where a tailored answer matters most[3]
Producing large volumes of polished-looking output that shifts the real work downstream instead of advancing it ("workslop")[4]
Filling a gap in what it knows with something fluent, rather than flagging that it doesn't know[2]
Missing a requirement it wasn't specifically asked to check for, rather than surfacing it as a gap in the work[5]
Weighing evidence and judging causation, as distinct from retrieving or summarizing a known fact[6]
It's not obvious when AI is failing
What makes AI particularly risky for high-stakes strategy and judgment is that good and bad output sounds identical to a non-expert—the same confident voice, the same volume of seemingly credible sources, no reliable signal for telling which is which. A large field study of consultants found the same tool producing opposite results depending on the task: strong gains on familiar work, and a real loss—worse than not using AI at all—on unfamiliar work just outside it, with no way for the consultants themselves to tell which kind of task they were on.[7]
AI mistakes in manufacturing quality or pharmacovigilance causation can be systematically checked and fixed, though that typically takes the resources of a larger company or specialized vendor. An increasing number of emerging companies instead use AI for development strategy, program prioritization, and drafting FDA submissions. And strategy failures don't leave the kind of evidence manufacturing and safety failures do: a decision that quietly points a program in the wrong direction produces no inspection finding at all—only, eventually, failure to meet an ill-advised endpoint.
Aviation treats every crash as a mandatory source of data: a flight recorder, an investigation, a public report, so the next flight is safer. Drug development has no equivalent requirement. When a program encounters a costly delay or dies altogether—especially in the case of an emerging company—the "recorder" is rarely found and analyzed to pinpoint that "an AI-built strategy caused this failure." Public case studies of decision failures like this aren't typically reported—not because they don't happen, but because there is no mechanism built into the system that would surface them even when they do.
What the data actually say
AI reasons well exactly where it has been fed enormous, information-rich datasets paired with fast, checkable feedback: millions of labelled photographs for image recognition, a program playing itself a million games of chess or Go and learning winning moves and principles. The recurrent pattern is volume in training data plus a clear immediate signal for what counts as correct.
Regulatory and clinical development has almost none of that structure. Except for a relative handful of the largest companies, a single company files a major application a handful of times in its entire existence. The outcome—approval or rejection—arrives years after the decisions that actually caused it, and even then it's rarely possible to isolate which specific choice mattered, because dozens of interacting factors (trial design, timing, which reviewer was assigned) fed into that result. And most of what would explain the outcome was never published in the first place. A model reasons only as well as what it was trained on, and in this field, the most useful data are often locked in closed loops.
FDA has the broadest possible internal vantage point on this industry—decades of applications, across every division and therapeutic area. Even so, its own internal AI tool, Elsa, built specifically for its staff and run inside the agency's secure systems, has been shown to hallucinate entire studies. Part of the reason appears to be mechanical, not incidental: the tool is walled off from proprietary and confidential submission material, and when it encounters those gaps in what it can see, there may not, at least currently, be a reliable way to flag them. Like other large language models, it tends to fill the gap with something fluent instead.[2] If the agency's own tool, with by far the strongest data position available to anyone, still does this, a generic model starting with none of that access is working from considerably less information, and filling in considerably more of it.
A large, established company can build an internal system trained on its own cumulative regulatory and clinical history. This is genuinely useful, but bounded to one company's experience. A generic cloud model available to a smaller company has neither that archive nor FDA's institutional one. It has the public record, which is demonstrably incomplete: for decades, FDA complete response letters - letters explaining why an application could not be approved as submitted - were generally confidential. Since 2025, FDA has published selected redacted letters, but the releases cover only recent years and a limited subset of applications, while the governing disclosure policy remains unsettled. They are a valuable new source for particular products and periods, but not a substitute for the far larger non-public record of submissions, agency interactions, and regulatory reasoning. A cloud model can learn from that public window; it cannot learn from the record it cannot see.
Experienced consultants occupy a distinct vantage point. Through repeated engagement with FDA, across many sponsors and many programs over years, they accumulate exposure to cross-sponsor, cross-division patterns that far exceed what any single company sees on its own, and much of that experience was never written down in a form a generic model could train on. Much of what a seasoned advisor knows about how a specific division responds to a specific kind of argument was learned by watching it happen, in person, over years of persuading, or failing to persuade, regulators.
There is also a reason this cannot simply be delegated, independent of how good any model eventually gets. FDA submissions require a named, responsible individual to personally certify their contents, and knowingly submitting false or misleading statements to the agency is a federal crime, not an internal compliance matter.[9] A company executive would not sign a higher-stakes contract generated by AI without understanding it with the help of experts, and this applies to regulatory submissions as well.
Why not save money by relying on AI research and drafting, then commissioning human expert review?
A reasonable-sounding plan: let AI draft, pay an advisor only to check it, get most of the speed for a fraction of the cost. In practice, it rarely works that way.
Software engineering has been running a version of this experiment for the past few years, with the advantage that review time is easy to measure. One large dataset of roughly 250,000 developers found that AI-assisted work got created 58 percent faster—and took 4.6 times longer to review.[10] The job of the 'expert human in the loop' didn't get easier because a machine wrote the first draft; a fluent draft doesn't announce which of its assumptions are wrong, so the reviewer has to check everything rather than skim the parts that look fine. The same dynamic plays out in regulatory and clinical strategy, at higher stakes: an AI-assisted draft handed to an advisor for "review" is really asking the advisor to first locate which parts of a plausible, internally consistent document are quietly wrong, before any actual strategic judgment can begin—a forensic pass through material that already reads as finished, layered in front of the same strategic work that would have been needed regardless.
The added cost runs wider than review time. AI generates faster than judgment can prune it, and more material reads as more thorough to someone who doesn't know better. Researchers have formalized a name for this: "workslop"—which, as noted earlier, shifts the burden of finishing work downstream. A 2025 study found 40 percent of employees had received it, at a cost of nearly two hours of rework each time, and that colleagues who send it are trusted less afterward.[4]
Curation and register are separate, harder-to-fake skills. FDA reviewers read hundreds of pages a day; knowing which three of seven possible arguments will actually move a specific reviewer—and getting the tone right, not just the content—is pattern recognition built from hundreds of submissions,[11] the same tacit judgment this piece has already argued a model doesn't have. An amateur pass includes everything defensible rather than what's persuasive—it can't tell which four arguments are noise, or which register will land.
None of this is new to AI—regulatory writing was expensive to fix long before generative models existed; one former FDA attorney called industry regulatory writing "horrific," describing billable hours spent decoding internal reports rather than editing them.[12] What AI changes is the volume of documents to review, and the gap between how a document reads and how much work it needs: a founder budgeting for a light edit is usually describing what a fluent draft looks like, not what it will take to fix.
AI in skilled hands does help, on the tasks listed above—a controlled study found it closed real ground for moderately experienced workers nearing expert level, but barely moved novices, since evaluating AI output takes judgment the tool doesn't supply.[13] Even in expert hands, AI-assisted work on a genuinely novel question tends to add review cycles, not remove them: the advisor is now validating both the strategy and the AI's execution of it. In inexpert hands, that cost doesn't disappear—it moves downstream, later and more expensively than if the advisor had been brought in from the start.
Why be penny-wise about the decisions that matter most?
Advancing a therapy from IND-enabling studies through a Phase II proof-of-concept is a high-stakes enterprise for an emerging biopharma company, representing the core use of its venture capital. The financial penalty for strategic missteps is unforgiving. Flawed regulatory logic or an ill-conceived clinical protocol rarely results in just a few days of delay; it triggers FDA clinical holds, costly mid-stream protocol amendments, or—worst of all—trials that finish on time but yield unpartnerable data because they missed the endpoints that actually matter to regulators or buyers. A single substantial protocol amendment in a Phase II trial costs a median of $141,000 in direct out-of-pocket fees and stalls trial progress for months.[15] If a strategic error forces a company to repeat a Phase II trial entirely, the financial loss typically runs between $7 million and $20 million.[16] Advisory fees for the expert regulatory and clinical strategy that prevents these unforced errors typically run in the five and six figures—roughly equivalent to the hard cost of just one avoidable protocol amendment. Scrimping to save a proportionately modest fee on an AI-generated first draft, which can quietly point a multi-million-dollar early development program toward a dead end, is exactly backwards.
So, what does an expert consultant actually offer that AI does not?
AI already does faster drafting, and an advisor who refuses to use it is behind, not careful. The value of expertise has shifted, as one recent analysis put it, from content to context[14]—and for regulatory and clinical strategy, that shift lands on specific places AI cannot reach on its own: where the relevant experience was never made public, so the training data simply doesn't exist; where a specific, named person has to understand and stand behind what gets submitted, not merely approve it; and where that judgment is applied before a plausible-sounding first draft exists, not after—because by the time a fluent, confident document is sitting in front of an advisor, a harder and more expensive part of the job has been commissioned.
That's the answer to the founder's question: AI can't know, or safely advise on, exactly the things worth knowing earliest, when they're cheapest to fix. Skimping there to save a modest fee on a multi-million-dollar project is backwards. Providing that early, trajectory-setting judgment is exactly what an expert is for.
Endnotes
1. Shaherah Yancy and Heath Naquin, "ChatGPT vs the FDA," originally published by University City Science Center; republished by RCI Global Partners, rciglobalpartners.com/chatgpt-vs-the-fda/ (accessed July 2026). Published transcript excerpt; the associated podcast episode is no longer accessible on Spotify.
2. CNN, "FDA's artificial intelligence is supposed to revolutionize drug approvals. It's making up nonexistent studies," July 23, 2025, cnn.com/2025/07/23/politics/fda-ai-elsa-drug-regulation-makary; U.S. Food and Drug Administration, "FDA Launches Agency-Wide AI Tool to Optimize Performance for the American People," press release, June 2, 2025, fda.gov/news-events/press-announcements/fda-launches-agency-wide-ai-tool-optimize-performance-american-people.
3. Angelo Romasanta, Llewellyn D.W. Thomas, and Natalia Levina, "Researchers Asked LLMs for Strategic Advice. They Got 'Trendslop' in Return," Harvard Business Review, March 16, 2026, hbr.org/2026/03/researchers-asked-llms-for-strategic-advice-they-got-trendslop-in-return. Six leading models tested against seven classic strategic trade-offs across more than 15,000 trials; the buzzword bias persisted even under adversarial prompting designed to defeat it.
4. Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano, and Jeffrey T. Hancock, "AI-Generated 'Workslop' Is Destroying Productivity," Harvard Business Review, September 22, 2025, hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity. Survey of 1,150 employees; 40 percent had received workslop, at an average cost of nearly two hours of rework per instance.
5. DLA Piper, "FDA Warning Letter Highlights Risks of Using AI in Drug Manufacturing," April 2026, dlapiper.com/en-us/insights/publications/2026/04/fda-warning-letter-highlights-risks-of-using-ai-in-drug-manufacturing.
6. Abate et al. (2025), cited in "Optimizing Large Language Models for Causality Assessment in Pharmacovigilance," arXiv, 2026, arxiv.org/pdf/2607.03704.
7. Fabrizio Dell'Acqua et al., "Navigating the Jagged Technological Frontier," field experiment with 758 Boston Consulting Group consultants, published in Organization Science, 2026; see also Ethan Mollick, "Centaurs and Cyborgs on the Jagged Frontier," oneusefulthing.org/p/centaurs-and-cyborgs-on-the-jagged. Consultants using AI outside its actual competence performed roughly 19 percentage points worse than consultants using no AI at all.
8. U.S. Food and Drug Administration, "FDA Embraces Radical Transparency by Publishing Complete Response Letters," press release, July 10, 2025, fda.gov/news-events/press-announcements/fda-embraces-radical-transparency-publishing-complete-response-letters; "FDA Announces Real-Time Release of Complete Response Letters, Posts Previously Unpublished Batch of 89," press release, September 4, 2025; PharmExec, "FDA Pauses Release of New CRLs Following Citizen's Petition," July 2026, pharmexec.com/view/fda-pauses-release-new-crl-following-citizen-petition.
9. 18 U.S.C. § 1001 (false statements to a federal agency), law.cornell.edu/uscode/text/18/1001; see also U.S. Food and Drug Administration, Form FDA 3674, "Certification of Compliance," which requires a named signatory and warns of criminal liability under this statute, fda.gov/media/74213/download.
10. Value Add VC, "AI Coding Productivity Study Data: What METR, McKinsey, and GitHub Found in 2026," citing Opsera's 250,000-developer dataset, 2026, valueaddvc.com/blog/ai-coding-productivity-study-data-what-metr-mckinsey-and-github-actually-found-in-2026.
11. WAYS, "Working with the FDA Reviewer in Mind: Why Reviewer-Friendly Submissions Matter," February 2026, waysps.com/post/working-with-the-fda-reviewer-in-mind-why-reviewer-friendly-submissions-matter.
12. Compliance Architects, "Writing For FDA Compliance Success: The Greatest Problem No One Talks About," March 2025, compliancearchitects.com/writing-for-fda-compliance/.
13. Harvard Business Review, "Gen AI Won't Make Your Employees Experts," March–April 2026, hbr.org/2026/03/gen-ai-wont-make-your-employees-experts. Controlled writing experiment at a fintech firm; AI narrowed the performance gap for moderately experienced workers approaching true expertise, but did little for novices.
14. Ravikiran Kalluri, "What's Your Edge? Rethinking Expertise in the Age of AI," MIT Sloan Management Review, October 20, 2025, sloanreview.mit.edu/article/whats-your-edge-rethinking-expertise-in-the-age-of-ai/.
15. Getz, K. A., et al., "The Impact of Protocol Amendments on Clinical Trial Performance and Cost," Therapeutic Innovation & Regulatory Science, vol. 50, no. 4, 2016, pp. 436-441. (Recent Tufts CSDD benchmark updates confirm these median direct costs persist, alongside tens of thousands of dollars in daily operating burn while a trial is stalled).
16. Sertkaya, A., et al., "Key cost drivers of pharmaceutical clinical trials in the United States," Clinical Trials, vol. 13, no. 2, 2016, pp. 117-126; supported by contemporary clinical research organization (CRO) benchmarking data placing average Phase II out-of-pocket costs between $7 million and $20 million depending on therapeutic area and trial complexity.
#RegulatoryAffairs #FDA #ClinicalDevelopment #GenerativeAI #BiotechStartup #RegulatorySubmissions
NOTE: A version of this article was originally published in Kinexions - the Kinexum Newsletter, https://mailchi.mp/kinexum.com/kinexions-summer-2026. It was previously posted on LinkedIn at: https://www.linkedin.com/feed/update/urn:li:ugcPost:7497651513054531584/
Thomas Seoh
Thomas Seoh is an entrepreneur/executive who has held senior leadership positions in public and private, pharmaceutical, biotech, and medical device companies for three decades.
Thomas holds an AB in Philosophy and History and a JD from Harvard University.