Whitepapers, conference research, and technical notes on AI in health economics and outcomes research, from the team building the systems.
Presented at R for HTA 2026. A published health economic model description is incomplete by design, assuming the reader brings the tacit modeling knowledge that was never written down. Hand that same description to an LLM and it fills those gaps silently, producing code that runs whether or not the interpretation is right. This talk lays out a decomposition framework that breaks the translation into discrete, inspectable stages and surfaces the points where a model description is ambiguous. Used this way, the LLM becomes a research companion that keeps the modeler in control of every judgment call and enables full transparency, traceability, and reproducibility, so the resulting R code is something you can actually trust.
This presentation, delivered at the ISPOR 2026 Workshop, explores a paradigm shift in health economic modeling, using AI not as a standalone tool, but as a reasoning engine within a structured multi-agent platform. It demonstrates how researchers can collaborate with AI to develop robust models in both Excel and R. Crucially, this approach generates more than just executable code, producing a comprehensive paper trail. By combining deterministic scaffolding with AI-driven extraction and human-in-the-loop oversight, the platform automatically generates full documentation of user interactions, model assumptions, data sources, design choices, and rigorous validation and verification results. This is a collaborative model development platform built for the scientific rigor of HEOR.
The widespread adoption of AI tools in HEOR has created a critical problem, because common evaluation metrics were designed for simpler tasks. This paper explores where traditional metrics fall short in complex, multi-step AI research workflows and what should complement them.
The field agrees that accuracy metrics cannot tell you whether an AI research tool is trustworthy. This paper proposes a confidence layer built on four properties that make reliability visible and verifiable, embedded during construction rather than bolted on.
This guide explores the shift from treating AI as a mere tool to a collaborative teammate. Learn how to effectively brief an AI by leveraging the delegation skills you already possess, setting the right context, and establishing clear output expectations.
When organizations evaluate AI-enabled health economic modeling tools, the question most frequently posed is whether the AI can replicate a published model. The appeal is understandable, but replication alone is an incomplete test of whether a model, and the workflow that produced it, can actually be trusted.
A hallucinated transition probability reads identically to an evidence-based one. This white paper argues that responsible LLM use in HEOR requires structural workflow architecture, not post-hoc review.
True power in Generative AI comes from asking the right questions, not just using the tools. This guide outlines the five foundational skills every Health Economics and Outcomes Research (HEOR) team must master to successfully navigate and leverage AI technologies.
This paper explores the hidden crisis of 'digital graveyards' where valuable research insights are lost. It argues that the true power of GenAI lies not just in automating tasks, but in revolutionizing how institutional knowledge is accessed and utilized by everyone.
Foundation LLMs present significant challenges for scientific research, from hallucinations to a lack of reproducibility. This presentation outlines a methodical framework for transforming general-purpose LLMs into reliable, high-quality research tools built for the scientific rigor of HEOR.
LLM hallucinations are a multifaceted challenge, not a single problem. This brief decodes the primary causes of AI-generated misinformation, from imperfections in training data to suboptimal context management, and outlines targeted, practical solutions for building more reliable applications.
From "One" to "Many": this presentation explores the shift from single-agent LLMs to multi-agent systems. It details common design patterns, such as sequential systems, parallel systems, and independent validation, to solve complex HEOR workflows effectively.
The term 'AI agent' is often misused, leading to confusion and poor tool selection. This guide clarifies the difference between structured AI workflows (like a systematic review) and adaptive AI agents (like qualitative research), providing a practical framework for health researchers to choose the right tool for their specific task.
GenAI is more than just a tool. It requires a new set of skills for researchers. This presentation outlines a practical upskilling journey, covering foundational principles, the balance between prompting and coding, and the importance of hands-on experimentation to truly harness the power of AI in a scientific context.
This practical presentation demonstrates a specific use case for augmenting the creation of health economic model reports. It showcases a concrete workflow where LLMs assist in drafting and reviewing report sections, integrating directly with existing processes to improve efficiency while maintaining essential human oversight.
The choice between off-the-shelf AI and a custom solution is a strategic decision between a consumer-grade product and an enterprise-grade system, not just a technical preference. This white paper decodes the critical risks of generalist tools, from their 'black box' logic to data integrity gaps, and outlines a framework for building purpose-built AI that delivers true control, integration, and auditability for specialized, high-stakes work.
Retrieval-Augmented Generation (RAG) is more than a simple pipeline. It is a sophisticated workflow that requires a deliberate architectural approach to succeed at scale. This white paper provides a blueprint for moving beyond naive prototypes to build advanced, enterprise-grade RAG systems, aimed at building AI that is accurate, efficient, and truly reliable.
The debate over Large Language Models isn't about hype versus fear. It is about a fundamental misunderstanding that leads organizations to treat them as standalone 'magic boxes,' resulting in staggering project failure rates. Using a powerful Formula 1 analogy, this white paper reframes the LLM as a high-performance engine and provides the systems-based framework for building the complete race car, one that delivers reliable workflow enhancement, strategic agility, and a durable competitive advantage.
Many organizations face what can be called the 'AI Trust Paradox,' where the potential efficiency of AI is often lost to the time-consuming manual verification needed to meet scientific standards. This whitepaper explores a framework called the 'LLM-as-a-Judge' model as one way to address this challenge.
This whitepaper argues that the 95% failure rate of enterprise GenAI projects is not due to model limitations, but rather a lack of 'Context Engineering,' a systematic architectural approach that prioritizes data governance, reliable retrieval, and structured memory over simple prompt crafting.
Nothing matches that combination. Clear the search or switch topics to see more.