Confidence by Design: Practical Patterns for Trustworthy AI Research Systems

The first three papers in this series built the argument for rethinking how trustworthiness is constructed in AI-enabled research. This paper closes the arc by demonstrating the confidence layer built into a working multi-agent modeling system, where transparency, traceability, selective attention, and calibrated human oversight are each a load-bearing architectural decision, and the assurance record is produced as a byproduct of doing the work.

View Abstract & Download →

Publication

A Decomposition Framework for Translating Published HE Model Descriptions into Executable R Code (R for HTA 2026)

Presented at R for HTA 2026. A published health economic model description is incomplete by design, assuming the reader brings the tacit modeling knowledge that was never written down. Hand that same description to an LLM and it fills those gaps silently, producing code that runs whether or not the interpretation is right. This talk lays out a decomposition framework that breaks the translation into discrete, inspectable stages and surfaces the points where a model description is ambiguous. Used this way, the LLM becomes a research companion that keeps the modeler in control of every judgment call and enables full transparency, traceability, and reproducibility, so the resulting R code is something you can actually trust.

Health Economic Modeling · Trust & Validation June 2026 View Abstract & Download →
Publication

Use of Agentic AI to Create Health Economic Models (ISPOR 2026 Workshop)

This presentation, delivered at the ISPOR 2026 Workshop, explores a paradigm shift in health economic modeling, using AI not as a standalone tool, but as a reasoning engine within a structured multi-agent platform. It demonstrates how researchers can collaborate with AI to develop robust models in both Excel and R. Crucially, this approach generates more than just executable code, producing a comprehensive paper trail. By combining deterministic scaffolding with AI-driven extraction and human-in-the-loop oversight, the platform automatically generates full documentation of user interactions, model assumptions, data sources, design choices, and rigorous validation and verification results. This is a collaborative model development platform built for the scientific rigor of HEOR.

Health Economic Modeling · Agentic & Multi-Agent Systems May 2026 View Abstract & Download →
White Paper

The Validation Gap: What Accuracy, Sensitivity, and Specificity Cannot Tell You About AI Research Tools

The widespread adoption of AI tools in HEOR has created a critical problem, because common evaluation metrics were designed for simpler tasks. This paper explores where traditional metrics fall short in complex, multi-step AI research workflows and what should complement them.

Trust & Validation April 2026 View Abstract & Download →
White Paper

The Confidence Layer: A Design Philosophy for Trustworthy AI in Evidence Generation

The field agrees that accuracy metrics cannot tell you whether an AI research tool is trustworthy. This paper proposes a confidence layer built on four properties that make reliability visible and verifiable, embedded during construction rather than bolted on.

Trust & Validation April 2026 View Abstract & Download →
White Paper

The Art of Delegation in the Era of AI: How to Brief an AI Like You'd Brief a Colleague

This guide explores the shift from treating AI as a mere tool to a collaborative teammate. Learn how to effectively brief an AI by leveraging the delegation skills you already possess, setting the right context, and establishing clear output expectations.

Skills & Training April 2026 View Abstract & Download →
White Paper

Beyond Replication: What It Actually Means to Evaluate an AI-Generated Health Economic Model

When organizations evaluate AI-enabled health economic modeling tools, the question most frequently posed is whether the AI can replicate a published model. The appeal is understandable, but replication alone is an incomplete test of whether a model, and the workflow that produced it, can actually be trusted.

Trust & Validation · Health Economic Modeling April 2026 View Abstract & Download →
White Paper

Invisible Errors: How to Know When an LLM Gets It Wrong

A hallucinated transition probability reads identically to an evidence-based one. This white paper argues that responsible LLM use in HEOR requires structural workflow architecture, not post-hoc review.

Trust & Validation · Health Economic Modeling March 2026 View Abstract & Download →
White Paper

5 Foundational Skills Every HEOR Team Needs for GenAI

True power in Generative AI comes from asking the right questions, not just using the tools. This guide outlines the five foundational skills every Health Economics and Outcomes Research (HEOR) team must master to successfully navigate and leverage AI technologies.

Skills & Training March 2026 View Abstract & Download →
White Paper

From Task Replacement to Workflow Revolution

This paper explores the hidden crisis of 'digital graveyards' where valuable research insights are lost. It argues that the true power of GenAI lies not just in automating tasks, but in revolutionizing how institutional knowledge is accessed and utilized by everyone.

AI Strategy & Architecture February 2026 View Abstract & Download →
Publication

Transforming LLMs into Reliable HEOR Research Tools

Foundation LLMs present significant challenges for scientific research, from hallucinations to a lack of reproducibility. This presentation outlines a methodical framework for transforming general-purpose LLMs into reliable, high-quality research tools built for the scientific rigor of HEOR.

AI Strategy & Architecture January 2026 View Abstract & Download →
White Paper

Decoding LLM Hallucinations: Beyond Just "Lack of Context"

LLM hallucinations are a multifaceted challenge, not a single problem. This brief decodes the primary causes of AI-generated misinformation, from imperfections in training data to suboptimal context management, and outlines targeted, practical solutions for building more reliable applications.

Trust & Validation December 2025 View Abstract & Download →
Publication

Understanding Multi-Agent Systems

From "One" to "Many": this presentation explores the shift from single-agent LLMs to multi-agent systems. It details common design patterns, such as sequential systems, parallel systems, and independent validation, to solve complex HEOR workflows effectively.

Agentic & Multi-Agent Systems November 2025 View Abstract & Download →
White Paper

AI Agents in Health Research: What Does "Agent" Actually Mean?

The term 'AI agent' is often misused, leading to confusion and poor tool selection. This guide clarifies the difference between structured AI workflows (like a systematic review) and adaptive AI agents (like qualitative research), providing a practical framework for health researchers to choose the right tool for their specific task.

Agentic & Multi-Agent Systems October 2025 View Abstract & Download →
Publication

Upskilling for GenAI: A New Era in Research

GenAI is more than just a tool. It requires a new set of skills for researchers. This presentation outlines a practical upskilling journey, covering foundational principles, the balance between prompting and coding, and the importance of hands-on experimentation to truly harness the power of AI in a scientific context.

Skills & Training September 2025 View Abstract & Download →
Publication

Augmenting Health Economic Model Report Creation with LLMs

This practical presentation demonstrates a specific use case for augmenting the creation of health economic model reports. It showcases a concrete workflow where LLMs assist in drafting and reviewing report sections, integrating directly with existing processes to improve efficiency while maintaining essential human oversight.

Health Economic Modeling August 2025 View Abstract & Download →
White Paper

Why Enterprise-Grade AI Requires a Purpose-Built Strategy

The choice between off-the-shelf AI and a custom solution is a strategic decision between a consumer-grade product and an enterprise-grade system, not just a technical preference. This white paper decodes the critical risks of generalist tools, from their 'black box' logic to data integrity gaps, and outlines a framework for building purpose-built AI that delivers true control, integration, and auditability for specialized, high-stakes work.

AI Strategy & Architecture July 2025 View Abstract & Download →
White Paper

A Deep Dive into Advanced Retrieval-Augmented Generation (RAG)

Retrieval-Augmented Generation (RAG) is more than a simple pipeline. It is a sophisticated workflow that requires a deliberate architectural approach to succeed at scale. This white paper provides a blueprint for moving beyond naive prototypes to build advanced, enterprise-grade RAG systems, aimed at building AI that is accurate, efficient, and truly reliable.

AI Strategy & Architecture June 2025 View Abstract & Download →
White Paper

Beyond Hype and Fear

The debate over Large Language Models isn't about hype versus fear. It is about a fundamental misunderstanding that leads organizations to treat them as standalone 'magic boxes,' resulting in staggering project failure rates. Using a powerful Formula 1 analogy, this white paper reframes the LLM as a high-performance engine and provides the systems-based framework for building the complete race car, one that delivers reliable workflow enhancement, strategic agility, and a durable competitive advantage.

AI Strategy & Architecture May 2025 View Abstract & Download →
White Paper

LLM-as-a-Judge

Many organizations face what can be called the 'AI Trust Paradox,' where the potential efficiency of AI is often lost to the time-consuming manual verification needed to meet scientific standards. This whitepaper explores a framework called the 'LLM-as-a-Judge' model as one way to address this challenge.

Trust & Validation April 2025 View Abstract & Download →
White Paper

Context Engineering: Why GenAI Projects Fail (or Work)?

This whitepaper argues that the 95% failure rate of enterprise GenAI projects is not due to model limitations, but rather a lack of 'Context Engineering,' a systematic architectural approach that prioritizes data governance, reliable retrieval, and structured memory over simple prompt crafting.

AI Strategy & Architecture March 2025 View Abstract & Download →