The widespread adoption of AI tools in HEOR has created a critical problem, because common evaluation metrics were designed for simpler tasks. This paper explores where traditional metrics fall short in complex, multi-step AI research workflows and what should complement them.