The field agrees that accuracy metrics cannot tell you whether an AI research tool is trustworthy. This paper proposes a confidence layer built on four properties that make reliability visible and verifiable, embedded during construction rather than bolted on.