Summary
An evaluation harness revealed that AI models, particularly large language models (LLMs), are often most confident when providing incorrect information, a flaw missed by qualitative reviews. This gap between 'sounds right' and 'is correct' poses significant risks for LLM-assisted enterprise tools that influence critical business decisions.