Artificial intelligence models can give users the wrong answer and do so with great confidence. They can also hedge and warn that they are unsure—even when they get the answer right.

A study led by UC Riverside computer scientists helps explain why. Researchers found that confidence and correctness can arise from different internal features within large language models, challenging the assumption that a model's confidence reliably indicates whether its answer is accurate.

Their discovery, published on the arXiv preprint server, could help build more reliable AI models. As large language models are increasingly used to inform decisions and complete tasks, developers need better ways to determine when their answers can be trusted.

To read more, click here.