World Wires · story 28624 · corroborated · 1 source(s)

PHRBench: Evaluating Post-Hallucination Reasoning in LLMs

The authors introduce PHRBench to evaluate how large language models resolve hallucinated premises at the response level.

Open in the desk

Coverage

What this site indexes