PHRBench: Evaluating Post-Hallucination Reasoning in LLMs
The authors introduce PHRBench to evaluate how large language models resolve hallucinated premises at the response level.
The authors introduce PHRBench to evaluate how large language models resolve hallucinated premises at the response level.