Detecting LLM Sabotage and Unverbalized Deception via Probes
Researchers developed a white-box probe architecture and deception dataset to detect hidden sabotage and unverbalized deception in frontier LLM agents.
Researchers developed a white-box probe architecture and deception dataset to detect hidden sabotage and unverbalized deception in frontier LLM agents.