From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
Researchers propose a unified taxonomy for large language model (LLM) misalignment, structured along three dimensions: degree of goal-directedness, object of deception, and mechanism. By applying this taxonomy to 50 existing benchmarks, they find that fabrication is well-represented, while pragmatic distortion, attribution, and capability self-knowledge are underrepresented, and strategic deception benchmarks are still emerging. The paper also offers recommendations for developers and regulators, including a reporting template for future work.
Why it matters: A unified taxonomy can help standardize research on LLM misalignment and highlight gaps in current evaluation methods, informing both development and regulation.
Full story at: arXiv Computers and Society ↗