← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models

Researchers formalize the concept of payload recoverability in LLM steganography and introduce new low-recoverability schemes using embedding-space mappings. Their experiments show that while these schemes reduce the ease of secret recovery, mechanistic interpretability methods—specifically linear probes on later-layer activations—can still detect steganographic content with up to 33% higher accuracy in fine-tuned models compared to base models.

Why it matters: This work highlights both the evolving threat of covert communication via LLMs and the potential for interpretability-based defenses, signaling a new dynamic in model security.

Full story at: arXiv Cryptography and Security