← Back to brief
ResearchOfficialPreprintarXiv Cryptography and Security

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

A new preprint introduces CPInj, a collaborative prompt injection attack that targets Textual Collaborative Prompt Optimization (TCPO), a decentralized method for improving LLM prompts. The attack injects malicious instructions into local prompts, which persist through aggregation and degrade downstream task performance, while evading current defenses. The authors also propose a defense method, APAgg, which partially mitigates the attack, but CPInj remains highly effective, exposing a significant vulnerability in TCPO systems.

Why it matters: This work exposes a critical and previously unexplored vulnerability in decentralized prompt optimization for LLMs, emphasizing the urgent need for more robust defenses against collaborative prompt injection attacks.

Full story at: arXiv Cryptography and Security