Prompt injection attacks usually aim to force a model to output harmful or attacker-chosen content. The ENDOPROMPT approach takes a different route: it degrades the model's performance on benign tasks without any need for harmful output. The authors argue that many existing attack objectives depend on task labels or predefined target responses, which can be a limitation.
ENDOPROMPT Shows Prompt Injection Can Silently Reduce Task Utility
A new white-box method uses victim-side pseudo-references to degrade benign task performance without triggering harmful content.
More in Research Digest
CaliPPer: Quantifying and predicting AI binding prediction performance
A new method called CaliPPer helps researchers anticipate and improve how well AI models perform on new binding prediction datasets.
Early Data on AI's Impact in Science: Excitement and Concerns
A new arXiv paper offers some of the first empirical insights into how AI is affecting scientific research, addressing both promise and worry.
Certifying When an Agent Knows Enough to Act
New framework lets autonomous agents decide which hidden states matter, how many probes are needed, and when to abstain.
SciWalker Synthesizes Scientific Coding Problems for LLM Training
A new method uses operator graphs and execution feedback to generate realistic scientific coding tasks, addressing the scarcity of high-quality training data for large language models.