Kai Ole Hartwig

name of the term: Prompt Injection
descriptions of the term:

Definition

Definition. In a prompt injection, an attacker smuggles instructions into text that a language model processes. The model does not cleanly separate commands from data. It then follows the injected text instead of its original task. The indirect variant hides the instruction in web pages, emails or documents that an agent reads.

Why it matters. Once an agent uses tools, for example via MCP on a CMS or a cluster, a wrong answer becomes a wrong action. No filter alone protects against this. What works are narrow permissions, approvals for write actions and workspaces instead of direct publishing.

Example. An agent summarises comments in the TYPO3 backend. One comment contains a hidden instruction to publish all pages. Without an approval step, the agent carries it out.

Related. MCP, RAG, Agent Readiness

Synonyms: Prompt injection attack, Indirect Prompt Injection
Type of term: definition
Language of the term (2 char ISO code): en
Back