Back to Blog

Explosive Prompts: CERT-AGID measures the prompt injection that waits for the right moment

·5 min read
Tomato Blue illustration on the theme of a delayed trigger: an apparently inert element hiding a mechanism, ready to fire when the trigger arrives.

Imagine an AI agent at your company reads an email, a PDF or a web page with an instruction hidden inside. The model does nothing unusual, and that is exactly why nobody notices. CERT-AGID, Italy's national CERT for the public administration, has just published a paper that measures how often that silence hides an instruction already absorbed and ready to fire later, on any trigger. The most useful result for governance teams, though, is not about the models: it is about the architecture around them.

What an Explosive Prompt is

A prompt injection is a malicious instruction planted in content the agent reads, in the hope that the model mistakes it for a legitimate order. In its classic form the effect is immediate. An Explosive Prompt is the delayed-effect variant: the instruction stays inactive in the context and takes effect only when a later trigger occurs.

The term does not originate at CERT-AGID. The paper takes it from the study by Szczepaniak, Feldman, Viner and Nassi, Defusing Explosive Prompts (arXiv:2609.22510), which reports that, across nine production agents, conditional injections performed markedly better than imperative ones. That figure comes from the cited study, not from a CERT-AGID experiment.

The practical consequence is uncomfortable: the absence of anomalous behavior while the document is being read is not enough to rule out the risk. A control that only looks at what happens right away misses the part that matters.

CERT-AGID's experiment

CERT-AGID built a minimal environment: a custom Python orchestrator, with no agentic framework. It tested four open-weight models (GLM-5.2, DeepSeek V4 Flash, Qwen 3.8 27B and Gemma 4 26B-A4B-IT) with identical inference parameters. The conversation lasts three turns and the trigger is a harmless phrase, "Perfetto, grazie." ("Perfect, thanks."). Each model was run 30 independent times.

Before the trigger, the tool was not called in any run. After the trigger, activation rates differ widely from model to model:

  • Gemma 4: 76.7%
  • GLM-5.2: 53.3%
  • DeepSeek V4 Flash: 46.7%
  • Qwen 3.8: 10%

One caveat, which is the paper's own: these values describe that single experiment, with 30 runs per condition, and are not a general measure of model safety. Using them to build a ranking would be a mistake.

The comparison that matters: manipulated model, compromised system?

The second experiment is the most interesting one for anyone deciding how to build an agent. The tool in play, set_protected_flag(value="TRUE"), changes a protected state, while the agent operates as viewer_agent, that is, without the privilege to do so. The same model (GLM-5.2) is tested in two authorization modes, 30 runs each:

ModeActivations after triggerTools executedState changes
Permissive (the system trusts the model)141414
Enforced (independent external control)1000

The model is manipulated in both cases: in enforced mode it tries to call the tool 10 times out of 30. But the action is not executed and the protected state never changes.

Mind how the table is read. The two series, 14 versus 10, are independent, stochastic batches: the paper warns that the difference should not be read as an effect of the policy on how likely the model is to "try". The policy acts on execution, not on the model's behavior. It is a detail that press summaries tend to lose.

Why not trust the model alone

Hence the paper's conclusion: a successful prompt injection against the model does not necessarily have to become a successful action against the system. Authorization decisions should not be left to the model alone. You need independent external controls, sitting outside the model, at system level.

For anyone integrating agents through the Model Context Protocol, or any other tool-calling mechanism, the principle is concrete: the model may ask for an action to be executed, but whether the action happens must be decided by a component that does not read the manipulated content. The paper points in this regard to work on zero-trust authorization for enterprise MCP (Li, Wang, Manoharan, arXiv:2609.22573).

What it means for compliance and governance (Tomato editorial reading)

This section is our own interpretation. The CERT-AGID paper is a technical proof of concept: it cites no regulation or standard, and nothing below should be attributed to AgID.

In our view, the principle "authorization sits outside the model" lends itself to being read alongside the topics a compliance officer already has on the table: risk management of artificial intelligence systems, access and authorization controls for information systems, operational resilience. These are the areas covered by the AI Act, the ISO/IEC 42001 and 27001 standards and, for those subject to them, NIS2 and DORA. The paper gives an experimental example of why an external technical control is worth more than a recommendation entrusted to the model.

It is a generic link, and we want it to stay generic. If you need specific requirements, articles or controls from any of these regulations, verify them against the official source before writing them into a policy.

What to do now

Four checks you can run this week on your agents:

  1. List the tools with real-world effects: writing data, sending communications, payments. They are the only places where an injection does damage.
  2. Move authorization outside the model: every tool call goes through a control that applies the caller's permissions, not those of the content that was read.
  3. Do not trust quiet early turns: content with no visible effect right away is not harmless for that reason.
  4. If you cite the percentages, cite the limit too: they are frequencies from a proof of concept, not a measure of model safety.

If you want to understand where your exposure points are, mapping the tools, permissions and external effects of your agents in production is the first step, and it costs far less before an incident than after one: Assess the risks of your AI agents and models in production. On the three conditions that make an agent dangerous, see also The lethal trifecta: the fire triangle of AI agents.

Sources