This post examines the security risks of relying on one AI model to police another. Known as model-as-a-judge, this approach adds a second model to review user requests, documents, or proposed actions before an AI application proceeds. The judge acts as a security checkpoint, approving what it considers safe and blocking what it considers malicious.

But attackers can target that checkpoint too. The content being inspected can contain instructions or misleading information designed to manipulate the judge’s decision. We look at research demonstrating this problem, including Check Point’s findings on Jev, and examine the tradeoffs in detection accuracy, cost, and response time. AI can support security, and protecting your business requires safeguards beyond another model’s approval.

The attacker can target the judge

In JudgeDeceiver, published at ACM CCS 2024, researchers examined a model tasked with picking the best answer from several candidates. They automatically crafted text to insert into one candidate, steering the judge toward that answer even when it was wrong. The attack worked with control over a single candidate.

Two copies of the same LLM-as-a-judge prompt. Without the attack the judge picks Output (a); with an optimized sequence appended to Output (b), it picks Output (b).
Left: the judge picks the correct answer, Output (a). Right: the same prompt with an optimized sequence appended to Output (b), and the judge now picks Output (b). Source: Shi et al., JudgeDeceiver, ACM CCS 2024, Figure 1.

The study extended this to search results, feedback used to train models, and tool selection. In the tool-selection experiment, an altered tool description persuaded the judge to choose an irrelevant tool over one suited to the task. The material under review became the attack against the reviewer.

The Attacker Moves Second, published at USENIX Security 2026, tested 12 defenses against jailbreaks and prompt injections. Researchers adapted their attacks to each defense and refined them over repeated attempts. Most defenses had originally reported near-zero attack success. Under these stronger evaluations, attacks exceeded 90% success against most.

Its model-filtering experiments tested the detector and protected model together. Success required both passing the filter and making the application follow the attacker’s objective. Researchers bypassed both in these experiments.

A successful injected text that reads like an ordinary company policy note asking to delete a tracking file.
A trigger found by the search attack that passed the detector. It reads like an ordinary policy note, not like an injection. Source: Nasr et al., The Attacker Moves Second, arXiv:2510.09023, Section 5.3 (CC BY 4.0).

These results show that transformer-based models remain susceptible to manipulation when used as security judges, with the degree of vulnerability varying across models.

A judge’s value as an independent defense depends on its own resistance to manipulation.

Smaller, larger, or local: you still pay

Self-hosting a smaller judge moves inference into your infrastructure budget. You still need memory, compute capacity, and enough headroom for peak traffic. Existing hardware may absorb the workload.

A March 2026 GovTech preprint reported 40 ms latency for two local encoder-based detectors on an Apple M2, compared with approximately 1.44 to 5.88 seconds for its LLM judges. The judges achieved better detection scores on that dataset. The study compared deployed configurations running on different hardware.

Table comparing latency, precision, recall and F1 of prompt-attack detectors: encoder models at 0.04 seconds with low F1, LLM judges at 1.44 to 5.88 seconds with F1 above 0.81.
Latency is in seconds. The encoder-based detectors answered in 0.04 s with F1 scores of 0.23 and 0.35; the LLM judges took 1.44 to 5.88 s and scored 0.81 to 0.87. Rows marked 1 were timed on a local Apple M2. Source: Le, Goh and Tang, GovTech Singapore, arXiv:2603.25176, Table 2 (CC BY 4.0).

That is the tradeoff to measure: missed attacks, blocked legitimate requests, and time spent waiting. Model size is only one factor in that assessment.

Use a larger model for a blocking check and allowed requests incur another inference cost and another wait. When the judge and main call have comparable costs and durations and run sequentially, both totals can approach 2x. Short verdicts and overlapping execution can reduce the overhead, and resistance to attack still requires separate testing.

Cheaper and still unreliable

Specialized classifiers give applications a quick security verdict: approve a request, flag a suspicious document, or block a detected attack. Models such as Meta’s Prompt Guard and Protect AI’s prompt-injection detectors use this approach. Their decisions become another target for attackers trying to get malicious content through.

In Bypassing LLM Guardrails, published at LLMSEC 2025, researchers tested six detection systems, including Meta’s original Prompt Guard and Protect AI’s DeBERTa models. They modified characters and wording in malicious prompts, causing detectors to label attacks as benign. Results varied across models and attack techniques, demonstrating the importance of testing each deployed configuration against deliberate evasion.

Table of twelve character injection techniques, each applied to the word Hello, such as H3110, hèllö and olleH.
The character injection techniques tested, each applied to the word “Hello”. Source: Hackett et al., Bypassing LLM Guardrails, LLMSEC 2025, Table 1 (CC BY 4.0).

Meta’s documentation for Prompt Guard 2 identifies adaptive attacks as a limitation. It also explains that detection performance depends on the application’s mix of legitimate and malicious inputs. A dedicated security model still requires testing against the conditions it will encounter in production.

The Limitations section of the Prompt Guard 2 model card, listing vulnerability to adaptive attacks and application-specific prompts.
The Limitations section of Meta’s own documentation. Source: Llama Prompt Guard 2 22M model card, Hugging Face.

Check Point demonstrated a related failure in Jev, a model that returns structured decisions. An attacker controlled part of an investment report and manipulated the model into recommending a high-risk investment. The strongest attacker succeeded in 25 of 27 runs, at approximately $0.50 per successful break. The model produced a correctly formatted decision based on attacker-supplied misinformation.

Table of successful breaks and attacker cost per break for three attackers at three difficulty levels; the strongest attacker broke 25 of 27 runs at 0.50 dollars per break.
Successful breaks out of 9 runs per difficulty level, and the attacker’s API cost per successful break. Source: Check Point, A First Look at Jev.

These findings put the focus on the decision itself. Whether the output is a label, a probability, or an approval, its reliability under attack determines the protection it provides. Business actions authorized by model judgments need safeguards that remain effective when those judgments are wrong.

What protects you when the judge is wrong?

Security decisions require safeguards beyond a transformer-based model’s interpretation of hostile content.

A useful detector can catch attacks. A security architecture still needs an answer for the attacks it misses, whatever the judge’s size or price.

At Milgram, we approach LLM security differently.

Talk to Milgram about protecting your AI workflows.