This post examines the security risks of relying on one AI model to police another. Known as model-as-a-judge, this approach adds a second model to review user requests, documents, or proposed actions before an AI application proceeds. The judge acts as a security checkpoint, approving what it considers safe and blocking what it considers malicious.
But attackers can target that checkpoint too. The content being inspected can contain instructions or misleading information designed to manipulate the judge’s decision. We look at research demonstrating this problem, including Check Point’s findings on Jev, and examine the tradeoffs in detection accuracy, cost, and response time. AI can support security, and protecting your business requires safeguards beyond another model’s approval.
The attacker can target the judge
In JudgeDeceiver, published at ACM CCS 2024, researchers examined a model tasked with picking the best answer from several candidates. They automatically crafted text to insert into one candidate, steering the judge toward that answer even when it was wrong. The attack worked with control over a single candidate.

The study extended this to search results, feedback used to train models, and tool selection. In the tool-selection experiment, an altered tool description persuaded the judge to choose an irrelevant tool over one suited to the task. The material under review became the attack against the reviewer.
The Attacker Moves Second, published at USENIX Security 2026, tested 12 defenses against jailbreaks and prompt injections. Researchers adapted their attacks to each defense and refined them over repeated attempts. Most defenses had originally reported near-zero attack success. Under these stronger evaluations, attacks exceeded 90% success against most.
Its model-filtering experiments tested the detector and protected model together. Success required both passing the filter and making the application follow the attacker’s objective. Researchers bypassed both in these experiments.

These results show that transformer-based models remain susceptible to manipulation when used as security judges, with the degree of vulnerability varying across models.
A judge’s value as an independent defense depends on its own resistance to manipulation.
Smaller, larger, or local: you still pay
Self-hosting a smaller judge moves inference into your infrastructure budget. You still need memory, compute capacity, and enough headroom for peak traffic. Existing hardware may absorb the workload.
A March 2026 GovTech preprint reported 40 ms latency for two local encoder-based detectors on an Apple M2, compared with approximately 1.44 to 5.88 seconds for its LLM judges. The judges achieved better detection scores on that dataset. The study compared deployed configurations running on different hardware.

That is the tradeoff to measure: missed attacks, blocked legitimate requests, and time spent waiting. Model size is only one factor in that assessment.
Use a larger model for a blocking check and allowed requests incur another inference cost and another wait. When the judge and main call have comparable costs and durations and run sequentially, both totals can approach 2x. Short verdicts and overlapping execution can reduce the overhead, and resistance to attack still requires separate testing.
Cheaper and still unreliable
Specialized classifiers give applications a quick security verdict: approve a request, flag a suspicious document, or block a detected attack. Models such as Meta’s Prompt Guard and Protect AI’s prompt-injection detectors use this approach. Their decisions become another target for attackers trying to get malicious content through.
In Bypassing LLM Guardrails, published at LLMSEC 2025, researchers tested six detection systems, including Meta’s original Prompt Guard and Protect AI’s DeBERTa models. They modified characters and wording in malicious prompts, causing detectors to label attacks as benign. Results varied across models and attack techniques, demonstrating the importance of testing each deployed configuration against deliberate evasion.

Meta’s documentation for Prompt Guard 2 identifies adaptive attacks as a limitation. It also explains that detection performance depends on the application’s mix of legitimate and malicious inputs. A dedicated security model still requires testing against the conditions it will encounter in production.

Check Point demonstrated a related failure in Jev, a model that returns structured decisions. An attacker controlled part of an investment report and manipulated the model into recommending a high-risk investment. The strongest attacker succeeded in 25 of 27 runs, at approximately $0.50 per successful break. The model produced a correctly formatted decision based on attacker-supplied misinformation.

These findings put the focus on the decision itself. Whether the output is a label, a probability, or an approval, its reliability under attack determines the protection it provides. Business actions authorized by model judgments need safeguards that remain effective when those judgments are wrong.
What protects you when the judge is wrong?
Security decisions require safeguards beyond a transformer-based model’s interpretation of hostile content.
A useful detector can catch attacks. A security architecture still needs an answer for the attacks it misses, whatever the judge’s size or price.
At Milgram, we approach LLM security differently.
Talk to Milgram about protecting your AI workflows.
