2 campaigns · target: google/gemma-3-27b-it · 120 simulations · 2 policy categories · generated 07 August 2026, 19:15 UTC
Every number below is the total across these runs. Attempts are grouped by policy category, so runs that covered the same policy share a row.
45cdcb07-a4c7-48f2-86fe-82a5931dee3e45cdcb07-a4c7-48f2-86fe-82a5931dee3eOne complete multi-turn conversation between our attack simulator and your model, aimed at a single harmful goal. Every number on this page counts attempts.
The share of attempts that ended in a policy violation, counting critical and moderate together. Higher is worse.
Average ASR is measured across every attempt in a category.
Best single strategy is the highest rate one attack approach reached on its own — a model can look safe on average and still be reliably breakable by one method.
Each bar is one policy category, filled by what share of its attempts ended in a violation. A shorter bar is a safer model.
Every attempt, worst first. Pick one to read the full conversation — turns the judge flagged are outlined in red, and where it quoted specific evidence, that text is highlighted inside the turn.
This report shows outcomes only — scores, goals, categories, and where things went wrong. The attack strategies themselves, how they evolve, and how we run them at scale aren't included here. Full transcripts, custom policies and remediation help are what we do next.