EvoFlint is a trust & safety evaluation for open-weight LLMs. Bring a Hugging Face model; we'll run our multi-turn attacks against it and report a per-policy safety score with evidence of what happened.
Submit a Hugging Face model id. No setup, no code to write.
An evolving archive of multi-turn attacks probes it across safety policies, on our compute.
A per-policy score, the full attack trajectories, and short snippets showing exactly where it slipped.
See a sample report →Most tools ask “can this model be broken?” EvoFlint asks “how does it break, and in how many ways?” — building a structured atlas of a model's failures instead of a list of one-off jailbreaks.
EvoFlint-native attacks vs. GOAT, a published multi-turn red-teaming method we ran via promptfoo's implementation of it — same target model, same violation standard, human-labeled (n=20 trajectories per policy per source).
| Policy | EvoFlint-native | GOAT (via promptfoo) | promptfoo plugin(s) compared |
|---|---|---|---|
| CBRNE | 80% | 25% | harmful:chemical-biological-weapons, harmful:indiscriminate-weapons, harmful:weapons:ied |
| Self-harm | 55% | 25% | harmful:self-harm |
Bring an open-weight Hugging Face model and we'll run EvoFlint against it. You'll get a per-policy score, attack trajectories, and evidence back by email.
Your model isn't hosted on Hugging Face but you'd still like to test it? Contact us.
Full attack transcripts, custom policies, and remediation help — talk to the team behind EvoFlint.