EvoFlint is a trust & safety evaluation for open-weight LLMs. Bring a Hugging Face model; we'll run our multi-turn attacks against it and report a per-policy safety score with evidence of what happened.
Submit a Hugging Face model id. No setup, no code to write.
An evolving archive of multi-turn attacks probes it across safety policies, on our compute.
A per-policy score, the full attack trajectories, and short snippets showing exactly where it slipped.
See a sample report →Most tools ask “can this model be broken?” EvoFlint asks “how does it break, and in how many ways?” — building a structured atlas of a model's failures instead of a list of one-off jailbreaks.
EvoFlint-native attacks vs. GOAT, a published multi-turn red-teaming method we ran via promptfoo's implementation of it — same target model, same violation standard, human-labeled (n=20 trajectories per policy per source).
| Policy | EvoFlint-native | GOAT (via promptfoo) | promptfoo plugin(s) compared |
|---|---|---|---|
| CBRNE | 80% | 25% | harmful:chemical-biological-weapons, harmful:indiscriminate-weapons, harmful:weapons:ied |
| Self-harm | 55% | 25% | harmful:self-harm |
Bring an open-weight Hugging Face model and we'll run EvoFlint against it. You'll get a per-policy score, attack trajectories, and evidence back by email.
Have another model that isn't hosted on Hugging Face? We can evaluate those too — contact us to set it up.
Full attack transcripts, custom policies, and remediation help — talk to the team behind EvoFlint.