swish-bench · Open AI Agent Security Benchmark

AI Agent Security Leaderboard

Empirical security evaluations comparing popular AI agent frameworks against OWASP LLM Top 10 and ASI01–10 agentic threat vectors.

5
Frameworks Scored
1,250
Payloads Tested
100%
Highest Pass Rate
76/100
Avg Industry Score
Loading benchmark evaluation data...
Benchmark Methodology & Reproducibility
Evaluations are conducted using agentic-redteam v1.1.0 over 250 standardized payloads per framework target. Scores are computed using weighted severity penalties (CRITICAL × 4, HIGH × 3, MEDIUM × 2). To run these evaluations locally or submit your agent framework for indexing, inspect our open repository at github.com/Muneeb7860/agentic-redteam.