Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Oct 8, 2026 - Python
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 61 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Agent Memory Integrity: a conformance test suite that measures whether AI agent memory and checkpoint stores notice tampering. Sixteen measured rows across LangGraph (SQLite, Postgres, Redis), OpenAI Agents SDK, LlamaIndex, Letta, Mem0 and six integrity tools. IETF draft-khandelwal-bmwg-agent-memory-integrity.
Local-first workbench to run, inspect, compare, report, and gate OpenAI Codex Security scans.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Deterministic security benchmark for tool-using AI agents
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
21 passive security evaluation inputs: conventional appsec and exploratory hardware, firmware, PLC, RTOS, database and policy domains. No scans started.
Security evaluation input | Database-engine source, SQL parser, storage, transaction and extension boundaries | production-control-no-vulnerability-claim
Production-grade microservices security benchmark featuring OWASP Top 10 logic exploits, automated remediation, custom Semgrep SAST rules, and CI/CD DevSecOps gates.
Internal PyPI SCA precision and recall benchmark corpus
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
Reproducible intentionally vulnerable JavaScript/TypeScript SAST accuracy benchmark
Security evaluation input | Cisco ACLs, zone firewall policy, iptables, topology, routing and rule changes | labeled-examples-in-library
An AI assistant that can't be hijacked by the documents it reads. The part that decides what to do never sees untrusted text, and every value is tagged with its origin, checked at each action. In a 1440-run offline benchmark: a normal agent was hijacked 73/73, this one 0/73.
Security evaluation input | Labeled Python injection, deserialization, crypto and trust-boundary cases | labeled-positive-negative
Security evaluation input | GraphQL authorization, queries, mutations, subscriptions, injection and query complexity and resource-exhaustion scenarios | intentional-vulnerabilities
Commit-pinned cal.diy source corpus; Vybscan ground-truth oracle available in benchmark-results
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."