โ Fleet-evaluation harness for security-agent benchmarks ๐๐๐ฏ๐ฒ๐ฟ๐๐๐บ-๐๐ฎ๐, ๐๐ ๐ฝ๐น๐ผ๐ถ๐๐๐๐บ, ๐๐ ๐ฝ๐น๐ผ๐ถ๐๐๐ฒ๐ป๐ฐ๐ต ๐ฉ๐ด, ๐๐ฉ๐-๐ฏ๐ฒ๐ป๐ฐ๐ต, run in real disposable Daytona sandboxes per trial โ with Modal-based KVM-capable sandboxes planned for kernel-exploit tasks. Measures agent capability and infra reliability, not simulated.
docker binaries modal fuzzing patch cve qemu-kvm net-snmp futex arvo otel msan libdwarf ubsan cybergym audit-loop libxml2-uaf exploit-gym asan-global-buffer-overflow aslr-entropy
-
Updated
Oct 3, 2026 - Python