Replicating ExploitBench Would Cost ~$59.3M in API Fees, Security Researcher Estimates
OwariDa · x · 2026-09-22
Security researcher Jhaddix points out that offensive/defensive cybersec people and threat actors can't match frontier labs. Napkin math from Hugging Face reports: ExploitBench's eval (10,000 concurrent agents, 8 days non-stop, no limiters) would cost $59.3M in API fees (Astra), or $3.4M for DeepSeek v4, $16.9M for Kimi K3, $9.3M for GLM 5.3 on cloud hardware — plus engineering a harness for 10,000 agents. Frontier labs also removed all guardrails and classifiers during testing, so outsiders can't replicate the setup.
More from Models
- Azure "oopsie" reportedly leaks GPT-6-Luna and GPT-6-Sol names — scaling01 · 2026-09-22
- GPT-6-Luna and GPT-6-Sol rumored imminent, names possibly leaked via Azure — scaling01 · 2026-09-22
- Qwen 3.8 27b fine-tune cuts verbose output by up to 40% with little quality loss — julianharris · 2026-09-22
- METR publishes independent investigation of OpenAI agents' multi-day Hugging Face hack — JeffLadish · 2026-09-22
- JevBench v1.3.0 Released: Original Jev Still Leads at 74.4, 47 Rivals Closing In — airesearch12 · 2026-09-22
- DeepSeek vs Jev: How an LLM Stacks Up on a System-One Probability Benchmark — frappuccinoCoin · 2026-09-22