OpenAegis doubles security-agent pass rate to 58.1% via CyberFactory trajectories
jiqizhixin · x · 2026-09-10
Beihang University, ELLIS, Singapore Management University, and IQuest Research present CyberFactory and OpenAegis. Key insight: real vulnerability data is not training data. CyberFactory rebuilds scattered public vulnerability materials into runnable, judgeable tasks by reconstructing pre/post-fix environments from repos and fix commits; a security-analysis Skill guides an agent through each task, producing operation trajectories that train OpenAegis instead of static QA pairs. Result: pass rate nearly doubles from 29.6% (Qwen3.5 base) to 58.1%, leading GLM 5.2 by 14.8 points and Kimi K2.7 by 6.4 points under a 1-hour limit with fewer parameters.
More from coding & agent
- How to run a hackathon entirely with ChatGPT Sites, all from your phone — gabrielchua · 2026-09-10
- Engineer stops reading AI-written code, shares 4 ways to keep judgment sharp — every · 2026-09-10
- Memory Doesn't Equal Expertise: Builders Ask Where Domain-Specific Agents Get Their Knowledge — sartomiki · 2026-09-10
- New Benchmark Idea: One Agentic Pass, 8 Hours, Model Makes Its Own Launch Video — Lowkey_LokiSN · 2026-09-10
- Google Simplifies Gemini API Skills, Boosting Correct API Code Generation to 87% — patloeber · 2026-09-10
- Do Persistent AI 'Teammates' Actually Work? Users Question Grok Bot's Real-World Value — No-Advantage8770 · 2026-09-10