LLMs Need More RL Training for Vulnerability Mining

teortaxesTex · x · 2026-07-20

Discussions point out that randomly sampled CVE (security vulnerability) tests have become too easy for current models to differentiate their capabilities, prompting the community to build more challenging, curated datasets.

Commenter teortaxesTex noted that models like GLM 5.2 and Kimi K3 are currently far from their capability ceilings due to insufficient reinforcement learning (RL) training. In comparison, models like GPT, Grok, and Opus perform more stably because of more comprehensive RL. Domestic models still need to increase their training steps to improve consistency.

Original post →

More from Models

Models channel →