LLMs Need More RL Training for Vulnerability Mining
teortaxesTex · x · 2026-07-20
Discussions point out that randomly sampled CVE (security vulnerability) tests have become too easy for current models to differentiate their capabilities, prompting the community to build more challenging, curated datasets.
Commenter teortaxesTex noted that models like GLM 5.2 and Kimi K3 are currently far from their capability ceilings due to insufficient reinforcement learning (RL) training. In comparison, models like GPT, Grok, and Opus perform more stably because of more comprehensive RL. Domestic models still need to increase their training steps to improve consistency.
More from Models
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Macaron V1 adds LoRA RL on GLM 5.2 and claims SOTA benchmark gains — Xianbao_QIAN · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22