421M fine-tuned model Gyra v0.2 catches destructive coding-agent commands, ships as Claude Code hook
CelebrationAble6568 · reddit · 2026-09-27
The author fine-tuned the Laya English checkpoint into Gyra v0.2, a 421M-parameter model that reviews coding-agent tool calls for destructive commands, hangs, and secret exposure, and judges tool output for masked failures or prompt injection.
On a frozen 365-case audit at hook thresholds: destructive commands caught 18/24 (vs 15/24 for base Laya), hangs 29/30 (vs 0/30), secret exposure 7/12 (still weak), and benign scripts left alone 37/39 (vs 18/39 — far fewer false kills). A separate deterministic rule layer caught 17/17 direct file and secret cases with no false alarms. It ships as a Claude Code / Codex hook combining model judgment with fixed rules. Code, weights, thresholds, and audit scripts are open-sourced, though the author notes the demo videos are scripted and the eval folder still contains v0.1 artifacts.
More from coding & agent
- Claude Opus 5.5 Has 5 Effort Levels — Here's When to Use Each — CodeByPoonam · 2026-09-27
- Supermemory open-sources its Slack-based company brain AI teammate with one-click self-host — _AustinCalvert_ · 2026-09-27
- Dev hands an Opus 5.5 his released web game to improve and cut a trailer — hullabaloo22 · 2026-09-27
- Grok Build v1.0.41 ships subagent config inheritance and long-reasoning fixes — mark_k · 2026-09-27
- Harrison Chase on building frontier harnesses: context layers, evals, online learning — VeryWellVersed · 2026-09-27
- Dev shares public Grok Bot 'chief of staff' template running a background specialist agent swarm — RachelVT42 · 2026-09-27