MetaLint: Qwen3-4B lifts code lint detection F-score 2.7x to 70.4%, matching o3-mini
dan_fried · x · 2026-09-29
A CMU team will present MetaLint at #COLM2026: a meta-learning framework framing code linting as instruction following, where models check code against natural-language best-practice specs—enabling test-time control and generalization to unseen rules without retraining. Trained only on synthetic data from automatic linters, Qwen3-4B gains 2.7x detection F-score (25.9% → 70.4%) with the highest recall and 26.7% localization F-score on a human-curated PEP-inspired hard benchmark, matching larger models like o3-mini. The author also presents work on evaluating slop in long-horizon coding agents.
More from coding & agent
- AlignOPSD fixes decision-timestamp mismatch in agent distillation, beating GRPO by 5.5-8.7% — Mingju Chen · 2026-09-29
- CompoWorld scales general agent training by composing environments from reusable services — AllSpark-Research · 2026-09-29
- Adaptive Consistency Graph lifts long-horizon agent success from 44.5% to 50.2% — Beihang · 2026-09-29
- ByteDance's TraceDance auto-builds agent behavior benchmarks from real deployment traces — ByteDance · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29
- ADHD is a superpower for juggling 10 AI agents, dinner, and shitposting at once — enggirlfriend · 2026-09-29