Hugging Face researcher proposes labs release small models, alignment recipes to rebuild trust
lvwerra · x · 2026-09-15
lvwerra of Hugging Face argues the AI safety and pacing debate is again dominated by a handful of people, and lays out a constructive alternative:
- Release small open variants of frontier models: lets the community red-team without API costs or ban risks; the small model acts as a canary for the large one, with labs documenting differences. He notes researcher churn means most secrets are already known in labs anyway.
- Share core parts of the post-training/alignment recipe: training stages, objectives, data descriptions, and key ablations at minimum.
- Publish tech reports beyond evals: reproducible evidence for scary internal findings, verified by an independent team, with sensitive details disclosed privately first.
He argues five years of "if you knew what we've seen" messaging has blurred the line between marketing and genuine concern, and this proposal would help rebuild trust.
More from AGI Musings
- Critics Slam AI Firms for Hyping Danger While Disclosing Zero Technical Detail — 1a3orn · 2026-09-16
- Virologist vs. AI safety advocate clash over whether AI biorisk is real or a regulation dodge — anshulkundaje · 2026-09-16
- Max Levchin: Publicly discussing agent takeover strategies feeds the training data — JosephJacks_ · 2026-09-16
- Nobody Got Fired for AI Skepticism in 2023 — Does That Still Hold in 2026? — dfinke · 2026-09-15
- Gary Marcus Decodes Altman: AI 'Pacing' Means Going as Fast as You Can Without Getting Sued — Gary Marcus · 2026-09-15
- Tsinghua-ByteDance 75-page paper maps why recursive AI self-improvement still stalls — alex_verem · 2026-09-15