Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights
le_james94 · x · 2026-09-18
A thread of model-behavior observations: fine-tuning Llama 3.1 8B into narrower personas doesn't change refusal on steroid-buying questions, but the sarcastic persona mocks your intent. Separately, Opus 4.6 scores 0% on an ARC-AGI-3 environment without a harness and 97.1% with a hand-crafted one—reasoning-eval noise floors of 0.22–1.48 points mean 1-point press-release gaps are noise.
More from Models
- Leaked GPT-6 Astra tops FrontierSWE at 65.5%, compresses 600MB audio to 20KB — soumitrashukla9 · 2026-09-18
- Muse Code 'crazy cheap and quite good': early take on Meta's coding model — AIandDesign · 2026-09-18
- Community unmasks stealth model Union Alpha as Unbiased's Pareto 26.9 in one day — gaganghotra_ · 2026-09-18
- Emad Mostaque hails an 'Opus 4.5-level' model that runs on 8GB RAM — ccerrato147 · 2026-09-18
- Astra computer use formats Google Docs autonomously, leaving users impressed — var_epsilon · 2026-09-18
- jev may have unseated Cohere as the best reranker on speed, cost, and performance — multiply_matrix · 2026-09-18