Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights

le_james94 · x · 2026-09-18

A thread of model-behavior observations: fine-tuning Llama 3.1 8B into narrower personas doesn't change refusal on steroid-buying questions, but the sarcastic persona mocks your intent. Separately, Opus 4.6 scores 0% on an ARC-AGI-3 environment without a harness and 97.1% with a hand-crafted one—reasoning-eval noise floors of 0.22–1.48 points mean 1-point press-release gaps are noise.

Original post →

More from Models

Models channel →