Research: Easy Model Identity Laundering via Simple Fine-Tuning

generativist · x · 2026-08-01

The quoted tweet outlines an experiment on LLM identity recognition.

Researchers took 1000 ordinary prompts and obtained answers from teacher models like GPT-4o or Claude 3.5 Sonnet, deliberately removing any mentions of model or lab names.

They then LoRA fine-tuned open-source models on this dataset and evaluated them on identity questions. The experiment demonstrates how easily a model's native identity can be altered or hidden through simple fine-tuning.

Original post →

More from Research

Research channel →