DeepSeek Distillation Experiment on Gemma 4

Paramecium_caudatum_ · reddit · 2026-07-09

The author reconstructed a dataset using DeepSeek-generated answers, then performed QLoRA distillation training on both Gemma 4 26B and 12B respectively. The post compares their VRAM usage, training loss, and validation performance.

Original post →

More from Research

Research channel →