CrossGMN: neural networks learn to compress other networks, 8.89x faster distillation

HaggaiMaron · x · 2026-10-05

A new paper by Adir Dayan, Haggai Maron and colleagues takes a first step toward neural networks that compress other neural networks. CrossGMN learns to transform trained models into different architectures directly in weight space, achieving up to an 8.89× speedup in distillation steps. The author unpacks the idea in an extended thread.

Original post →

More from Models

Models channel →