Debate: Direct Projections May Replace Embedding Tables
A technical discussion suggests that over-determined direct projections from low to high dimensions (e.g., 8 to 4096) may outperform per-token embedding tables in representation learning, with architectures like Gemma 4 12B possibly moving toward direct linear projection of raw bits, similar to modern ViTs.
2026-08-19 ~ 2026-08-19 · 3 related posts
- Removing embedding tables: ViT-style projections for tokens — kalomaze · 2026-08-19
- Discussion: Over-determined projections may outperform embedding tables in representation learning — kalomaze · 2026-08-19
- Technical discussion: BPE padding and representation learning — kalomaze · 2026-08-19