Solar Open 2 Employs Selective Weight Transfer and Efficient Pretraining
Solar Open 2 utilizes selective weight transfer from its predecessor, reusing 5.69B parameters while re-initializing new expert layers. Its pretraining compresses data to 12T tokens and extends context to 1M.
2026-07-26 ~ 2026-07-26 · 2 related posts
- Solar Open 2 reuses 5.69B weights from Solar Open 1, then randomizes experts — keunwoochoi · 2026-07-26
- Solar Open 2’s pretraining trims 20T tokens to 10T and stretches context to 1M — keunwoochoi · 2026-07-26