Post says an attention-heavy model has 104B active parameters

zephyr_z9 · x · 2026-07-28

A short post quotes a note about a "very heavy attention backbone" and says the model has 104B active parameters, much larger than expected.

With no further context, the key signal is the scale claim itself: the architecture appears to use a much heavier active subset than the commenter anticipated.

Original post →

More from Models

Models channel →