giffmana pitches multimodal scoring model under 1B params with calibrated image-text scores
giffmana · x · 2026-10-01
- giffmana (known vision encoder researcher) proposes a model that takes any image plus freeform texts and returns how well each text fits the image.
- Could support yes/no judgments via sigmoid heads, with calibrated scores.
- Ambition: open-weight it at under 1B params (400M), though he "might need to raise xxxM".
- He frames himself as both an inventor of vision encoders and one of their biggest critics.
Related event: Researchers Envision a Sub-1B Multimodal Scoring Model(2 posts)→
More from Models
- GPT-6 Astra Ultrafast now roughly as fast as Gemini 3.5 Flash Lite — Angaisb_ · 2026-10-01
- Why the same Llama 3.2 1B comes in different file sizes: quantization explained — night-alien · 2026-10-01
- Ant's Ling-3.1-flash: 560B MoE with 25B active, 1M context, open-source soon — _AndrewZhao · 2026-10-01
- Grok 4.8 spotted in xAI's official Grok Build repo as release prep begins — mark_k · 2026-10-01
- 27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster — bjivanovich · 2026-10-01
- Sol 6.1 Day Two Impressions: Strong at Scheduled Tasks and Data Analysis, Cheap — bindureddy · 2026-10-01