Researcher pitches a multimodal scoring model: any image, any freeform texts, calibrated fit scores

ducha_aiki · x · 2026-10-01

Researcher duchaaiki, replying to @giffmana, half-jokingly sketches a 'multimodal Jevons' idea: given any image and any set of freeform texts, the model would return how well each text fits the image, plus calibrated yes/no outputs via sigmoid. He quips it might take hundreds of millions in funding, and teases giffmana to survive a few years on softmax before jumping to sigmoid.

Related event: Researchers Envision a Sub-1B Multimodal Scoring Model(2 posts)→

Original post →

More from Fun

Fun channel →