Researcher explains how to build applied multimodal projects: own datasets, metrics, models

mariyaivasileva · x · 2026-09-21

Responding to a question about breadth vs. depth, researcher Mariya Ivasileva walks through her applied vision-language project: since no off-the-shelf solution existed, she sourced her own dataset, defined and validated her own metrics, and proposed her own model. Her approach: decompose the problem step by step — how to represent abstract visual relationships, learn embeddings that encode them, find where they fall short, and improve efficiency — with the goal of a solution that satisfies the application's constraints rather than deep coverage of every subproblem.

Original post →

More from Research

Research channel →