QuixiAI teases a small, fast omni embedding model for image, video, audio, and text
QuixiAI · x · 2026-07-21
QuixiAI says its next project will be quixi-embed, an omni embedding model for image, video, audio, and text.
The project is described as small and fast, and it will use the same custom optimized kernel treatment as the author’s other work. The reply notes a benchmark run with embeddinggemma-300M-qat-q40-GGUF at a 2k context window, but the main point is the upcoming multi-modal embedding model.
More from Multimodal
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21
- Night-party video demo uses Seedance 2.0, timecode prompts and 4K upscaling — gen_ericai · 2026-07-21