QuixiAI teases a small, fast omni embedding model for image, video, audio, and text

QuixiAI · x · 2026-07-21

QuixiAI says its next project will be quixi-embed, an omni embedding model for image, video, audio, and text.

The project is described as small and fast, and it will use the same custom optimized kernel treatment as the author’s other work. The reply notes a benchmark run with embeddinggemma-300M-qat-q40-GGUF at a 2k context window, but the main point is the upcoming multi-modal embedding model.

Original post →

More from Multimodal

Multimodal channel →