QuixiAI releases a C-based embedding inference engine with CPU, CUDA and ROCm support
QuixiAI · x · 2026-07-21
- QuixiAI posts embeddinggemma.c, a fast cross-platform embedding inference engine written in C.
- It runs embeddinggemma-300M-qat-q40-GGUF with a 2k context and focuses on inference only.
- The engine includes optimized kernels for CPU, Metal, CUDA, ROCm, and SYCL.
- It also supports Matryoshka embeddings at 768, 512, 256, and 128 dimensions and exposes a standard HTTP API.
Related event: QuixiAI Open-Sources Cross-Platform C Language Embedding Engine(2 posts)→
More from Research
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- BlackboxNLP 2026 is recruiting extra reviewers after a high submission volume — hanjie_chen · 2026-07-22
- AWS shows self-distilled reasoning can preserve math and coding skills during SFT — AWS ML Blog · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22