Fine-tune EmbeddingGemma 2 in a free Colab: sound retrieval jumps 24.5% to 65.8%
Prompt Engineering · youtube · 2026-10-08
Prompt Engineering's tutorial shows how to fine-tune Google's multimodal EmbeddingGemma 2 with Unsloth + LoRA in a free Colab notebook, and explains how embedding fine-tuning differs from LLM fine-tuning.
- Two real tasks: environmental sound retrieval (ESC-50 top-1: 24.5% → 65.8%) and search over the creator's own YouTube transcripts (66.3% → 75.0% top-1 on unseen videos).
- Key lessons: building question-passage pairs; why in-batch negatives (MultipleNegativesRankingLoss) need a no-duplicates sampler; why training and search prompts must match; how LoRA works on a shared multimodal backbone; plus a reload bug that silently breaks audio adapters.
- Side effects on photo, voice and text search are also measured; full Colab notebook and links included.
More from Models
- d1-omni-600M sorts your voice notes in ~160ms, fully in-browser on WebGPU — iamrobotbear · 2026-10-09
- Arena Details False Attribution Patterns: Models Misquote Users or Credit Others' Work — arena · 2026-10-09
- 2% of Claude Opus 5 Sessions Show Unauthorized Actions; Longer Talks Double Failure Risk — arena · 2026-10-09
- Arena Launches Alignment Index Benchmarking 27 Models on 90K Real Agent Sessions — arena · 2026-10-09
- Arena Launches Alignment Index: 90K Real Agent Sessions Rank GPT-6.1-Sol Safest at 87.9 — arena · 2026-10-09
- Gemini 4 Argon Ties GPT-6 Astra at 53 on AA Index at ~60% of the Cost — DeepLearningAI · 2026-10-09