TinyAya v0.3 open speech-to-speech model adds text supervision and 36-layer LoRA
irombie · x · 2026-07-23
TinyAya v0.3 is an open Turkish↔Hindi speech-to-speech translation model built on a Cohere2 backbone with LoRA fine-tuning and a frozen Moshi depth decoder.
Key details from the thread:
- The corpus includes 840,426 × 2 word-level alignments.
- The project shifted from audio-only training to a text + audio multitask setup, and this is the first time the project supervised text.
- A 6-arm recipe re-validation sweep at a 5K-step horizon found that LoRA adapters on all 36 layers worked best; top-layer adapters turned out to be the most useful text lever.
- The authors added a pipeline-validation gate before scaling, which caught a real off-by-one bug in the eval harness.
- They also fixed 9 XLA blockers to make layer-scan training work on v6e-8 slices.
The post says all recipes, configs, W&B runs, and postmortems are open.
More from Research
- ACM says it will open its digital library to LLMs in a controversial AI-written notice — TobyWalsh · 2026-07-23
- A meme lines up AI’s favorite metaphors, from “stochastic parrot” to “blurry JPEG” — SchoeneggerPhil · 2026-07-23
- The Stack v3 releases 5T code tokens across 700+ programming languages — LoubnaBenAllal1 · 2026-07-23
- OpenMed marks one year as healthcare stays the weak spot in open-source AI — MaziyarPanahi · 2026-07-23
- Fable is used to help disprove a 50-year-old algebraic geometry conjecture — Dr_Singularity · 2026-07-23
- Scientists find sperm whales change vowel-like sounds near boats — begusgasper · 2026-07-23