Apertus 1.5 is a 70B European foundation model with native image and speech support
AxSaucedo · x · 2026-07-27
Apertus 1.5, a 70B European foundation model, adds native image and speech capabilities
ETH Zurich, EPFL, and the Swiss National Science Foundation released Apertus 1.5, a 70B foundation model from Europe. The post says it includes native image understanding and speech processing.
The chart in the image compares its visual performance on 33 image benchmarks:
- Apertus 1.5 8B: 55.0
- Apertus 1.5 70B: 53.6
- Qwen 3 VL 8B: 61.7
- Gemma 4 12B: 57.1
- Molmo 2 8B: 57.1
- Gemma 3 27B: 53.1
- Pixtral 12B: 50.6
- EuroVLM 9B: 46.4
The post frames Apertus as part of a broader wave of European foundation model activity.
More from Multimodal
- Grok Imagine is being praised for natural motion and synced audio in AI video — XFreeze · 2026-07-27
- Designer builds beginner-friendly guides to FLUX, Stable Diffusion and ComfyUI — Masha-AI · 2026-07-27
- Local vs. Cloud: Evaluating image generation costs for indie game devs — Simple-Evidence-9125 · 2026-07-27
- Midjourney shows a new painterly style that turns people into blue-and-orange smoke — azed_ai · 2026-07-27
- Zombie Genesis trailer looks like a cinematic AI-generated apocalypse concept — HomeRemedyHealer · 2026-07-27
- Emad says AI 3D workflows are the downstream of a “TikZ unicorn” — emollick · 2026-07-27