Apertus 1.5 is a 70B European foundation model with native image and speech support
AxSaucedo · x · 2026-07-27
Apertus 1.5, a 70B European foundation model, adds native image and speech capabilities
ETH Zurich, EPFL, and the Swiss National Science Foundation released Apertus 1.5, a 70B foundation model from Europe. The post says it includes native image understanding and speech processing.
The chart in the image compares its visual performance on 33 image benchmarks:
- Apertus 1.5 8B: 55.0
- Apertus 1.5 70B: 53.6
- Qwen 3 VL 8B: 61.7
- Gemma 4 12B: 57.1
- Molmo 2 8B: 57.1
- Gemma 3 27B: 53.1
- Pixtral 12B: 50.6
- EuroVLM 9B: 46.4
The post frames Apertus as part of a broader wave of European foundation model activity.
More from Multimodal
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- ComfyUI trick: aux preprocessor + Qwen transfers poses across characters with one prompt — Acceptable-Work8202 · 2026-09-23
- Same portrait prompt across Midjourney V6.1, V7 and V8.2: do older models look better? — tisch_eins · 2026-09-23
- Testing AI character consistency across a 20-image travel sequence — SiennaVaire · 2026-09-23
- Midjourney v8.2 Faces: New Portrait Generation Samples Shared — azed_ai · 2026-09-23
- One-sentence prompt generates lifelike dog video, shown side-by-side with the real one — wgrathwohl · 2026-09-23