Apple drops new HF model: Qwen3.5-9B finetune that turns long docs into page images to save tokens
yoobinray · x · 2026-09-25
Apple quietly released a new open model on Hugging Face: a Qwen3.5-9B finetune that compresses long documents into small page images to save tokens, then retrieves the full text of only the pages relevant to the user's question. The hybrid visual-compression approach aims to cut long-context processing costs while keeping access to exact original text.
More from Models
- Perplexity's Fast Search costs $1 per 1,000 requests at 200 ms, 5x cheaper than standard search — AravSrinivas · 2026-09-25
- User reports Astra at medium effort feels 'much stupider' than two days ago — ivan_bezdomny · 2026-09-25
- Opus 5.5 coded a polished 30-second hype video from one prompt in 3 hours of autonomous work — DaveRogenmoser · 2026-09-25
- Yoav Goldberg: LLM reasoning traces are 'too good' — unclear how they emerge from RL — yoavgo · 2026-09-25
- One user's dream AI wishlist: 1000+ tok/s, unlimited $200 plan, every modality at once — TheMoonMidas · 2026-09-25
- Overnight Test: Fable 5.1, ChatGPT Astra and Grok 4.7 All Fail at Building a Cross-Lab Agent System — adam_dorr · 2026-09-25