TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference
Hanzhi Zhang · hf · 2026-08-25
TileMix routes attention score tiles to mixed FP16 or INT8 precision within fused dense attention. It recovers long-context accuracy while improving prefill throughput without retraining.
More from Infra
- Fal releases post-trained H3 model co-optimized with custom inference stack — isidentical · 2026-08-25
- Optimizing Qwen3-ASR Latency to 70ms to Beat Deepgram — Comprehensive_Quit67 · 2026-08-25
- Agent Memory System Refuses Hallucinations with Provable Deletion and Auditability — External-Fee-8920 · 2026-08-25
- $3 ESP32-C3 Runs TFLite for Industrial Anomaly Detection — blaizedsouza · 2026-08-25
- LayerStoRm: Run Frontier-Scale MoE LLMs on Consumer GPUs via PCIe Streaming — CharacterBumblebee99 · 2026-08-25
- 4070 Ti Benchmarks: Running Kimi K3, DeepSeek V4, and Qwen 122B Locally — JayB_Official · 2026-08-25