Apple's Embedding Atlas and the Modular LLM Inference Handbook: two practical learning resources
techNmak · x · 2026-09-04
Entries 11-12 of @techNmak's visual learning resource thread:
- Apple Embedding Atlas: embeddings are much easier to reason about via neighborhoods, clusters, and outliers than raw vectors — it lets you interactively explore large embedding datasets.
- Modular LLM Inference Handbook: bridges "how LLMs work" and "how LLM systems serve requests," covering prefill, decode, KV cache, batching, scheduling, quantization, prefix caching, and speculative decoding.
More from Infra
- Will combining multiple GPUs' VRAM for local LLMs ever work out of the box? — PusheenHater · 2026-09-05
- Declarative Attention lets LLMs declare their own focus, cutting 52% of KV cache reads — eigenlaplace · 2026-09-05
- Agent outputs die when the VM sleeps: octomind's design for deliverables that survive — donk8r · 2026-09-05
- Japan to develop AI-powered satellites — AIFlow_ML · 2026-09-05
- Hybrid Compute on Mac ships with open-sourced local inference engine and PII classifier — andrewgwils · 2026-09-05
- Tesla's RIM process kills the paint shop, shrinking Cybercab factory footprint ~50% — elonmusk · 2026-09-05