Pre-training a Foundation Model Whose Tokenizer Is PTX, Not Natural Language
rickasaurus · x · 2026-09-28
BenKoska describes an unusual experiment made possible by having too many GPUs: pre-training a foundation model with constrained decoding where the entire language is PTX — its tokenizer is PTX, its thoughts are PTX, its output is all PTX, with no natural language involved. A quirky exploration of constraining pre-training to a formal hardware language.
More from Infra
- Fireworks' Ember-1 post-trains Kimi K3 to reason 40% more concisely at same quality — isidentical · 2026-09-28
- One AMD driver flag boosts dual-GPU Vulkan LLM inference up to 4x — tabletuser_blogspot · 2026-09-28
- PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090 — dl_weekly · 2026-09-28
- Developer Plans to Let Codex Pick Which Tests Run, Slashing CI Costs — sull · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Is a vector database enough for production AI agents? Reddit debates storage design — OkShirt9372 · 2026-09-28