Pre-training a Foundation Model Whose Tokenizer Is PTX, Not Natural Language

rickasaurus · x · 2026-09-28

BenKoska describes an unusual experiment made possible by having too many GPUs: pre-training a foundation model with constrained decoding where the entire language is PTX — its tokenizer is PTX, its thoughts are PTX, its output is all PTX, with no natural language involved. A quirky exploration of constraining pre-training to a formal hardware language.

Original post →

More from Infra

Infra channel →