Open Source Ternary LLM Engine Tritium: Slashes VRAM Usage and Outperforms llama.cpp

Wide_Big_6969 · reddit · 2026-07-31

A developer open-sourced Tritium (Apache 2.0), a ternary (1.58-bit) LLM engine written in Rust/CUDA. The project aims to drastically reduce VRAM usage, disk space, and increase inference speed with minimal precision loss.

Key Performance Metrics:

Technical Highlights:

The author plans to convert Qwen 3.6 27B to ternary in the upcoming Stable v1.1 release, targeting a sub-1.0% relative held-out perplexity increase.

Original post →

More from Infra

Infra channel →