Dev porting $65/1B-token model to CUDA Rust to build and optimize a custom inference stack
idanbeck · x · 2026-09-18
- Developer idanbeck is porting a model priced at roughly $65 per 1B tokens to CUDA and Rust, and is building/optimizing an inference stack for it to see how fast it can go.
- A community-side inference engineering effort; benchmark results worth watching.
More from Infra
- A20 Pro runs Apple on-device model at 150 tok/s, half the TTFT of A19 Pro — rxwei · 2026-09-18
- Ai2's $152M NSF OMAI open AI infrastructure goes online with NVIDIA Blackwell Ultra systems — allen_ai · 2026-09-18
- NVIDIA BioNeMo Inference Runtime Delivers 2.9x Faster Structure Prediction for Boltz-2 — PyTorch · 2026-09-18
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- Single RTX 3090 + Qwen 3 8B with deepseek harness 'game changing', dev reports — tlpta · 2026-09-18
- Ben Bajarin on why Credo's full-stack interconnect ownership (SerDes, DSP, photonics) drives its next growth phase — BenBajarin · 2026-09-18