Indie dev offers to train a fully open 9.4B dense model, tuned for a single GPU
NineThreeTilNow · reddit · 2026-09-13
A LocalLlama user is offering to train and open-source a 9.4B parameter dense model, asking the community what to target:
- Architecture: 1/2/3-layer Engram table injection, Moonshot-style AtRes modeling, 3:1 RoPE/NoPE layering; starts from the Llama 3 tokenizer and LM head, with logit-level data extracted from a Llama 3 teacher model.
- He produced all training data on a single 4090 plus rented hardware; the code is deeply optimized for one RTX 6000 Pro card, aiming to beat existing 9B models and add open-standard 'thinking'.
- Qwen's recent release showed a single Engram table captures most of the benefit, so he cut the design from two tables to one.
- Code and data are already in public repos; along the way he found and reported vLLM prompt-loading code that could be 10-100x faster.
More from Models
- DeepSeek V4.1 costs less to serve but prices 2x higher per output token, margins likely up — teortaxesTex · 2026-09-13
- Motorhomelabs Teases New Release Amid Prediction Open Models Face Strict Regulation — aiamblichus · 2026-09-13
- DeepSeek has no single research direction, but 'infinite context' is the goal — teortaxesTex · 2026-09-13
- User shows how GPT flatters your beliefs and hallucinates premises in arguments — GlenBradley · 2026-09-13
- Orchestrator Error Reveals Mystery Model 'Daybreak': 'astra Is Not Allowed to Access Those Resources' — LeopardBernstein · 2026-09-13
- kalomaze: Opus 5 is "such a bad model" — kalomaze · 2026-09-13