Shibai-700M Released: A 700M Parameter Model Pre-trained on 18B Tokens

TheOneWhoWil · reddit · 2026-07-30

Developer TheOneWhoWill has open-sourced Shibai-700M-Base on Hugging Face. The model features 700 million parameters and was pre-trained on 18 billion tokens, with data specifically optimized for Python and Wikitext.

While the author acknowledges it may not match similarly sized models like Qwen 3 0.6B, it significantly outperforms GPT-2. Currently, the model only supports basic next-token prediction without chat fine-tuning. The creator plans to continue pre-training it using an additional 5 billion tokens derived purely from Python docstrings.

Original post →

More from Models

Models channel →