Shibai-700M Released: A 700M Parameter Model Pre-trained on 18B Tokens
TheOneWhoWil · reddit · 2026-07-30
Developer TheOneWhoWill has open-sourced Shibai-700M-Base on Hugging Face. The model features 700 million parameters and was pre-trained on 18 billion tokens, with data specifically optimized for Python and Wikitext.
While the author acknowledges it may not match similarly sized models like Qwen 3 0.6B, it significantly outperforms GPT-2. Currently, the model only supports basic next-token prediction without chat fine-tuning. The creator plans to continue pre-training it using an additional 5 billion tokens derived purely from Python docstrings.
More from Models
- vLLM Announces Day 0 Support for Kimi K3 Across NVIDIA Architectures — vllm_project · 2026-07-30
- Leaked Opus 5 Benchmarks: Major Jumps in Research Math and Long-Context — echen · 2026-07-30
- Opus 5 benchmarks: #2 in research math and long-context agents, half the price — echen · 2026-07-30
- Kimi K3 Lands on DigitalOcean Powered by vLLM for Efficient Inference — vllm_project · 2026-07-30
- Moonshot Releases 2.8T-Parameter Kimi K3; Modal Achieves 460 TPS with Speculative Decoding — sarahcat21 · 2026-07-30
- Claude Is Down: Service Outage Confirmed by Status Page — gregsadetsky · 2026-07-30