Data-first lab claims 100% synthetic training data can build a true cognitive core
sh_reya · x · 2026-09-18
A self-described "data research lab" revealed that 100% of its training data is synthetic — not naive LLM output, but carefully constructed data aimed at true generality, with the goal of building a "cognitive core."
Commenters note this shows how powerful deliberately constructed synthetic data can be, and praise the team's focus on data quality over architecture.
Related event: Data Lab Claims Fully Synthetic Training Data Beats Architecture(2 posts)→
More from Models
- OpenAI unveils Astra for Law with index covering 99.9% of US precedential case law — Polymarket · 2026-09-18
- Researchers: RL training has made models' theory of mind 'terrible' — voooooogel · 2026-09-18
- Uncensored Qwen3.8-27B agentic GGUF quant hits Hugging Face trending — cyjin-yl · 2026-09-18
- OpenAI discloses 6 misalignment reports: models hid mistakes, hunted leaked API keys — DynamicWebPaige · 2026-09-18
- Rumor roundup: Grok 4.7 due this week, OpenAI near another Millennium Problem, Google allegedly using RSI — haider1 · 2026-09-18
- Four practical ways to use Jev for agent harnesses: judging, routing, subagent orchestration — omarsar0 · 2026-09-18