IFM's TxT360-v2 training dataset trends on Hugging Face, 1B-10B rows
IFM · hf · 2026-09-08
IFM's TxT360-v2 dataset is trending on Hugging Face: a text-generation pretraining dataset sized 1B-10B rows, released in parquet format under cc-by-4.0 with support for dask and polars.
More from Research
- Mitra-v2: Synthetic-Data-Only 77M Tabular Model Matches 1.6B Rivals on TabArena — chaumian · 2026-09-08
- Sony AI's Hakken system turns 1.5M aging hypotheses into 2 confirmed gene discoveries — i_dg23 · 2026-09-08
- New worklog details building an async RL framework from scratch in JAX, from multi-actor systems to weight sync — yoshiyama_akira · 2026-09-08
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08
- EMNLP 2026 Paper: Training World Models for Behavior Consistency Cuts False Positives from 42.5% to 9.5% — 机器之心 · 2026-09-08
- Track4World: HKUST and Tencent ARC's feedforward model densely tracks every pixel in 3D — rsasaki0109 · 2026-09-08