OpenDataArena's Spark-234K, a 234K-sample machine-generated English text dataset, trends on Hugging Face
OpenDataArena · hf · 2026-09-19
OpenDataArena released Spark-234K on Hugging Face, where it is currently trending.
- 234K samples for text-generation tasks
- Annotations and content are machine-generated, English, monolingual
- Apache-2.0 license, parquet format, loadable via the datasets library
More from Research
- Specter Core: an open-source multi-agent governance experiment with no final authority, not even the ASI — AIGODSEND · 2026-09-19
- AgentVLN: a 3B VLM brain tops R2R-CE and RxR-CE in real time on Jetson — jiqizhixin · 2026-09-19
- willccbb: self-play + self-calibration is how you do recursive self-improvement — willcb · 2026-09-19
- Split Monolithic Agents Into a Generation Layer Plus Small Classifiers — aigclink · 2026-09-19
- A Lean-verified proof can still prove the wrong theorem: scrutinizing OpenAI's Navier-Stokes claim — MasterWitcher69 · 2026-09-19
- SE veteran calls for de-anonymized peer review to catch AI-generated paper spam — moarbugs · 2026-09-19