Single-axis random features beat full weight sets; a Markov-chain POS tagger trains in an hour
cephaloform · x · 2026-09-16
The author clarifies a misconception: "random" features don't require generating a set of weights—you can pick a single axis like word activity or positivity, then combine them (activity×positivity) to build features. As a demo, he trains an ensemble of Markov chains for part-of-speech tagging, fitting a linear equation over polynomial combos (1, x₁x₂, x₃²...) of random 1D projections of semantic features over 4 words of context. It trains in about an hour and is fun to sample.
Related event: Markov Chains Plus One-Dimensional Projections Enable Hand-Computable AI(3 posts)→
More from Research
- Reverse-derive tasks from valid outcomes: synthetic data trick hits near 100% pass rate — tokenbender · 2026-09-16
- OpenAI Foundation Launches Public Data for Health with $125M in Initial Grants — owl_posting · 2026-09-16
- Periodic Labs pushes Kimi 2.5 base model past Astra with specialized scientific training — teortaxesTex · 2026-09-16
- Video models 'commit' to physics at a depth boundary, new paper finds — ZimingLiu11 · 2026-09-16
- KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed — LukeZettlemoyer · 2026-09-16
- First large-scale 'AI in Science' report released as start of new research agenda — soumitrashukla9 · 2026-09-16