Single-axis random features beat full weight sets; a Markov-chain POS tagger trains in an hour

cephaloform · x · 2026-09-16

The author clarifies a misconception: "random" features don't require generating a set of weights—you can pick a single axis like word activity or positivity, then combine them (activity×positivity) to build features. As a demo, he trains an ensemble of Markov chains for part-of-speech tagging, fitting a linear equation over polynomial combos (1, x₁x₂, x₃²...) of random 1D projections of semantic features over 4 words of context. It trains in about an hour and is fun to sample.

Related event: Markov Chains Plus One-Dimensional Projections Enable Hand-Computable AI(3 posts)→

Original post →

More from Research

Research channel →