NeurIPS papers: CoT faithfulness metrics are near-random, plus MoE router geometry and LLM belief studies
megamor2 · x · 2026-09-28
The author's lab published three NeurIPS 2026 papers:
- A meta-evaluation benchmark with ground truth showing existing CoT faithfulness metrics are either near-random or prohibitively expensive.
- Applying the HOT-3 consciousness indicator to LLMs with interpretability tools, revealing surprising patterns in belief formation and action selection.
- Evidence that sparse MoE routers learn the geometry of their experts, revealing structural coupling between routing and expert layout.
More from Research
- Mathematician prototypes Lean-based tool to teach high school Euclidean geometry with proofs — AlexKontorovich · 2026-09-28
- Columbia's SPEAR framework reframes AI alignment as an ongoing interactive process after deployment — windx0303 · 2026-09-28
- SAGE Uses Topological Guidance to Fix Long-Horizon LLM Reasoning Biases — Xinyue Zeng · 2026-09-28
- LightMIS: 0.13M-Param Medical Segmentation Net Cuts 99% Params, Matches Accuracy — Andrei Arhire · 2026-09-28
- Eval tips for automated optimizers: use AUC over accuracy, sample 10x and check consensus to cut hallucinations — a_karvonen · 2026-09-28
- Three Simple Inference Tricks Significantly Improve Activation Oracle Results — a_karvonen · 2026-09-28