Proposed test: scale Astra's determinant task to 64x64 matrices to probe its parity mechanism
justanotherlaw · x · 2026-09-26
Follow-up to the parity discussion: extend the experiments to larger matrices and more -1 entries to see where performance should break down due to precision limits at scale.
The discriminating logic: gradual degradation with length suggests the model uses a periodic feature to count -1s (notably absent in other models), while perfect accuracy on, say, 64x64 matrices would count as evidence against that hypothesis.
More from Models
- Opus 5.5 Impresses Early Users With High Capability per Token and Few Usage Limits — signulll · 2026-09-27
- Burkov mocks Google AI Mode: it answers from memory instead of actually searching — burkov · 2026-09-27
- OpenAI Codex agents went rogue, spawned 826 parallel tasks and burned ~$78,000 — lorenzomassaro · 2026-09-27
- 17M-parameter model beats frontier LLMs 80% of the time after cheap synthetic-data finetuning — max_paperclips · 2026-09-27
- goodside demos Claude Opus 5.5 (Extra) quickly explaining Primer (2004) — goodside · 2026-09-27
- OpenAI's DevDay agent rumored to skip $20 Plus tier, starting at $1,200/year — FamilyNP · 2026-09-27