True Evaluation Should Focus on Real-World Collaboration
curious_vii · x · 2026-07-11
This reply pushes back against the generalization that "chat makes us dumb": what truly matters isn't abstract testing, but whether a model can help a neighbor achieve a beneficial goal under real-world conditions of uncertainty.
Furthermore, because human-model inputs and outputs are increasingly text-based, searchable, and traceable, every sentence and paraphrase becomes more significant, meaning high-quality thinking will lead to better outcomes.
More from AGI Musings
- ControlAI CEO says an international ban on superintelligence is needed to avert extinction risk — zetalyrae · 2026-07-22
- Gary Marcus says LLMs still cannot really do math on their own — GaryMarcus · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- AI may make digital work infinitely leveraged while offline life gets more human — illscience · 2026-07-22
- Better AI math could save researchers time by killing false conjectures earlier — prateekj · 2026-07-22
- AI’s economic forecasts are split by nearly a quadrillion dollars by 2035 — bittingthembits · 2026-07-22