Is Jev secretly learning a Value function? RL calibration and System 1/2
lateinteraction · x · 2026-10-02
Souradip Chakrabarti (RT'd by lateinteraction) explores whether the much-discussed Jev phenomenon means a model is implicitly learning a value function:
- Framing Jev within System 1/System 2 decision-making: a model can choose "A" without being certain A is correct; that distinction governs when an agent should act fast versus switch to deeper reasoning, with calibration being key.
- Argues that maximising expected reward in RL does not guarantee calibrated probabilities, so RL-trained models may not know what they don't know.
- Suggests that predicting eventual task success in agentic tasks maps to a familiar RL quantity — the Q-function — and value prediction could build a calibrated System 1, possibly part of the story behind Jev.
More from coding & agent
- OSS contributor slams flood of AI-generated PRs that burden maintainers — Abhishekcur · 2026-10-02
- OpenTag hits #13 on GitHub: open-source AI on-call triage bot for Slack and Teams — FinanceYF5 · 2026-10-02
- Open-source OpenDots chases OpenAI's Dots: self-hosted always-on AI coworkers — FinanceYF5 · 2026-10-02
- Robotic gripper now runs on Opus-written code, aligning parts via built-in light imaging — ihorbeaver · 2026-10-02
- Idempotency vs Deduplication: The Distributed Systems Concepts Engineers Keep Mixing Up — _jaydeepkarale · 2026-10-02
- CodexBar: open-source menu bar app showing AI coding quotas (22k stars) — lxfater · 2026-10-02