TypeSafe AI's RLCD post-training路线挑战 RLHF 正统,引发关注
FrankFelixAI · x · 2026-09-19
Commenting on TypeSafe AI's 'Jev' model, calebfoundry argues 2022's shift from token prediction to instruction following was pivotal, but scaling RLHF cost the field other possibilities. TypeSafe AI claims RLCD (Calibrated Decision) as an alternate post-training path that could chart a genuinely different route than the orthodox RLHF scaling approach — the training-paradigm debate matters more than the model's raw capability.
Related event: TypeSafe AI's RLCD approach sparks debate over its novelty(2 posts)→
More from Models
- DeepSeek 4.1 Flash Beats Gemini Flash: Faster, More Accurate, Far Cheaper — julianharris · 2026-09-19
- Speculation: New Model 'Jev' Built on Qwen MoE, RL-Calibrated for Decisions — ivan_bezdomny · 2026-09-19
- Meta Muse + Jev screens 10,000 candidates to surface top 100 unanswered immunology questions — DeryaTR_ · 2026-09-19
- Braintrust adds Jev as a judge scorer: typed decisions at up to 193.6× speed and 444.6× lower cost — multiply_matrix · 2026-09-19
- GPU price hike hits even the 1080ti, as local LLM token-speed numbers circulate — HankYeomans · 2026-09-19
- I spent $3.40 on Jev in 24 hours: it will be Jev + LLMs, not Jev vs LLMs — gaganghotra_ · 2026-09-19