PostTrainBench v1.2 发布:Fable 5.1 以 44.6% 登顶,Opus 5.5 居次
dejavucoder · x · 2026-10-02
PostTrainBench v1.2 更新发布,新版本榜单头名易主:Fable 5.1 以 44.6% 的成绩成为排名第一的 agent,Claude Opus 5.5 位居第二。作者还宣布该 benchmark 现已支持通过 Harbor 框架自行复现运行,详细说明见原帖线程。
「模型」频道最新
- Cohere、Perplexity 等密集发布新嵌入模型,multi-vector 生产端持续扩展 — lateinteraction · 2026-10-02
- 微软发布 MAI-Transcribe-2-Streaming:流式转写准确率登顶,快 55% 便宜 60% — mustafasuleyman · 2026-10-02
- GPT-6.1 Sol 3D 谜题翻车:agent 偏爱俯视,不爱转镜头 — patience_cave · 2026-10-02
- GPT-6.1 Sol 在 MazeBench 仅得 9%,险胜 Claude Opus 5.5 — patience_cave · 2026-10-02
- LlamaIndex 发布 Extract v2.5,文档抽取性价比超 Claude 与 GPT — llama_index · 2026-10-02
- Claude API 与平台出现性能降级,信用额度到账延迟 — ClaudeAI-mod-bot · 2026-10-02