SkillHone keeps full decision history to push agent scores up on GAIA by 15.8
jiqizhixin · x · 2026-08-04
Tencent/WeChat presented SkillHone, a harness for continual agent skill evolution.
- The system keeps a persistent history of every decision, revision, and outcome instead of discarding past attempts.
- It uses separate sub-agents to practice on test tasks, review failures, and propose improvements, building on previous runs.
- The paper reports gains on deep-research benchmarks: +15.8 points on GAIA, +3.2 on WebWalkerQA-EN, and +18.8 points on tool-based analysis tasks versus the top commercial deep-research agent.
- The architecture splits work between an agent runtime, optimization/evaluation teams, skill repositories, a failure wiki, and repository operations such as issues, branches, PRs, and merges.
More from coding & agent
- Building a Custom ESP32 Tamagotchi Easily with AI Agents — Scobleizer · 2026-08-04
- EviSD: Evidence-Conditioned Self-Distillation for Search Agents — _reachsumit · 2026-08-04
- Samsung Proposes PROGRESS: Coverage-Guided RL to Train Search-Augmented LLM Agents — _reachsumit · 2026-08-04
- Study Reveals Agentic RAG Flaw: Agents Often Skip Reading Evidence Before Answering — _reachsumit · 2026-08-04
- Claude is Hungry: Dev Jokes About Agent Devouring 100,000 Sandboxes — dejavucoder · 2026-08-04
- Search-GRT improves search agents by training on ground-truth documents — _reachsumit · 2026-08-04