HF daily papers: DeepSeek KV cache at 890 bytes/token, auto-research loop cuts agent tokens ~45-49%
ThomasAger · reddit · 2026-09-19
The author highlights the top 3 papers on HF Daily Papers, all relevant to local LLMs and agent harness optimization:
- DeepSeek-V4.1-Flash: KV cache compression — cross-layer KV reuse plus FP4 KV caching brings global KV cache to 890 bytes per token, roughly a quarter of V4-Flash.
- SoL-Pi: recursively scaling auto-research loops — auto-research loops that improve the agent harness, cutting token traffic by 44.7–49.0% at comparable performance.
- Empirical study of harness design for coding agents — varies planning, action space, and context management across 176 settings to quantify what each component actually contributes.
An unusually dense trio covering KV compression, automated harness optimization, and harness design ablations.
More from Research
- JevBench v1 puts nine typed-decision models head-to-head across 242 decisions — airesearch12 · 2026-09-19
- FlashREINFORCE debuts: critic-free, single-rollout async RL for agentic LLMs — CatAstro_Piyush · 2026-09-19
- AI companies are conquering math — and exposing a discipline built on competition, not understanding — danbri · 2026-09-19
- CMU Autonomous Science Lab to Feature at Enamine Drug Discovery Conference — olexandr · 2026-09-19
- AI & Science feature in Scientific American gets a positive write-up — JMateosGarcia · 2026-09-19
- Terence Tao: If Math Is More Than Proof, We Must Celebrate the Rest of It — num42 · 2026-09-19