Caveman Renders AI Agent Skills as Images, Cutting Token Cost by 61%
bigaiguy · x · 2026-10-07
The open-source caveman tool (part of a 103k-star project) tackles a hidden token cost: every agent skill is prompt text reloaded on each call. caveman convert renders skill bodies into PNG pages that the model reads as images instead — the caveman skill itself dropped from 1,069 to 415 estimated tokens, a 61% cut.
Safety design:
- Only converts when image pages actually beat text
- Failed conversions leave files byte-identical and name the gate that rejected it
- --dry-run shows token math for every installed skill with no writes
- --revert restores the original byte for byte from SKILL.orig.md
More from coding & agent
- Running 2-3 parallel AI coding sessions beats 5-6, founder finds — IgorCarron · 2026-10-07
- Apple paper: a single well-prompted agent with shell beats multi-agent ML harnesses, 62.5% vs 47.1% Kaggle medal rate — rohanpaul_ai · 2026-10-07
- SOTA models quietly nerfed ~15 times in 2025, dev suggests dual Codex+Claude subscriptions — StewartalsopIII · 2026-10-07
- Kev: an open-source reimplementation of Jev built on Qwen3.5 that runs 100% locally — Arindam_1729 · 2026-10-07
- Google Releases SAM, a P2P Network Letting AI Agents Discover and Call Each Other's Tools — thisguyknowsai · 2026-10-07
- Open-Source PhoneUse Agent Runs Real Android Tasks With Observe-Act-Verify Loop — 381654729fuck · 2026-10-07