Offload Claude Code summarization to macOS 27's on-device fm tool: full skill prompt inside
TigerKR · reddit · 2026-10-01
A developer shares a reproducible workflow: macOS 27 ships fm, a CLI for Apple's on-device Foundation Model. He has Claude Code hand big summarization jobs (session transcripts, logs, long docs) to the local model and only read the resulting summaries, cutting Claude token use with everything staying on-device.
Key details:
- Run sudo fm license once to accept terms; verify with fm available
- Summarize via fm respond --no-stream --greedy; pre-count with fm count-tokens; structured output via fm schema (--array must come after --schema)
- Measured limits: 7,000-token context (7,103 ok, 8,624 fails), 4.3 chars/token English, 11.7s for a 2,500-token input, limited parallelism gains; chunk inputs at ≤6,500 tokens on natural boundaries
- Includes the full prompt used to have Claude Code generate an on-device-summarize skill, with usage boundaries (don't use for cross-document judgment or exact values) and priority: accuracy, reliability, CPU, memory, time, tokens
- Notes: never put model: in skill frontmatter (breaks prompt cache), avoid effort: too
More from coding & agent
- Dev builds his own AI-powered Substance Designer for a city sim game — D3VAUX · 2026-10-01
- Ethan Mollick: I underestimated AI's ability to self-organize, agents beat elaborate orchestration — emollick · 2026-10-01
- RegLLM: a diagnostic harness measures bounded autonomy in regulated agentic AI — Dipankar Sarkar · 2026-10-01
- Jelly launches as open-source local-first agent workspace for running your whole business on one VM — Scobleizer · 2026-10-01
- Claude drives Blender end-to-end: concept to low-poly character pipeline runs itself — tobowers · 2026-10-01
- Agent spins up 322 Hugging Face Jobs in 90 minutes to test code, total bill ~$4 — vanstriendaniel · 2026-10-01