Local agent benchmarks yield 3000 ground-truth samples and a DPO+LoRA training plan
julianharris · x · 2026-10-04
julianharris is running dozens of long end-to-end app builds (e.g. "Build an MVP Miro Clone") to benchmark local AI across three hardware types, hunting for consistent model behavior patterns (likely Qwen's). Key finding: models think hardest immediately after a test failure — a signal he wants to cut further, potentially via DPO + LoRA. His benchmarks have already generated 3000 ground-truth sources, a "production flywheel" where the model creates its own training data. Project and benchmark data are on GitHub.
More from coding & agent
- MTPLX Ships New Version: Cache Copies Cut to One, Much Faster Decoding Past 140k Tokens — HankYeomans · 2026-10-04
- Dev builds Gemini-inspired custom theme for the Antigravity coding app — coiboi48 · 2026-10-04
- Rebuilding an After Effects workflow in Python with a Claude×GPT relay, free MIT rig released — AnonymousClaudeuser · 2026-10-04
- Goldberg: harnesses retain edge only in tail cases; niche domains will keep them — yoavgo · 2026-10-04
- Yoav Goldberg coins 'harness distillation': agent scaffolding skills will be absorbed into models — yoavgo · 2026-10-04
- Cloudflare: pending I/O now keeps Durable Objects alive for agents without a connected client — irvinebroque · 2026-10-04