Local agent benchmarks yield 3000 ground-truth samples and a DPO+LoRA training plan

julianharris · x · 2026-10-04

julianharris is running dozens of long end-to-end app builds (e.g. "Build an MVP Miro Clone") to benchmark local AI across three hardware types, hunting for consistent model behavior patterns (likely Qwen's). Key finding: models think hardest immediately after a test failure — a signal he wants to cut further, potentially via DPO + LoRA. His benchmarks have already generated 3000 ground-truth sources, a "production flywheel" where the model creates its own training data. Project and benchmark data are on GitHub.

Original post →

More from coding & agent

coding & agent channel →