Aiden Bai clarifies: ReactBench uses Harbor, but eval tooling woes are industry-wide
aidenybai · x · 2026-10-10
Follow-up to his post on eval tooling: Aiden Bai clarifies that his team uses the Harbor framework for ReactBench/data work, but stresses the eval tooling problems are not a Harbor-specific issue — they are industry-wide.
Related event: Aiden Bai: AI eval tooling falls short industry-wide(2 posts)→
More from coding & agent
- Setting /autocompact to 400k saved 29% of weekly Claude usage without hurting performance — rickasaurus · 2026-10-10
- Eazo hands-on: one-prompt apps with 6 design variants, Stripe payments, 50-70% referral cuts — vista8 · 2026-10-10
- Nested parallelism: running Hermes agents inside Grok bot VMs for compute offload — alexcovo_eth · 2026-10-10
- Brian Holt sold his 42U homelab — a Framework Desktop plus coding agents does more — film_girl · 2026-10-10
- Group-Evolving Agents wins COLM 2026 award, hits 71% on SWE-bench Verified — xwang_lk · 2026-10-10
- Asking an agent to fix a bug you don't understand is continuous paperclip maxxing — brandon_xyzw · 2026-10-10