Blogger runs Namedrop eval in 1.5 hours of focused attention
maxsloef described his Namedrop eval as a vibe-research experiment done in about 1.5 hours of attention over three days, using self-guided style references chosen in separate calls; he found LLM judges misalign with human taste, which he suspects is why style weaknesses persist.
2026-09-27 ~ 2026-09-27 · 4 related posts
- Eval design note: models self-select style references in a separate call to isolate the effect — maxsloef · 2026-09-27
- LLM graders can't match human taste on style eval — maybe why the gap persists — maxsloef · 2026-09-27
- A vibe-research experiment: one pairwise comparison study shipped on ~1.5 hours of attention — maxsloef · 2026-09-27
- Vibe-research experiment: Namedrop eval took just 1.5 hours of attention over three days — maxsloef · 2026-09-27