FW-Icon: An Expert-Graded Benchmark for SVG Icon Creativity and Execution
andpoul · x · 2026-10-02
andpoul's team launched FW-Icon, an expert-graded benchmark measuring LLM creativity and execution on SVG icon generation. For transparency, they published the prompts, model outputs, and grades on their website, and are inviting labs training agents for judgement and idea diversity to collaborate. It's a rare expert-scored creativity eval built around a concrete design task.
More from Research
- EMNLP Findings paper: explicit reasoning hurts pointwise rerankers, disabling it helps — lintool · 2026-10-02
- Paper learns Alzheimer's signatures from EEG via spiking neural networks and biophysical simulation — giorgiodidio · 2026-10-02
- ETH Zurich: Repulsive Self-Distillation Destabilizes Training, Contrastive Distillation Wins — arkrause · 2026-10-02
- Google team's new paper: complex multi-step reasoning emerges in Neural Cellular Automata — pbaylies · 2026-10-02
- Ofir Press: a good benchmark needs scalable data collection — the hardest step yet — OfirPress · 2026-10-02
- NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite — rohanpaul_ai · 2026-10-02