Training Models on the Pelican Benchmark: A Fun Use Case with TRL and HF Jobs

SergioPaniego · x · 2026-07-30

Developer Sergio Paniego shared an engaging experiment: transforming Simon Willison's famous "pelican riding a bicycle" prompt—often dubbed a deeply unscientific benchmark—into a functional AI evaluation and training environment.

The article details how to set up this system using OpenEnv, TRL, and HF Jobs. The benchmark scores AI-generated SVG drawings across three layers and provides two core commands: one to evaluate existing models on an HF Space, and another to fine-tune a model against the benchmark on a GPU.

Original post →

More from coding & agent

coding & agent channel →