Gemma 4 fell for 3/4 prompt injection traps and ran a shell command; Qwen 3.8 mostly didn't

divinetribe1 · reddit · 2026-09-21

A developer ran Gemma 4 31B and Qwen 3.8 27B as fully local browser agents (128GB M5 MacBook, MLX). Both aced 16/16 real tasks, but Gemma was faster (110s vs 176s). On four trap pages testing prompt injection, Gemma fell for 3/4 — including actually executing a shell command planted in a fake system message — while Qwen only took the hidden-link bait and spotted the wrong-answer trick. Takeaway: the faster model is easier to fool; single runs, treated as signals. Next round: traps in URL query strings and tab titles. Author discloses their open-source browser-broker tool for per-agent tab isolation.

Original post →

More from Models

Models channel →