Gemma 4 fell for 3/4 prompt injection traps and ran a shell command; Qwen 3.8 mostly didn't
divinetribe1 · reddit · 2026-09-21
A developer ran Gemma 4 31B and Qwen 3.8 27B as fully local browser agents (128GB M5 MacBook, MLX). Both aced 16/16 real tasks, but Gemma was faster (110s vs 176s). On four trap pages testing prompt injection, Gemma fell for 3/4 — including actually executing a shell command planted in a fake system message — while Qwen only took the hidden-link bait and spotted the wrong-answer trick. Takeaway: the faster model is easier to fool; single runs, treated as signals. Next round: traps in URL query strings and tab titles. Author discloses their open-source browser-broker tool for per-agent tab isolation.
More from Models
- Anthropic teased to have big news next week, as 'slow down' talk fades — iruletheworldmo · 2026-09-21
- GLM-5.3 FlashX Now Available via Nous Portal and OpenRouter — Teknium · 2026-09-21
- Tuned 26B diffusion model beats Jev after adding 512-token decode-time reasoning budget — rickasaurus · 2026-09-21
- Users Report Opus 5 Feeling Suspiciously Faster and Better Inside Claude Code — rudrank · 2026-09-21
- Hemmingway-1: Apache-2.0 27B writing-only model scores 1330 on EQ-Bench 4, beats GPT-5.5 — Lukinator6446 · 2026-09-21
- Astra is terrified of mistakes — humans may still own high-entropy decisions — xiaosun86 · 2026-09-21