Pangram Can Be Defeated by a Fine-Tuned LLM
theshawwn · x · 2026-07-19
The author cautions: Pangram can indeed be defeated by a fine-tuned LLM. Even though the model was fine-tuned for other purposes, this result is still highly noteworthy.
The attached image shows Pangram's detection interface: a document is labeled as Human Written, with the right panel stating 100% of this text is Human Written. The author presents this as an interesting case study demonstrating that such detection systems are not invulnerable to adversarial attacks.
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11