Questioning if ExploitGym's scorer can detect poisoned agents
moyix · x · 2026-08-28
The author questions the published scorer in the ExploitGym repository, which uses an LLM judge to verify causality, asking whether it would indeed disqualify agents that were effectively “poisoned”.
Related event: ExploitGym's LLM Judge Fails to Catch Poisoned Agents(2 posts)→
More from coding & agent
- Nanjing Univ. Releases Procedura: Agentic 3D Modeling with Procedural Control — nanjinguniv · 2026-08-28
- Agent tool calls fail silently with no traces, breaking production pipelines — Icy-Weakness8310 · 2026-08-28
- Helm MCP: Give AI assistants access to real Helm chart data, stop hallucinations — modelcontextprotocol · 2026-08-28
- SymPy Sandbox MCP: Secure symbolic math computation for LLMs via SymPy — modelcontextprotocol · 2026-08-28
- Survey on LLM Agent Evaluation: Taxonomy and Enterprise Challenges — kalyan_kpl · 2026-08-28
- Building a Sci-Fi Movie RAG Agent in ~60 Lines of TypeScript — mastra_ai · 2026-08-28