Grok 4.5 Shows More Persistence in Code Audits
avaitopiper · x · 2026-07-10
The forwarded content mentions that Composio found Grok 4.5 to be one of the most "persistent" agent models they have ever tested during their trials.
In a sample task where three models were asked to search code for hardcoded credentials in a GitHub repository, both GPT-5.5 and GLM-5.2 stopped at the first page. In contrast, Grok 4.5 continuously paginated until results were exhausted, successfully completing the audit.
More from coding & agent
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- ty adds first-class Pydantic support, including strict and lax field handling — charliermarsh · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22