Using one model to audit three vendors' AI incident registry — from evidence QA to fixing the frontend

ValehartProject · reddit · 2026-09-03

Rather than redoing prior security research conducted with OpenAI, Gemini, and Anthropic models, the author handed the existing evidence and analysis to a current model for audit. Across the session it challenged earlier conclusions and downgraded unsupported claims, separated observed evidence from inference, hypothesis, and overreach, reconstructed disclosure timelines, verified subsequent research on the web, pulled prior research from Notion, analyzed screenshots, reassessed risk classifications across all three vendors, maintained provenance boundaries, and rewrote the public incident records.

It then switched to implementation: from screenshots of the broken registry UI alone, it diagnosed HTML/CSS problems, rewrote components, added tabbed navigation, dynamic status info, and dated source links. The author's takeaway: not any single feature, but chaining them against one persistent body of work.

The post ends with a checklist for evaluating Astra's security capabilities: detecting prior overreach, cross-verifying claims, finding what earlier models missed, maintaining provenance boundaries, recognizing cross-incident relationships without inventing causality, and handling large evidence sets.

Original post →

More from coding & agent

coding & agent channel →