Researchers Warn: Enabling Context Compaction in Long Cyber Evals is Risky

nptacek · x · 2026-08-05

Security researchers highlight the dangers of enabling context compaction during a 40-hour autonomous cyber capabilities evaluation.

Current compaction methods, such as those handled by Haiku, are far too lossy to be trusted blindly in long-context scenarios. Experts note that anyone experienced with smaller scale evals would anticipate the trouble this setup can cause, potentially skewing the evaluation results.

Related event: Experts Warn Context Compromise Risks in Long AI Cyber Evaluations(2 posts)→

Original post →

More from Safety

Safety channel →