Sandboxing AI evals across harnesses and models is hard, researcher jokes

kevinnbass · x · 2026-08-06

Kevin Bass tweets that sandboxing evals across harnesses and models to ensure scientific accuracy is hard, and jokes that he is slowly going insane. This highlights the technical challenges in AI evaluation.

Original post →

More from Research

Research channel →