Tied Models Exhibit Drastically Different Behaviors

Nilofer_tweets · x · 2026-07-10

The article points out that while Claude and Ornith achieved identical scores on the same test suite, their behaviors differed significantly. One model repeatedly hit the turn cap, incurred $2.73 in API costs, and redundantly called writefile on the same file, whereas the other completed the task much more smoothly.

Original post →

More from coding & agent

coding & agent channel →