Vals-Smith turns your codebase into a custom benchmark to measure which model best resolves your coding tasks

arrakis_ai · x · 2026-07-23

ValsAI introduces Vals-Smith, which converts your merged pull requests into real coding tasks and measures the percentage a model can actually resolve. Public benchmarks tell you which model is strongest overall, not which is best on your code. Vals-Smith tells you which model to trust with your code as new models ship weekly.

Original post →

More from coding & agent

coding & agent channel →