Jev tested: prompt tips and an open 27B model matches it
Testing TypeSafe's Jev yielded two prompt lessons (e.g., always add a 'none of the above' option), and a local open-source 27B model matched Jev's accuracy within 1.4 points with even better calibration on sentiment tasks.
2026-09-21 ~ 2026-09-21 · 2 related posts
- Open 27B model on a single GPU matches Jev's accuracy with better calibration — colinmcnamara · 2026-09-21
- Testing decision model Jev: add "none of these" options and never do algebra across questions — colinmcnamara · 2026-09-21