Thinkymachines Expands Evaluation Team Focused on Model Utility
ziqiao_ma · x · 2026-08-31
Thinkymachines is growing its evaluation team, focusing on whether models are actually useful and if measurements serve as a valid y-axis for scaling research. The role spans a wide range of problems, including building signal-bearing internal evals for predictive scaling, turning user feedback and product workflows into evals that close usability gaps, auditing graders and harnesses, and developing new benchmarks for customizability.
More from Companies & People
- Personal Site Analytics: AI Referrals Only 0.19% in August — gaganghotra_ · 2026-08-31
- Speculation on Permanent Sandbagging by OpenAI and Anthropic — scaling01 · 2026-08-31
- Sapient Observes Cross-Industry Demand for Autonomous R&D — JaynitMakwana · 2026-08-31
- AI Agent Builders Struggle to Translate Customer Feedback into Features — Srinidhi_Murali · 2026-08-31
- Startup idea: A push-based Google Alerts 2.0 with deep research — garrytan · 2026-08-31
- SF isn't dead: OpenAI and Anthropic are the new pillars — garrytan · 2026-08-31