Apodex reports process-verifier repair loop beats SOTA in AAV capsid design (+0.155 over 434 retried runs)
omarsar0 · x · 2026-09-04
A proof case from Apodex: the "process score" does more than grade — it drives a repair loop. When the process verifier flags a run as deficient, that diagnosis becomes a repair note, and the solver retries the same task without seeing the answer.
Across 434 trajectories marked deficient, repaired reruns scored 0.155 higher on average per Apodex's evaluation. The company reports a specific model capability in AAV capsid design surpassed the best previously published method in the field — results are self-reported and not independently verified.
More from Research
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05
- Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals — a_karvonen · 2026-09-05
- Video DeltaNet open-sources hybrid-attention VDN-H3, 14.4s video in 11.2s on 8 B200s — BigWideBaker · 2026-09-05
- Deep Learning Weekly #471: Claude Fable 5.1 launch, production-parity LLM evals, alignment paper — dl_weekly · 2026-09-05
- CoRL 2026 paper decisions out on OpenReview; conference heads to Austin this November — yukez · 2026-09-05