Critique of GeneralistAI Demo: Lack of OOD Evidence in Behavior Prompting
chris_j_paxton · x · 2026-08-21
Technical critique of the Generalist AI Gen 1.5 demo:
- Strengths: The demo is impressive, with few-shot finetuning results potentially outperforming In-Context Learning (ICL).
- Criticism: Tasks appear to be in-distribution (ID). Since the main benefit of behavior prompting is handling unseen tasks, the reviewer calls for more Out-Of-Distribution (OOD) examples, especially those hard to specify via text.
- ICL Speculation: Suggests that the emergence of ICL might relate to episode cutting strategies during data collection.
More from Models
- Developer Warns Uncensored Qwen 3.8 27B Model on Mac Immediately Explains How to Make Meth — Polymarket · 2026-08-21
- Adding Mermaid Support Becomes a Touchstone for Model Capabilities — oran_ge · 2026-08-21
- Liquid AI releases DSpark draft models, boosting inference speed by up to 3.18x — JosephJacks_ · 2026-08-21
- Opinion: Fable 5 is the only usable model in Anthropic's latest generation — bindureddy · 2026-08-21
- Anthropic's internal AECI index suggests minimal gains for next model — ChrisGPT · 2026-08-21
- Autoregressive models can beat diffusion in image generation — cloneofsimo · 2026-08-21