Discussion on model exploratory behavior and pass@k metrics
scaling01 · x · 2026-08-21
Discussion on model performance regarding pass@k metrics, noting that OpenAI was previously seen as 'exploration maxxing' while Anthropic was 'exploitation maxxing,' but recent Anthropic models have become more exploratory. The author suggests this exploration maxxing is mostly RL-related, but pre-training builds the base for pass@k performance.
Related event: Debate Over pass@k: OpenAI's Exploration vs Anthropic's Exploitation(2 posts)→
More from Models
- Rant: GPT flags Docker container for Bluetooth audio sink as ToS violation — cargsl · 2026-08-21
- Meta Unveils Muse Spark 1.2: Vision-to-Code, Robot Navigation, Audio-Visual Understanding — AIatMeta · 2026-08-21
- Muse Spark 1.2 shows strong performance across multimodal benchmarks — alexandr_wang · 2026-08-21
- Google releases Awesome Gemma list with 16 variants and fine-tuning guides — _philschmid · 2026-08-21
- Are benchmarks useful or broken? Compass vs. Certificate — sanmikoyejo · 2026-08-21
- Google launches Gemini 3.7 Flash: Faster, cheaper, and smarter — andrew_n_carr · 2026-08-21