Salesforce's 31B open Gemma web agent scores 74.6% on WebArena, beating Gemini 3 Flash
dair_ai · x · 2026-10-08
Salesforce AI Research's new CLIFT paper shows a 31B open Gemma web agent scoring 74.6% on the 9-app WebArena Infinity set, above Gemini 3 Flash with browser use at 70.1% — without calling a frontier judge at every step or at deployment.
How it works:
- The agent answers verification questions about its own rollouts (CLIFT);
- A conformal certifier keeps only questions whose answers agree with a training-time judge, weights them by trustworthiness, and folds them into per-step rewards;
- At test time, the same frozen question bank picks between a greedy rollout and a few retries, with no external judge.
Results: +12.8 points over the base model, winning 7 of 9 apps. The question bank also transfers to GPT-5.5 at test time on VisualWebArena, and a translated bank improves a live-web agent on Online Mind2Web with no training.
More from Models
- Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD — HankYeomans · 2026-10-08
- Cascade model beats Toto 2.0 4M on TIME MASE/CRPS with 15.4% of the training tokens — const_reborn · 2026-10-08
- LiquidAI's d1-3B Edge Image-Text-to-Text Model Trends on Hugging Face — LiquidAI · 2026-10-08
- Burkov: frontier models can't vibe-code Photoshop; video data mostly worthless for visual reasoning — burkov · 2026-10-08
- Anthropic's new model priced below DeepSeek and GLM flash, matching them on Terminal Bench 4.0 — op7418 · 2026-10-08
- Reddit challenge: can any LLM write a complete Risch algorithm implementation? — big_hole_energy · 2026-10-08