webAI ships TwIL LM3 Pro: a 3.66B formal-logic model that matches Qwen3-8B at half the size
CommonMinimum587 · reddit · 2026-10-10
webAI released TwIL LM3 Pro on September 30 — a 3.66B model built on IBM's Granite 4.2 3B, tuned purely for formal logic:
- Training recipe: LoRA SFT + checkpoint merging + RL against a programmatic verifier; the same recipe lifted VibeThinker-3B from 37.4 to 54.1.
- Results: 55.4 on their logic composite vs 43.1 for the Granite base, 41.2 for VibeThinker-3B, and 53.4 for Qwen3-8B (the card itself calls the Qwen gap sampling noise — a tie at under half the size). It clearly leads on strict multiple-choice logic (41% vs 17% for the next model) and BBH logic (95.4%).
- Weaknesses: gpt-oss-120b still wins rule induction, entailment, and Lean formalization; general benchmark average is 79.0 vs 84.9 for Qwen3-8B. It thinks long (1,900 tokens per answer) and a quarter of answers hit the length cap.
- Deployment: Q4KM GGUF is 2.09 GiB, runs on CPU or 4 GB VRAM via Ollama/LM Studio/llama.cpp; use temperature 0 and ≥2048 tokens. Official scores are BF16; Q4 is untested. License is non-commercial.
The author plans to test it as a pre-action logic checker in agent pipelines and is soliciting Q4 benchmarks.
More from Models
- Google Internally Testing Gemini 4 Checkpoint "Carbon" That Matches Opus 5.5 in Coding — koltregaskes · 2026-10-10
- OpenAI explains why Codex defaults to 300k context instead of 1M — prd_008 · 2026-10-10
- Anthropic models now outpace DeepSeek-Flash on speed, but API TTFT looks suspicious — teortaxesTex · 2026-10-10
- Qwen teases new model as dev backs sparse MoE: small experts favor inference — lxfater · 2026-10-10
- Dev burns through Grok credits in hours, eyes pricier Super Grok Heavy tier — alexcovo_eth · 2026-10-10
- repligate: Opus 3 can weave whole worlds with superintelligent subagents — repligate · 2026-10-10