Microsoft: LLMs Are Already Jev-Style Decision Models, Fine-Tuning Isn't Always Needed
microsoft · hf · 2026-10-07
Microsoft's LLM-as-Jev framework examines whether general-purpose LLMs can serve directly as Jev-style decision models—returning calibrated categorical distributions over predefined options instead of free-form text—by reading next-token probabilities over bracketed numeric identifiers. It offers both a training-free inference recipe and a fine-tuning objective using a tree-factorized listwise loss with KL anchors to the base model.
Key findings on Qwen3.5-4B and Qwen3-0.6B:
- The 4B model, with no training, matches community Jev-style models on the same backbone, beats letter-logit readouts, supports arbitrary option counts, and natively handles image-based multimodal decisions.
- Fine-tuning delivers targeted rather than universal gains: big improvements for weaker models and specific tasks like many-option intent routing, diminishing returns for strong backbones.
- KL anchors prevent degradation in conversational generation; LoRA works best on capable models.
More from Research
- Self-improving agent Steve masters Minecraft unaided, out-progressing 97% of human players — julianweisser · 2026-10-07
- LLM2Vec-Gen: frozen LLMs generate answer embeddings in one forward pass, SOTA self-supervised — sivareddyg · 2026-10-07
- Researchers pitch World Editing: modifying existing worlds instead of generating new ones — yuntiandeng · 2026-10-07
- New paper asks: when agents act for you, whose side are they on? — ZacharyHuang12 · 2026-10-07
- AI's Top 10 research list: Spurious Rewards tops RL-heavy ranking — ShayneRedford · 2026-10-07
- SciConBench Team to Rerun Evaluations Every Two Months, Seeks Funding — manoelribeiro · 2026-10-07