Small LLMs Hit Only 15% Rank-Score Consistency in Financial Analysis Tests vs 75% for Frontier Models
Rough_Practice7631 · reddit · 2026-10-07
The author tested Gemma 3 27B and Qwen3 32B against Opus 5 and GPT-5.6 Sol on judging companies from real financial data:
- Rank vs score consistency: Opus 75%, GPT 65%, but Gemma only 25% and Qwen 15% when the same judgment was elicited via ranking vs 0-100 scoring.
- Order sensitivity: Simply reversing the company list changed Qwen's top pick in 6 of 10 sets.
- Analyst opinion anchoring: A bearish analyst note lowered ratings 81% of the time for Qwen and 71% for Gemma vs 43% for GPT; asking models to identify the opinion and reason independently restored original ratings in 85% (Gemma) and 68% (Qwen) of cases.
- Summary vs full text: Frontier models kept ratings 84% of the time when scoring a summary; small models changed ratings in 25-31% of cases.
Samples are small, so results should be read as a methodology demonstration rather than proof.
More from Models
- Ex-Googler suspects Project Astra voice got quantized and downgraded — joannejang · 2026-10-07
- Running quantized Swift 1.5 in 16GB VRAM at 50 tps: can lower quants help agentic coding? — royalflash417 · 2026-10-07
- OpenAI Reportedly Makes Serious Breakdown Toward Riemann Hypothesis — ipeirotis · 2026-10-07
- Jev-as-a-Judge: New Model Boosts LLM Judge Reliability for Agent Evals — omarsar0 · 2026-10-07
- X rolls out @bot tagging: reply to any post to save to Notion, set reminders via Grok — tetsuoai · 2026-10-07
- Researcher: OpenAI's math results are 'bonkers' — labs should push it to medicine next — Afinetheorem · 2026-10-07