Small LLMs Hit Only 15% Rank-Score Consistency in Financial Analysis Tests vs 75% for Frontier Models

Rough_Practice7631 · reddit · 2026-10-07

The author tested Gemma 3 27B and Qwen3 32B against Opus 5 and GPT-5.6 Sol on judging companies from real financial data:

Samples are small, so results should be read as a methodology demonstration rather than proof.

Original post →

More from Models

Models channel →