Fine-tuned Qwen3-14B for finance: FinQA up to 12.6%, average score dropped
finxapp · reddit · 2026-10-10
MaxInt open-sourced two finance-adapted Qwen3-14B checkpoints with a documented trade-off:
- FinQA accuracy rose from 1.9% to 12.6%
- But the six-task average declined from 0.441 to 0.399, with substantial regressions on financial sentiment and NER
- They documented the training pipeline, data filtering, 13-gram benchmark decontamination, and a potential evaluation-format issue
The author asks the domain-adaptation community what to prioritize next: chat-template-aware re-evaluation, direct SFT without continued pretraining, or a better-balanced training mixture.
More from Research
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Claude Code agents beat torch.compile on CUDA kernels: 2.5x fused GELU, 1.57x matmul via fp32-splitting trick — lmoroney · 2026-10-10
- ResLRB: constant-memory differentiable light tracing via stochastic graph compression — ssh4net · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- SpliceAI2 deep dive: splice graphs and DP reconstruct top transcripts for 82% of held-out genes — anshulkundaje · 2026-10-10
- OpenAI publishes 722 math manuscripts covering 372 open-problem results from internal model — FinanceYF5 · 2026-10-10