Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric: benchmark scores, task accuracy, or one-off qualitative reviews are treated as evidence of readiness. In financial settings, this is insufficient. We take the position that financial LLM systems should not be approved for produ...
arXiv cs.CL
·Burak Payzun, \.Irem Demirta\c{s}, Simona Scala, Elena Ferretti, Se\c{c}il Arslan
·