Study finds AI models struggle with complex financial tasks
A study by Saturn involving 18 AI models, including ChatGPT, Claude, Copilot, Grok, and Gemini, revealed high error rates in financial queries. While models averaged a 57% error rate across general queries, accuracy dropped significantly on complex multi-step calculations, with error rates reaching 88% and up to 99% for some models. Claude Opus 5 performed best with a 39% error rate, while some AI advice could have resulted in significant tax penalties.
Summaries are written by AI from the original article. Not investment advice.