Approved blog size (16)

Banks Score Far Lower With AI Than With People. The Sources Explain Why.

Corporate Reputation02 Sep, 2026

When a bank does something wrong, a regulator, a court or a customer writes it down. Sometimes, it's all three. When a bank does right by people for ten years, the only account of it is the one the bank writes about itself.

Though this difference has always been true, its impact on corporate reputation is growing in a world where AI is becoming an increasingly adopted information channel. Our 2026 American Banker study illustrates this point.

While our study showed banks had another strong year when it comes to their reputations with the Informed General Public, we also measured their reputations with AI as a Stakeholder. The result? A much harsher rating.

One possible explanation: People, as opposed to LLMs, form impressions from experience, from the people they know, and from a general sense of how a company has behaved. LLMs don't; all they have is the information trail that lives online.

Scores like these only matter if somebody reads the output, so it's worth noting what the survey found about usage. About a third of the people we surveyed said they're very likely to use AI tools for banking or personal finance questions, and among 18- to 24-year-olds that figure rises to 43%. The most common use they reported is comparing banks and financial products.

Where AI's Sources Change, So Do the Scores

We put ChatGPT and Gemini through the same questionnaire the Informed General Public answers, on seven US banks. Those seven scored between 55.5 and 71.5 with people. They scored between 19.4 and 52.1 with the machines. Every gap ran the same direction, from about 19 points to more than 51.

For context, banking's Reputation Score across the full study came in at 70.5 this year, down 0.4 points from 70.9. That's inside the Strong band, and it's a fair summary of what people think.

The gaps don't run evenly across the seven drivers, and what separates them isn't how much has been written about a bank. It's who wrote the good news.

At one end, AI rated banks 9.2 points higher than the public did on Performance, and 1.8 points higher on Innovation. Profitability scored 91.3, the highest single factor in the exercise. Banks publish results quarterly, analysts write them up, ratings agencies grade them, and technology budgets and product launches all generate coverage. When a bank performs well on money or technology, somebody other than the bank produces the paperwork, and much of it is positive. That also rules out the simplest alternative explanation, which is that the training data is broadly negative about banks. General negativity would pull Performance and Innovation down alongside everything else.

At the other end, Conduct came back at 26.1 against 68.0 with the public, the largest gap in the exercise. Products & Services came back at 42.0 against 71.0, a gap of 29. Those two lead the set by a wide margin.

Consider what a positive conduct record would even look like. The positive material we found the platforms drawing on was codes of conduct, purpose reports and pricing statements, all published by the banks themselves. The negative material was guilty pleas, penalty figures, consent orders and complaint volumes, all published by somebody else. It isn't that the platforms disbelieve the bank's own account. It's that almost nothing else is carrying it.

Good conduct is a claim a bank makes. Bad conduct is a finding somebody else recorded. A bank that has treated customers well for a decade has no way to demonstrate it, because the evidence was never created, so its score gets set by whatever went wrong with nothing to weigh against it.

The pattern holds across all seven drivers

reptrak driver gap chart

The middle column is our characterization, not something the study measured. And we ranked it after seeing the gaps, not before. That turns the argument into something testable rather than something to agree with: if the explanation holds, a bank that builds a real conduct record should watch that gap narrow while the other six stay roughly where they are.

The commercial stakes sit in the ordering. Products & Services and Conduct take first and second place in what the public weights most, for customers and non-customers both. Performance and Innovation, where the machines are generous, sit in the bottom half for both. The record is thinnest exactly where the audience is paying most attention.

The two platforms don't read the same sources, and they don't agree

Company-owned material accounts for about 9% of what ChatGPT reads and about 9% of what Gemini reads. Everything else — news and editorial media, regulators, complaint boards, rankings, market data, litigation records — is written by somebody other than the bank. Where those parties document good performance, a bank's positive record reaches the platform the same way its negative one does. Where they don't, the bank's own account is competing against roughly ten times its volume in outside material.

The two platforms don't read the same mix, and the difference is large enough to matter operationally. One bank in our set scored 29.8 with ChatGPT and 74.5 with Gemini in the same week. A 45-point spread on one institution is wider than the spread between most of the banks we tested. News and editorial media are about 26% of ChatGPT's sources and about 45% of Gemini's, while government and regulatory sources run the other way, about 21% for ChatGPT against about 10% for Gemini. ChatGPT also reads litigation dockets and class-action trackers, which barely appear for Gemini. So a bank with a regulatory history has a ChatGPT problem, while a bank with a complaint pattern or thin press coverage has a Gemini problem, and measuring one gives you the wrong answer about the other.

The specific sites are more useful than the categories. Both platforms lean on consumerfinance.gov and justice.gov, with ChatGPT adding sec.gov and Gemini adding occ.gov. Both lean on bbb.org and trustpilot.com, with Reddit in ChatGPT's top three and consumeraffairs.com in Gemini's. Workplace perception runs substantially through Glassdoor on both, and Products & Services runs heavily through J.D. Power.

None of that has a sense of time. In the narratives behind these scores, settlements from a decade ago appeared alongside results from the last quarter, because both are documents and neither expires.

Measuring a platform is a different job from monitoring it

Monitoring tools and answer-engine optimization work will tell you what a platform said about you and how visible you are in it. They won't tell you how much it matters. Knowing that ChatGPT raised an old settlement doesn't tell a bank whether that sits on the topic its customers rank first or sixth.

So AI as a Stakeholder runs the standard RepTrak questionnaire rather than a custom one. Overall reputation, seven drivers, 23 factors, scored the way the survey scores them. That's what lets you put the human reading and the machine reading side by side on the same drivers and see where they separate. It's also what makes every comparison in this article possible.

It reports what sits underneath the score as well: the positive and negative narratives each platform is drawing on, and the specific sources behind them, separated by platform. The source map shows which filings, articles and review sites each platform drew on for each driver, and whether the bank's own material appears among them.


Related Blog Post stories