500 investment bankers review AI outputs and find none ready for client delivery

TL;DR AI
2 min readKey summary
Handshake AI and McGill University launched BankerToolBench, a benchmark built from real junior investment banking workflows.
About 500 bankers helped design and grade 100 tasks spanning Excel models, PowerPoint decks, SEC filings, and tool use.
Nine leading AI models were tested; none produced work judged ready for client delivery.
GPT-5.4 performed best, but still needed substantial human rework, while some other models failed outright.
The findings suggest frontier AI is not yet reliable enough to replace or directly support high-stakes banking analysis without heavy review.



