shkspr.mobi
Are LLMs still surprisingly bad at some simple tasks?
Last year I ran an experiment to test the ability of modern LLMs to correctly answer a relatively straightforward question. Every single one of them got it wrong. Some missed information, some made up false statements, none were right. Of course the fanbois variously claimed that I was holding it wrong, my prompts were shit, I should have chosen better defaults, and - my favourite - that it…