When I worked on vision DNNs (mostly CNNs with some LSTM), when we interrogated the layers of trained classifiers you could see that the models "learned" how to do edge detection, sharpening, and other image enhancement tasks. That LLMs would learn a screwed up way to do math doesn't shock me.