Interestingly, when I run Gemma locally on my machine, it seems to be able to answer this type of question about as reliable as a language model seems to be able to do so. I wonder if the larger versions of these models that run on data centers are tokenising the prompts differently?