If a colleague had the reliability rate of an LLM, they would be fired. If ordinary* software did, it would be uninstalled.
LLMs require the chat format, because chat makes us party to the conversation. We insert the meaning ourselves - and hold it to a different standard. It's the mirror effect.
If a human told you things that were correct 80% of the time but claimed, flat out, with absolute confidence, that they were correct 100% of the time, you would dislike them & never trust a word they say. All I'm really suggesting is for people to treat chatbots with that same distrust & antagonism.