Here, the same "help me" request generates two different movements because, a spray can with paint should be shaken, but a spray can with oil needs no shake: the multimodal AI gets this entirely from context (i.e., sensors, POV image from smart glasses, etc) = no custom code! 5/7