Change a single number in a math problem, and a human who understands the underlying logic will still get the right answer. Do the same to a large language model, and accuracy can fall off a cliff.