I mean this should be expected, models learn from us, and "WE" game the metrics time after time. I also don't think it will just disappear just because we clean pretraining data. I believe the deeper reason is optimization, if you point any optimizer at a proxy objective it finds the cheapest path to the number, whether or not the corpus ever contained "examples of cheating."
And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.
With the exception of using divs for things that have a more appropriate element. Buttons for example - for the love of all that is holy - when I see an onclick handler and aria attributes on divs, I think, "couldn't you have restyled the button??"
And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.