One could certainly contemplate versions of RL with non-linear orderings. I guess the reason people care about linear ordering is because you want the agent to at least understand "this outcome is better than that outcome". How would we hope for a good nonlinear-RL agent to behave in an environment with 2 buttons, one of which always gives reward X, and the other of which always gives reward Y, where X and Y are incomparable?
That's understandable. But many decisions in life are like that. Do you want to be close to your roots and your parents, or do you want a high-flying career in a remote city? Choices involve sacrifices as well as gains, and many meaningful outcomes are incomparable among themselves.