Spark: Sort an Rdd by Multiple Values in a Tuple / Columns
So I Have an Rdd as Follows Rdd[(String, Int, String)] and as an Example ('B', 1, 'A') ('A', 1, 'B') ('A', 0, 'B') ('A', 0, 'A') the Final Result Should Look...
So I have an RDD as follows
RDD[(String, Int, String)]
And as an example
('b', 1, 'a')
('a', 1, 'b')
('a', 0, 'b')
('a', 0, 'a')
The final result should look something like
('a', 0, 'a')
('a', 0, 'b')
('a', 1, 'b')
('b', 1, 'a')
How would I do something like this?
1 Answer
Try this:
rdd.sortBy(r => r)
If you wanted to switch the sort order around, you could do this:
rdd.sortBy(r => (r._3, r._1, r._2))
For reverse order:
rdd.sortBy(r => r, false)