Skip to main content

How to sort rows in Apache Hop

What this shows​

Sort Rows puts a stream in order. It sorts on as many fields as you name, in the order you name them, and it will spill to disk rather than give up when the rows outnumber the memory you have given it. Both of those are worth understanding before you need them.

Reference documentationThe complete list of Sort Rows options is documented in the Apache Hop manual, which is the authoritative reference.hop.apache.org →

Sample pipeline​

The pipeline in the video is transforms/sort-rows-basic.hpl, in the Putki tutorial samples project.

Transcript​

Full transcript

Sort Rows puts a stream in order. It sorts on as many fields as you name, in the order you name them, and it will spill to disk rather than give up when the rows outnumber the memory you have given it. Both of those are worth understanding before you need them.

A Data Grid supplies eight books, Sort Rows puts them in order, and a Dummy transform stands in for whatever would come next.

The table at the bottom is the sort itself. Every line is one key, and the order of the lines is the order of the comparison: rows are sorted by author first, and only rows with the same author are then separated by title. Ascending is per key rather than per sort, so you can order by one field upwards and another downwards in the same pass.

The three columns after Ascending decide what counts as order. Case sensitive compare comes off by default, so upper and lower case sort together. Sorting on the current locale is the one to reach for when the data is not English: without it, comparison is by character code, which puts accented letters after Z rather than beside the letter they belong to.

Sorting needs every row before it can emit the first one, so a large stream will not fit in memory. Sort size is how many rows are held at once; past that, a block is written to the sort directory and merged back at the end. Compressing those files trades processor time for disk. The free memory threshold is the safety net that spills early when the machine is running out, whatever the row count says.

This one is easy to misread, and the wording on the dialog tells you why. It removes duplicates, but it compares only the fields being sorted on. Two rows that agree on every key and differ everywhere else count as duplicates here, and one of them is dropped.

Running the pipeline reads all eight rows, orders them, and passes them on.

The authors now run Achebe, Butler, Le Guin, Simenon. Look at the two Achebe rows to see the second key doing its work: There Was a Country comes before Things Fall Apart, which is not the order they arrived in and is what sorting on title after author means. The three Le Guin rows are separated the same way.