How to run a pipeline from an Apache Hop workflow
What this shows
This is the action that joins the two halves of Hop together. Pipelines move and change data; workflows decide what runs, in what order, and what happens when something fails. The Pipeline action is where one calls the other, and almost every real project is built out of it.
Reference documentationThe complete list of Pipeline options is documented in the Apache Hop manual, which is the authoritative reference.hop.apache.org →Sample workflow
The workflow in the video is workflows/pipeline-action-basic.hwf, in the Putki tutorial samples project.
Transcript
Full transcript
This is the action that joins the two halves of Hop together. Pipelines move and change data; workflows decide what runs, in what order, and what happens when something fails. The Pipeline action is where one calls the other, and almost every real project is built out of it.
Start, one pipeline, and a Success. The pipeline it runs is the Sort Rows sample from another tutorial in this series, unchanged: a pipeline does not know or care that a workflow called it.
Two settings carry most of the meaning. The first is which pipeline to run, written with a variable rather than an absolute path so the workflow still works on somebody else's machine. The second is the run configuration, which decides where it runs: locally here, but this is the field you change to run the same pipeline on a remote server or on Beam.
Wait for the pipeline to complete is on by default and should usually stay on: with it off, the workflow moves to the next action immediately and the hop that checks for success has nothing to check. Execute for every result row is the interesting one. It runs the pipeline once per row the workflow is carrying, which turns a list of files or customers into a loop without any looping construct.
Parameters are how the workflow tells the pipeline what to work on. Passing all of them down is the usual choice and is what makes a pipeline reusable: the same pipeline runs against yesterday's data or today's, depending on what it was handed.
Running the workflow hands control to the pipeline, and waits.
The workflow does not see rows. What comes back from a pipeline is a count: eight lines read, eight written, which are the eight books the Sort Rows sample sorts. That is the level a workflow works at, and it is why a scene about a workflow shows numbers where a scene about a pipeline shows a table.