How pipeline run configurations work in Apache Hop
What this shows
A pipeline says what to do with data. It does not say where that happens. A run configuration is where, and keeping the two apart is what lets the same pipeline run on a laptop while you build it and on a cluster when it matters.
Reference documentationThe complete list of Pipeline run configuration options is documented in the Apache Hop manual, which is the authoritative reference.hop.apache.org →Transcript
Full transcript
A pipeline says what to do with data. It does not say where that happens. A run configuration is where, and keeping the two apart is what lets the same pipeline run on a laptop while you build it and on a cluster when it matters.
Engine type is the setting everything else follows from. Local runs the pipeline in the process that started it, which is what you want while you are working. Change this one field and the same pipeline runs somewhere else entirely, on Beam or on a remote server, with nothing else about it edited.
The settings underneath belong to the engine you picked, so they change when it does. Row set size is the buffer between one transform and the next: too small and they wait on each other, too large and you are holding rows in memory for no reason. Safe mode checks that every row has the layout it should and costs enough that you would not leave it on.
Variables set here apply to every pipeline run with this configuration. That is the seam between a pipeline and its environment: the same pipeline reads a development database or a production one depending on which configuration ran it, and nothing in the pipeline knows the difference.