Skip to main content

How workflow run configurations work in Apache Hop

What this shows​

Workflows have run configurations too, and they are simpler than a pipeline's. A workflow orchestrates rather than moves data, so there are no row buffers to size and no engines to choose between beyond where it should run.

Reference documentationThe complete list of Workflow run configuration options is documented in the Apache Hop manual, which is the authoritative reference.hop.apache.org →

Transcript​

Full transcript

Workflows have run configurations too, and they are simpler than a pipeline's. A workflow orchestrates rather than moves data, so there are no row buffers to size and no engines to choose between beyond where it should run.

The name is what a workflow refers to, so it is the only part other files depend on. Engine type decides where the workflow itself runs - which is a separate question from where the pipelines it calls run, because each Pipeline action names its own configuration. A workflow can sit on one machine and drive work on another.

Transactional is the one worth understanding. With it on, every database action in the workflow shares a single transaction, so either all of them commit or none do. That turns a sequence of separate steps into one unit of work, which is what you want when a half-finished load is worse than no load at all.