Run pipelines and workflows
Where a pipeline or workflow runs is not part of its design: a run configuration decides that when you launch it. This guide shows how to choose a run configuration, run on a remote Hop Server, in containers, on Spark or Beam, and on a schedule.
It assumes you have a project with at least one pipeline and workflow. If you don't, start with Your first Apache Hop project.
Choose a run configuration
Run configurations are metadata items in your project. Every new project gets a local one for
pipelines and for workflows. Add others in the Metadata perspective, under
Pipeline Run Configuration and Workflow Run Configuration.
| Engine | Runs on | Use it for |
|---|---|---|
| Native local | The machine that launches the work: your laptop, a server or a container | Development, and most production work |
| Native remote | A Hop Server | Running on a central server, or close to the data |
| Native Spark | Apache Spark 4, including Databricks | Large batch workloads |
| Beam Spark, Beam Flink, Beam Dataflow | A Spark or Flink cluster, or Google Cloud Dataflow, through Apache Beam | Streaming, or scaling out on a cluster you already run |
| Beam Direct | Your own machine | Testing a Beam pipeline before sending it to a cluster |
Workflows run on the native engine, locally or on a Hop Server. A workflow can still start pipelines that use a Spark or Beam run configuration.
For how the options inside a run configuration work, see the pipeline run configuration and workflow run configuration tutorials.
Run from Hop GUI
Open the pipeline or workflow, click the run button in its toolbar (or press F8), pick a run configuration and a log level, and click Launch. The Execution Results panel at the bottom shows the metrics and the log. See Check logs and execution information for what they tell you.
Run from the command line
hop-run runs any pipeline or workflow without the GUI:
- Windows
- Linux / macOS
hop-run.bat -j my-project -e my-env -r local -f C:\path\to\my-project\main.hwf -p INPUT_DATE=2024-03-01
./hop-run.sh -j my-project -e my-env -r local -f /path/to/my-project/main.hwf -p INPUT_DATE=2024-03-01
| Option | Sets |
|---|---|
-j, --project | The project |
-e, --environment | The environment, and with it the environment's variables |
-r, --runconfig | The run configuration |
-f, --file | The pipeline (.hpl) or workflow (.hwf) to run |
-p, --parameters | Parameter values, comma-separated: NAME=value,OTHER=value |
-l, --level | The log level: NOTHING, ERROR, MINIMAL, BASIC, DETAILED, DEBUG or ROWLEVEL |
hop-run exits with 0 when the run succeeded and 1 when the pipeline or workflow failed, so a
script or scheduler can act on the result. See hop-run
for all options and exit codes.
Run on a remote Hop Server
- Start a Hop Server; see Install Apache Hop.
- In the Metadata perspective, add a Hop Server item with the server's host name, port, username and password.
- Add a pipeline run configuration with the Native remote engine. Select the Hop Server, and
under Run Configuration the configuration the server itself should use, usually
local. - If the pipeline or workflow calls other pipelines or workflows, enable Export linked resources to server? so they are sent along.
- Do the same for workflows with a Workflow Run Configuration.
Launch with the remote run configuration from Hop GUI or hop-run. The work runs on the server,
and its status and log show up both in Hop GUI and on the server's status page.
In Hop GUI you launch flights-processing.hwf with a Native remote run configuration. Export linked resources is off.
Run in containers and on Kubernetes
The apache/hop image runs one pipeline or workflow and stops, which makes it a good fit for
schedulers and Kubernetes jobs:
docker run --rm \
--env HOP_PROJECT_FOLDER=/files \
--env HOP_PROJECT_NAME=my-project \
--env HOP_ENVIRONMENT_NAME=prod \
--env HOP_ENVIRONMENT_CONFIG_FILE_NAME_PATHS=/files/config/prod-config.json \
--env HOP_FILE_PATH='${PROJECT_HOME}/main.hwf' \
--env HOP_RUN_CONFIG=local \
-v /path/to/my-project:/files \
apache/hop:<version>
Two practices keep containers predictable in production:
- Pin the image version instead of using
latest, so a new Hop release never changes what runs until you decide it should. - Keep secrets out of the project. Put passwords and hostnames in the environment's configuration file, mounted from outside the project, or in a variable resolver that reads them from a secret store.
For the full list of variables the image reads, see Hop in Docker.
Run on Spark or Beam
Spark and Beam run your pipelines on a cluster instead of on one machine. The pipeline design stays the same; what changes is the run configuration and how you ship the work to the cluster:
- Create a pipeline run configuration with the matching engine (Native Spark, Beam Spark, Beam Flink or Beam Dataflow) and fill in the cluster settings.
- For a cluster, package Hop and your project:
hop-confbuilds a fat jar with--generate-fat-jarand exports your project's metadata to one JSON file with--export-metadata. - Submit the jar to the cluster, with the metadata file, the pipeline and the run configuration as arguments.
Test locally first: Native Spark runs with a local Spark master, and Beam Direct runs a Beam pipeline on your own machine. Not every transform supports every engine; each transform's page in the Apache Hop manual lists the engines it works on.
The details depend on the engine and the cluster:
- Getting started with the native Spark engine, including Databricks
- Getting started with Apache Beam, for Spark, Flink and Google Cloud Dataflow
Run on a schedule
Hop has no scheduler of its own; any scheduler that can start a command or a container works.
The exit code of hop-run tells the scheduler whether the run succeeded.
-
cron (Linux, macOS), for example every night at 02:00:
0 2 * * * /opt/hop/hop-run.sh -j my-project -e prod -r local -f /data/my-project/main.hwf >> /var/log/hop/main.log 2>&1 -
Windows Task Scheduler: create a task that runs
hop-run.batwith the same options. -
Kubernetes: a CronJob that runs the
apache/hopimage with the variables shown above. -
Apache Airflow: run the
apache/hopimage from a Docker or Kubernetes operator; see Run Hop in Apache Airflow. -
CI/CD (Jenkins, GitLab, Forgejo or GitHub Actions): call
hop-runor the container as a job step.
Run the whole process as one workflow that calls your pipelines, rather than scheduling each pipeline separately. The workflow then decides the order and what happens on failure, and the scheduler only has one thing to start and one exit code to check.
Related
- Apache Hop components: what each tool is for.
- Check logs and execution information: see what a run did.
- New to Apache Hop? Read What is Apache Hop: the complete guide: what it is, how it compares to other ETL tools, and how Putki runs it in production.