Skip to main content

Run pipelines and workflows

Where a pipeline or workflow runs is not part of its design: a run configuration decides that when you launch it. This guide shows how to choose a run configuration, run on a remote Hop Server, in containers, on Spark or Beam, and on a schedule.

It assumes you have a project with at least one pipeline and workflow. If you don't, start with Your first Apache Hop project.

Choose a run configuration​

Run configurations are metadata items in your project. Every new project gets a local one for pipelines and for workflows. Add others in the Metadata perspective, under Pipeline Run Configuration and Workflow Run Configuration.

EngineRuns onUse it for
Native localThe machine that launches the work: your laptop, a server or a containerDevelopment, and most production work
Native remoteA Hop ServerRunning on a central server, or close to the data
Native SparkApache Spark 4, including DatabricksLarge batch workloads
Beam Spark, Beam Flink, Beam DataflowA Spark or Flink cluster, or Google Cloud Dataflow, through Apache BeamStreaming, or scaling out on a cluster you already run
Beam DirectYour own machineTesting a Beam pipeline before sending it to a cluster

Workflows run on the native engine, locally or on a Hop Server. A workflow can still start pipelines that use a Spark or Beam run configuration.

For how the options inside a run configuration work, see the pipeline run configuration and workflow run configuration tutorials.

Run from Hop GUI​

Open the pipeline or workflow, click the run button in its toolbar (or press F8), pick a run configuration and a log level, and click Launch. The Execution Results panel at the bottom shows the metrics and the log. See Check logs and execution information for what they tell you.

Run from the command line​

hop-run runs any pipeline or workflow without the GUI:

hop-run.bat -j my-project -e my-env -r local -f C:\path\to\my-project\main.hwf -p INPUT_DATE=2024-03-01
OptionSets
-j, --projectThe project
-e, --environmentThe environment, and with it the environment's variables
-r, --runconfigThe run configuration
-f, --fileThe pipeline (.hpl) or workflow (.hwf) to run
-p, --parametersParameter values, comma-separated: NAME=value,OTHER=value
-l, --levelThe log level: NOTHING, ERROR, MINIMAL, BASIC, DETAILED, DEBUG or ROWLEVEL

hop-run exits with 0 when the run succeeded and 1 when the pipeline or workflow failed, so a script or scheduler can act on the result. See hop-run for all options and exit codes.

Run on a remote Hop Server​

  1. Start a Hop Server; see Install Apache Hop.
  2. In the Metadata perspective, add a Hop Server item with the server's host name, port, username and password.
  3. Add a pipeline run configuration with the Native remote engine. Select the Hop Server, and under Run Configuration the configuration the server itself should use, usually local.
  4. If the pipeline or workflow calls other pipelines or workflows, enable Export linked resources to server? so they are sent along.
  5. Do the same for workflows with a Workflow Run Configuration.

Launch with the remote run configuration from Hop GUI or hop-run. The work runs on the server, and its status and log show up both in Hop GUI and on the server's status page.

What Hop GUI sends to a remote Hop ServerHop GUI launches the workflow flights-processing.hwf, which calls clean-transform.hpl and aggregate.hpl, with a Native remote run configuration. With Export linked resources off, only the workflow reaches the Hop Server; its first Pipeline action cannot find clean-transform.hpl unless the project is on the server at the same path, and fails. With the option on, Hop GUI sends one archive with the workflow and both pipelines, and the run succeeds. Status and log come back to Hop GUI either way.sendsstatus and logHop GUIyour projectflights-processing.hwfclean-transform.hplaggregate.hplRun configurationremote · Native remoteExport linked resourcesHop ServerReceivednothing yetRunclean and transformaggregate

In Hop GUI you launch flights-processing.hwf with a Native remote run configuration. Export linked resources is off.

Run in containers and on Kubernetes​

The apache/hop image runs one pipeline or workflow and stops, which makes it a good fit for schedulers and Kubernetes jobs:

docker run --rm \
--env HOP_PROJECT_FOLDER=/files \
--env HOP_PROJECT_NAME=my-project \
--env HOP_ENVIRONMENT_NAME=prod \
--env HOP_ENVIRONMENT_CONFIG_FILE_NAME_PATHS=/files/config/prod-config.json \
--env HOP_FILE_PATH='${PROJECT_HOME}/main.hwf' \
--env HOP_RUN_CONFIG=local \
-v /path/to/my-project:/files \
apache/hop:<version>

Two practices keep containers predictable in production:

  • Pin the image version instead of using latest, so a new Hop release never changes what runs until you decide it should.
  • Keep secrets out of the project. Put passwords and hostnames in the environment's configuration file, mounted from outside the project, or in a variable resolver that reads them from a secret store.

For the full list of variables the image reads, see Hop in Docker.

Run on Spark or Beam​

Spark and Beam run your pipelines on a cluster instead of on one machine. The pipeline design stays the same; what changes is the run configuration and how you ship the work to the cluster:

  1. Create a pipeline run configuration with the matching engine (Native Spark, Beam Spark, Beam Flink or Beam Dataflow) and fill in the cluster settings.
  2. For a cluster, package Hop and your project: hop-conf builds a fat jar with --generate-fat-jar and exports your project's metadata to one JSON file with --export-metadata.
  3. Submit the jar to the cluster, with the metadata file, the pipeline and the run configuration as arguments.

Test locally first: Native Spark runs with a local Spark master, and Beam Direct runs a Beam pipeline on your own machine. Not every transform supports every engine; each transform's page in the Apache Hop manual lists the engines it works on.

The details depend on the engine and the cluster:

Run on a schedule​

Hop has no scheduler of its own; any scheduler that can start a command or a container works. The exit code of hop-run tells the scheduler whether the run succeeded.

  • cron (Linux, macOS), for example every night at 02:00:

    0 2 * * * /opt/hop/hop-run.sh -j my-project -e prod -r local -f /data/my-project/main.hwf >> /var/log/hop/main.log 2>&1
  • Windows Task Scheduler: create a task that runs hop-run.bat with the same options.

  • Kubernetes: a CronJob that runs the apache/hop image with the variables shown above.

  • Apache Airflow: run the apache/hop image from a Docker or Kubernetes operator; see Run Hop in Apache Airflow.

  • CI/CD (Jenkins, GitLab, Forgejo or GitHub Actions): call hop-run or the container as a job step.

Run the whole process as one workflow that calls your pipelines, rather than scheduling each pipeline separately. The workflow then decides the order and what happens on failure, and the scheduler only has one thing to start and one exit code to check.