Skip to main content

Apache Hop components

Apache Hop is not one program but a set of tools that share one idea: you design your data work once, Hop stores that design as files, and any of its tools can run it on any supported engine. This page explains what each part is for, how they fit together, and which one to reach for when.

The building blocks​

Every tool below works with the same few building blocks.

Building blockWhat it is
PipelineDoes the data work: reads, transforms and writes data. Stored as an .hpl file.
TransformOne step in a pipeline, such as reading a table, joining, filtering or writing a file. All transforms in a pipeline run at the same time and stream rows to each other.
WorkflowDoes the orchestration: runs pipelines and other tasks in order, and decides what happens on success or failure. Stored as an .hwf file.
ActionOne step in a workflow, such as running a pipeline, checking for a file or sending a mail. Actions run one after the other.
HopThe connection between two transforms or two actions: it carries rows in a pipeline, and decides the next step in a workflow.
MetadataDefinitions shared across a project: database connections, run configurations, servers, logging. Each is a small JSON file.
ProjectA folder holding a set of pipelines, workflows and metadata. The unit you put in Git.
EnvironmentThe settings for one stage of a project, such as development or production. It points the same project at different systems without changing its files.
Run configurationA metadata item that decides where a pipeline or workflow runs: which engine, on which machine or cluster.

The key consequence: what a pipeline does and where it runs are separate decisions. A pipeline built on a laptop runs unchanged on a server, in a container or on a Spark cluster; only the run configuration differs.

Where you design: Hop GUI and Hop Web​

Hop GUI is the desktop application where you build, test and debug pipelines and workflows. It is also where you preview data at any transform, browse the history of past runs and manage projects, environments and metadata. It runs on Windows, Linux and macOS.

Hop Web is the same Hop GUI, served from a container and used in a browser. It suits teams that want one shared development environment, or organisations where installing software on every laptop is not an option. Because people save work in it, its project files and configuration must live on persistent storage.

Both produce the same files. A pipeline designed in Hop Web opens in Hop GUI and the other way around.

What runs your work​

The engines​

The engine is what actually executes a pipeline or workflow, and the run configuration selects it.

  • The native engine is Hop's own. It runs pipelines and workflows on one machine: locally, or on a Hop Server. It handles most workloads, including production.
  • Native Spark runs pipelines on Apache Spark 4, including Databricks, for data volumes beyond one machine.
  • Apache Beam runs pipelines on Spark, Flink or Google Cloud Dataflow. It is the route for streaming, and for clusters you already operate.

Workflows always run on the native engine. They can start pipelines that run on Spark or Beam.

The ways to start a run​

ComponentWhat it isTypical use
Hop GUIRuns what you have open, with results in the Execution Results panelDevelopment and testing
hop-runCommand-line tool that runs one pipeline or workflow and exits with a status codeScripts, schedulers, CI/CD
Hop ServerLong-running service that accepts work from Hop GUI, hop-run or its REST APIA central execution server, running close to the data, exposing pipelines as web services
apache/hop imageContainer that runs one pipeline or workflow and stops, or runs Hop ServerProduction on Docker and Kubernetes; the most common way to run Hop in production
apache/hop-web imageContainer that serves Hop WebShared, browser-based development

Tools for configuration and housekeeping​

These come with every Hop installation, next to hop-run:

ToolWhat it does
hop-confCreates and manages projects and environments from the command line, sets configuration variables, and packages Hop for Spark and Beam clusters
hop setup (Hop 2.20 and later)Moves your Hop configuration out of the installation folder, so upgrades keep your projects and settings
hop-searchSearches all metadata in a project, for example every pipeline that uses a given database connection
hop-encryptEncrypts or obfuscates passwords for use in metadata and configuration files
hop-importConverts Pentaho (Kettle) jobs and transformations into Hop workflows and pipelines
hop-docGenerates documentation from your pipelines and workflows
MarketplaceAdds and removes plugins, including ones that are not bundled with Hop

How it fits together​

A typical team uses the components like this:

How the Apache Hop components fit togetherA pipeline is designed in Hop GUI or Hop Web and saved in a project in Git. Environments hold the settings for development, test and production. A scheduler starts hop-run or the apache/hop container, and the run configuration sends the work to the native engine, a Hop Server, Native Spark or Apache Beam. Execution information from every run comes back to Hop GUI.DesignHop GUI · Hop WebProject, in Gitpipelines/*.hplworkflows/*.hwfmetadata/*.jsonEnvironmentdevtestprodStart a runscheduler or CI/CDhop-runapache/hopEnginefrom the run configurationLocalnative engineHop Servernative engineSparkNative SparkBeamFlink · Spark · DataflowExecution information, back in Hop GUI

Developers design and test in Hop GUI or Hop Web. The work is saved as files in a project, in Git.

  1. Developers design and test in Hop GUI or Hop Web, in a project stored in Git.
  2. Settings that differ between development, test and production live in environments, not in the pipelines.
  3. In production, a scheduler starts hop-run or the apache/hop container, with the production environment and a run configuration.
  4. That run configuration sends the work to the native engine, a Hop Server, or Spark or Beam for large volumes.
  5. Every run records execution information, so anyone can check what ran, what failed and why, from Hop GUI.

Which one to use for which task​

You want toUse
Design, test and debug pipelines and workflows on your own machineHop GUI
Let a team design in the browser, without installing anything locallyHop Web (apache/hop-web image)
Run a pipeline or workflow from a script, a cron job or a CI/CD stephop-run
Run in production on Docker or KubernetesThe apache/hop image, one container per run
Send work to a central server, or run it close to the dataHop Server, with a native remote run configuration
Expose a pipeline as a web serviceHop Server
Process more data than one machine can handleA Native Spark or Beam run configuration
Process streaming dataA Beam run configuration, on Flink, Spark or Dataflow
Point the same project at development, test and production systemsEnvironments
Create projects and environments in a script or on a serverhop-conf
Keep your configuration when you upgrade Hophop setup, or HOP_CONFIG_FOLDER in earlier versions
Find every pipeline that uses a table or connectionhop-search
Keep passwords out of plain texthop-encrypt, or a variable resolver that reads from a secret store
Move from Pentaho Data Integrationhop-import
Document what your pipelines dohop-doc
Schedule runsAny scheduler that can start hop-run or a container: cron, Kubernetes CronJobs, Airflow, CI/CD

Where Putki fits​

Putki builds on Apache Hop and doesn't replace any of these components: your pipelines, workflows and projects stay standard Hop files, designed in Hop GUI or Hop Web and run by the same engines. Putki adds what production teams need around them: images with security patches, monitoring and central logs, governance and lineage, maintained connectors and support. See the Putki section.

Next steps​