Apache Hop components
Apache Hop is not one program but a set of tools that share one idea: you design your data work once, Hop stores that design as files, and any of its tools can run it on any supported engine. This page explains what each part is for, how they fit together, and which one to reach for when.
The building blocks
Every tool below works with the same few building blocks.
| Building block | What it is |
|---|---|
| Pipeline | Does the data work: reads, transforms and writes data. Stored as an .hpl file. |
| Transform | One step in a pipeline, such as reading a table, joining, filtering or writing a file. All transforms in a pipeline run at the same time and stream rows to each other. |
| Workflow | Does the orchestration: runs pipelines and other tasks in order, and decides what happens on success or failure. Stored as an .hwf file. |
| Action | One step in a workflow, such as running a pipeline, checking for a file or sending a mail. Actions run one after the other. |
| Hop | The connection between two transforms or two actions: it carries rows in a pipeline, and decides the next step in a workflow. |
| Metadata | Definitions shared across a project: database connections, run configurations, servers, logging. Each is a small JSON file. |
| Project | A folder holding a set of pipelines, workflows and metadata. The unit you put in Git. |
| Environment | The settings for one stage of a project, such as development or production. It points the same project at different systems without changing its files. |
| Run configuration | A metadata item that decides where a pipeline or workflow runs: which engine, on which machine or cluster. |
The key consequence: what a pipeline does and where it runs are separate decisions. A pipeline built on a laptop runs unchanged on a server, in a container or on a Spark cluster; only the run configuration differs.
Where you design: Hop GUI and Hop Web
Hop GUI is the desktop application where you build, test and debug pipelines and workflows. It is also where you preview data at any transform, browse the history of past runs and manage projects, environments and metadata. It runs on Windows, Linux and macOS.
Hop Web is the same Hop GUI, served from a container and used in a browser. It suits teams that want one shared development environment, or organisations where installing software on every laptop is not an option. Because people save work in it, its project files and configuration must live on persistent storage.
Both produce the same files. A pipeline designed in Hop Web opens in Hop GUI and the other way around.
What runs your work
The engines
The engine is what actually executes a pipeline or workflow, and the run configuration selects it.
- The native engine is Hop's own. It runs pipelines and workflows on one machine: locally, or on a Hop Server. It handles most workloads, including production.
- Native Spark runs pipelines on Apache Spark 4, including Databricks, for data volumes beyond one machine.
- Apache Beam runs pipelines on Spark, Flink or Google Cloud Dataflow. It is the route for streaming, and for clusters you already operate.
Workflows always run on the native engine. They can start pipelines that run on Spark or Beam.
The ways to start a run
| Component | What it is | Typical use |
|---|---|---|
| Hop GUI | Runs what you have open, with results in the Execution Results panel | Development and testing |
| hop-run | Command-line tool that runs one pipeline or workflow and exits with a status code | Scripts, schedulers, CI/CD |
| Hop Server | Long-running service that accepts work from Hop GUI, hop-run or its REST API | A central execution server, running close to the data, exposing pipelines as web services |
apache/hop image | Container that runs one pipeline or workflow and stops, or runs Hop Server | Production on Docker and Kubernetes; the most common way to run Hop in production |
apache/hop-web image | Container that serves Hop Web | Shared, browser-based development |
Tools for configuration and housekeeping
These come with every Hop installation, next to hop-run:
| Tool | What it does |
|---|---|
| hop-conf | Creates and manages projects and environments from the command line, sets configuration variables, and packages Hop for Spark and Beam clusters |
| hop setup (Hop 2.20 and later) | Moves your Hop configuration out of the installation folder, so upgrades keep your projects and settings |
| hop-search | Searches all metadata in a project, for example every pipeline that uses a given database connection |
| hop-encrypt | Encrypts or obfuscates passwords for use in metadata and configuration files |
| hop-import | Converts Pentaho (Kettle) jobs and transformations into Hop workflows and pipelines |
| hop-doc | Generates documentation from your pipelines and workflows |
| Marketplace | Adds and removes plugins, including ones that are not bundled with Hop |
How it fits together
A typical team uses the components like this:
Developers design and test in Hop GUI or Hop Web. The work is saved as files in a project, in Git.
- Developers design and test in Hop GUI or Hop Web, in a project stored in Git.
- Settings that differ between development, test and production live in environments, not in the pipelines.
- In production, a scheduler starts hop-run or the
apache/hopcontainer, with the production environment and a run configuration. - That run configuration sends the work to the native engine, a Hop Server, or Spark or Beam for large volumes.
- Every run records execution information, so anyone can check what ran, what failed and why, from Hop GUI.
Which one to use for which task
| You want to | Use |
|---|---|
| Design, test and debug pipelines and workflows on your own machine | Hop GUI |
| Let a team design in the browser, without installing anything locally | Hop Web (apache/hop-web image) |
| Run a pipeline or workflow from a script, a cron job or a CI/CD step | hop-run |
| Run in production on Docker or Kubernetes | The apache/hop image, one container per run |
| Send work to a central server, or run it close to the data | Hop Server, with a native remote run configuration |
| Expose a pipeline as a web service | Hop Server |
| Process more data than one machine can handle | A Native Spark or Beam run configuration |
| Process streaming data | A Beam run configuration, on Flink, Spark or Dataflow |
| Point the same project at development, test and production systems | Environments |
| Create projects and environments in a script or on a server | hop-conf |
| Keep your configuration when you upgrade Hop | hop setup, or HOP_CONFIG_FOLDER in earlier versions |
| Find every pipeline that uses a table or connection | hop-search |
| Keep passwords out of plain text | hop-encrypt, or a variable resolver that reads from a secret store |
| Move from Pentaho Data Integration | hop-import |
| Document what your pipelines do | hop-doc |
| Schedule runs | Any scheduler that can start hop-run or a container: cron, Kubernetes CronJobs, Airflow, CI/CD |
Where Putki fits
Putki builds on Apache Hop and doesn't replace any of these components: your pipelines, workflows and projects stay standard Hop files, designed in Hop GUI or Hop Web and run by the same engines. Putki adds what production teams need around them: images with security patches, monitoring and central logs, governance and lineage, maintained connectors and support. See the Putki section.
Next steps
- Install Apache Hop
- Your first Apache Hop project
- Run pipelines and workflows
- Check logs and execution information
- The Apache Hop manual: concepts and tools
- New to Apache Hop? Read What is Apache Hop: the complete guide: what it is, how it compares to other ETL tools, and how Putki runs it in production.