Your first Apache Hop project
In this tutorial you build a small but complete Apache Hop project. You start from a file of flight records, clean it and enrich it in one pipeline, summarise it per airline in a second pipeline, and tie both together in a workflow. Then you run that workflow three ways: from Hop GUI, from the command line and in Docker.
It takes about 45 minutes. You need Apache Hop installed; see Install Apache Hop if you haven't done that yet. Docker is only needed for the last run.
1. Create a project
A project is a folder that holds everything that belongs together: pipelines, workflows and the metadata they share, such as database connections and run configurations. It is also what you put under version control.
- Start Hop GUI.
- Create the project:
- Hop 2.20 and later: click the project name next to the briefcase icon in the status bar, at the bottom left of the window, and choose Add project… → In a folder.
- Earlier versions: click the Add a new project button next to the project list in the main toolbar.
- Name it
flights, choose a new, empty folder as its home folder, and click OK.
Hop creates the folder, writes a project-config.json into it and switches to the project.
Inside the project folder, create two subfolders: data and output.
2. Add an environment
An environment holds the settings for one stage of a project's life, such as development or production, so the same project can run against different systems without changing any files.
- Create the environment:
- Hop 2.20 and later: click the environment next to the cube icon in the status bar and choose Add environment….
- Earlier versions: click the Add a new environment button next to the environment list in the main toolbar.
- Name it
flights-dev, set its purpose to Development and link it to theflightsproject. - Click OK.
This tutorial doesn't need environment variables yet, but every pipeline you run from now on
runs in flights-dev, and that is where you will add them later.
3. Add the data
Create a file data/flights.csv in the project folder with this content:
FlightDate,Operating_Airline,Origin,Dest,DepDelay,ArrDelay
2024-03-04,AA,JFK,LAX,-3,-10
2024-03-04,DL,ATL,SEA,12,25
2024-03-05,UA,ORD,SFO,0,4
2024-03-06,AA,DFW,MIA,45,52
2024-03-11,DL,ATL,BOS,-5,-12
2024-03-11,UA,EWR,DEN,8,0
2024-03-12,AA,LAX,ORD,-1,3
2024-03-13,DL,SEA,ATL,95,110
2024-03-18,UA,SFO,IAD,-2,-8
2024-03-19,AA,MIA,JFK,15,9
Each line is one flight: the date, the airline, where it left from and went to, and how many minutes it departed and arrived late. A negative delay means it was early.
4. Build the first pipeline: clean and transform
A pipeline does the data work. Its transforms all run at the same time and pass rows to each other over hops. This first pipeline adds a week number and two on-time flags to every flight.
Create a new pipeline (File → New → Pipeline) and save it in the project folder as
clean-transform.hpl. Then add four transforms. To add a transform, click on an empty part of
the canvas and search for it by name. To connect two transforms with a hop, hold Shift and
drag from one to the other.
-
Text file input, named
read flights- File tab: add
${PROJECT_HOME}/data/flights.csvto the list of files. - Content tab: set the separator to
,and keep Header checked. - Fields tab: click Get Fields. Set the type of
FlightDateto Date with formatyyyy-MM-dd, and the types ofDepDelayandArrDelayto Integer.
- File tab: add
-
Calculator, named
week of year, connected fromread flights. Add one row:- New field
WeekOfYear, calculation Week of year of date A, Field AFlightDate, value type Integer.
- New field
-
JavaScript, named
on-time flags, connected fromweek of year. Enter this script:var OnTimeDep = DepDelay <= 0 ? 'Y' : 'N';
var OnTimeArr = ArrDelay <= 0 ? 'Y' : 'N';Click Get variables so both variables become output fields, and check that their type is String.
-
Text file output, named
write transformed, connected fromon-time flags- File tab: filename
${PROJECT_HOME}/output/transformed-flights, extensioncsv. - Fields tab: click Get Fields.
- File tab: filename
Before running anything, check the result: right-click on-time flags, choose
Preview & debug output, then click Quick Launch. You should see all ten flights, each with a WeekOfYear value and Y or N
in OnTimeDep and OnTimeArr.
Save the pipeline.
Step-by-step tutorials for these transforms, with screenshots, are in the tutorials section: Text file input, Calculator and Text file output.
5. Build the second pipeline: aggregate
The second pipeline reads the transformed file and summarises it per airline.
Create a new pipeline and save it as aggregate.hpl. Add:
- Text file input, named
read transformed, reading${PROJECT_HOME}/output/transformed-flights.csv. Set the separator to,, then click Get Fields on the Fields tab. - Memory group by, named
per airline, connected fromread transformed- Group field:
Operating_Airline. - Aggregates: a field
flightsof type Number of rows (without field argument), and a fieldavg_dep_delayof type Average (Mean) onDepDelay.
- Group field:
- Add sequence, named
rank, connected fromper airline: name of valueid, start at1, increment by1. - Text file output, named
write aggregated, connected fromrank: filename${PROJECT_HOME}/output/aggregated-flights, extensioncsv, and Get Fields on the Fields tab.
Save the pipeline. You can't preview it yet: the file it reads only exists once the first pipeline has run. That is exactly what the workflow takes care of.
6. Build the workflow
A workflow orchestrates: it decides what runs, in which order, and what happens when something fails. It doesn't touch the data itself.
Create a new workflow (File → New → Workflow) and save it as flights-processing.hwf. It
starts with a Start action. Add:
- A Pipeline action named
clean and transform, with the pipeline set to${PROJECT_HOME}/clean-transform.hpl. Connect Start to it. - A second Pipeline action named
aggregate, with the pipeline set to${PROJECT_HOME}/aggregate.hpl. Connectclean and transformto it.
The hop between the two pipelines is a success hop (green): aggregate only runs when
clean and transform finished without errors. That is the behaviour you want, so leave it.
The workflow starts at Start. Its actions run one after the other.
Save the workflow.
7. Run it from Hop GUI
With the workflow open, click the run button in its toolbar (or press F8), choose the
local run configuration and click Launch.
The Execution Results panel opens at the bottom of the window. Each action gets a green
check mark when it succeeds, and the Logging tab shows what happened. Your output folder
now holds two files. aggregated-flights.csv has one line per airline, for example:
Operating_Airline,flights,avg_dep_delay,id
AA,4,14.0,1
DL,3,34.0,2
UA,3,2.0,3
The order of the airlines and the number formatting can differ, but the values should match.
8. Run it from the command line
The same workflow runs without the GUI through hop-run, which is how you run it from a script
or a scheduler. From the Hop installation folder:
- Windows
- Linux / macOS
hop-run.bat -j flights -e flights-dev -r local -f C:\path\to\flights\flights-processing.hwf
./hop-run.sh -j flights -e flights-dev -r local -f /path/to/flights/flights-processing.hwf
-j selects the project, -e the environment, -r the run configuration and -f the file.
The log appears in the terminal, and hop-run exits with code 0 when everything succeeded.
9. Run it in Docker
In production, Hop usually runs in the apache/hop container image. Mount the project folder
and tell the container what to run:
docker run -it --rm \
--env HOP_LOG_LEVEL=Basic \
--env HOP_PROJECT_FOLDER=/files \
--env HOP_PROJECT_NAME=flights \
--env HOP_FILE_PATH='${PROJECT_HOME}/flights-processing.hwf' \
--env HOP_RUN_CONFIG=local \
-v /path/to/flights:/files \
apache/hop:latest
The container registers the project, runs the workflow, prints the log and stops. The output
files appear in your project's output folder, because that folder is mounted into the
container.
10. Put the project under version control
Everything you built is a plain file in the project folder, so it fits in Git like any other code. In Hop GUI, open the File Explorer perspective (Ctrl+Shift+E). If the project folder is a Git repository, Hop shows the status of every file in colour, and you can add, commit and push from the toolbar.
Leave the output folder out of the repository: it holds results, not project files.
What you learned
- A project holds your pipelines, workflows and metadata in one folder; an environment holds the settings for one stage of its life.
- A pipeline transforms data; a workflow decides what runs, in which order, and what happens on failure.
- The same workflow runs unchanged from Hop GUI, from
hop-runand in Docker.
Next steps
- Apache Hop components: which tool to use for which task.
- Run pipelines and workflows: run on Hop Server, in containers, on Spark or Beam, and on a schedule.
- Check logs and execution information: keep a history of every run.
- Set up a database connection and swap the CSV files for tables.
- New to Apache Hop? Read What is Apache Hop: the complete guide: what it is, how it compares to other ETL tools, and how Putki runs it in production.