Skip to main content

Your first Apache Hop project

In this tutorial you build a small but complete Apache Hop project. You start from a file of flight records, clean it and enrich it in one pipeline, summarise it per airline in a second pipeline, and tie both together in a workflow. Then you run that workflow three ways: from Hop GUI, from the command line and in Docker.

It takes about 45 minutes. You need Apache Hop installed; see Install Apache Hop if you haven't done that yet. Docker is only needed for the last run.

1. Create a project​

A project is a folder that holds everything that belongs together: pipelines, workflows and the metadata they share, such as database connections and run configurations. It is also what you put under version control.

  1. Start Hop GUI.
  2. Create the project:
    • Hop 2.20 and later: click the project name next to the briefcase icon in the status bar, at the bottom left of the window, and choose Add project… → In a folder.
    • Earlier versions: click the Add a new project button next to the project list in the main toolbar.
  3. Name it flights, choose a new, empty folder as its home folder, and click OK.

Hop creates the folder, writes a project-config.json into it and switches to the project. Inside the project folder, create two subfolders: data and output.

2. Add an environment​

An environment holds the settings for one stage of a project's life, such as development or production, so the same project can run against different systems without changing any files.

  1. Create the environment:
    • Hop 2.20 and later: click the environment next to the cube icon in the status bar and choose Add environment….
    • Earlier versions: click the Add a new environment button next to the environment list in the main toolbar.
  2. Name it flights-dev, set its purpose to Development and link it to the flights project.
  3. Click OK.

This tutorial doesn't need environment variables yet, but every pipeline you run from now on runs in flights-dev, and that is where you will add them later.

3. Add the data​

Create a file data/flights.csv in the project folder with this content:

FlightDate,Operating_Airline,Origin,Dest,DepDelay,ArrDelay
2024-03-04,AA,JFK,LAX,-3,-10
2024-03-04,DL,ATL,SEA,12,25
2024-03-05,UA,ORD,SFO,0,4
2024-03-06,AA,DFW,MIA,45,52
2024-03-11,DL,ATL,BOS,-5,-12
2024-03-11,UA,EWR,DEN,8,0
2024-03-12,AA,LAX,ORD,-1,3
2024-03-13,DL,SEA,ATL,95,110
2024-03-18,UA,SFO,IAD,-2,-8
2024-03-19,AA,MIA,JFK,15,9

Each line is one flight: the date, the airline, where it left from and went to, and how many minutes it departed and arrived late. A negative delay means it was early.

4. Build the first pipeline: clean and transform​

A pipeline does the data work. Its transforms all run at the same time and pass rows to each other over hops. This first pipeline adds a week number and two on-time flags to every flight.

Create a new pipeline (File → New → Pipeline) and save it in the project folder as clean-transform.hpl. Then add four transforms. To add a transform, click on an empty part of the canvas and search for it by name. To connect two transforms with a hop, hold Shift and drag from one to the other.

  1. Text file input, named read flights

    • File tab: add ${PROJECT_HOME}/data/flights.csv to the list of files.
    • Content tab: set the separator to , and keep Header checked.
    • Fields tab: click Get Fields. Set the type of FlightDate to Date with format yyyy-MM-dd, and the types of DepDelay and ArrDelay to Integer.
  2. Calculator, named week of year, connected from read flights. Add one row:

    • New field WeekOfYear, calculation Week of year of date A, Field A FlightDate, value type Integer.
  3. JavaScript, named on-time flags, connected from week of year. Enter this script:

    var OnTimeDep = DepDelay <= 0 ? 'Y' : 'N';
    var OnTimeArr = ArrDelay <= 0 ? 'Y' : 'N';

    Click Get variables so both variables become output fields, and check that their type is String.

  4. Text file output, named write transformed, connected from on-time flags

    • File tab: filename ${PROJECT_HOME}/output/transformed-flights, extension csv.
    • Fields tab: click Get Fields.

Before running anything, check the result: right-click on-time flags, choose Preview & debug output, then click Quick Launch. You should see all ten flights, each with a WeekOfYear value and Y or N in OnTimeDep and OnTimeArr.

Save the pipeline.

tip

Step-by-step tutorials for these transforms, with screenshots, are in the tutorials section: Text file input, Calculator and Text file output.

5. Build the second pipeline: aggregate​

The second pipeline reads the transformed file and summarises it per airline.

Create a new pipeline and save it as aggregate.hpl. Add:

  1. Text file input, named read transformed, reading ${PROJECT_HOME}/output/transformed-flights.csv. Set the separator to ,, then click Get Fields on the Fields tab.
  2. Memory group by, named per airline, connected from read transformed
    • Group field: Operating_Airline.
    • Aggregates: a field flights of type Number of rows (without field argument), and a field avg_dep_delay of type Average (Mean) on DepDelay.
  3. Add sequence, named rank, connected from per airline: name of value id, start at 1, increment by 1.
  4. Text file output, named write aggregated, connected from rank: filename ${PROJECT_HOME}/output/aggregated-flights, extension csv, and Get Fields on the Fields tab.

Save the pipeline. You can't preview it yet: the file it reads only exists once the first pipeline has run. That is exactly what the workflow takes care of.

6. Build the workflow​

A workflow orchestrates: it decides what runs, in which order, and what happens when something fails. It doesn't touch the data itself.

Create a new workflow (File → New → Workflow) and save it as flights-processing.hwf. It starts with a Start action. Add:

  1. A Pipeline action named clean and transform, with the pipeline set to ${PROJECT_HOME}/clean-transform.hpl. Connect Start to it.
  2. A second Pipeline action named aggregate, with the pipeline set to ${PROJECT_HOME}/aggregate.hpl. Connect clean and transform to it.

The hop between the two pipelines is a success hop (green): aggregate only runs when clean and transform finished without errors. That is the behaviour you want, so leave it.

The flights workflow running its two pipelinesThe workflow flights-processing.hwf starts, then runs the Pipeline action clean and transform. Its pipeline reads flights.csv, all four transforms work at the same time, and it writes transformed-flights.csv. A success hop leads to the aggregate action, whose pipeline reads that file and writes aggregated-flights.csv with one line per airline. When the on-time flags transform fails instead, the first action fails, the success hop is not followed, aggregate does not run, and the workflow fails.flights-processing.hwfStartclean and transformaggregateclean-transform.hplaggregate.hpl.csv.csv.csvflights.csvtransformed-flights.csvaggregated-flights.csv10 rows

The workflow starts at Start. Its actions run one after the other.

Save the workflow.

7. Run it from Hop GUI​

With the workflow open, click the run button in its toolbar (or press F8), choose the local run configuration and click Launch.

The Execution Results panel opens at the bottom of the window. Each action gets a green check mark when it succeeds, and the Logging tab shows what happened. Your output folder now holds two files. aggregated-flights.csv has one line per airline, for example:

Operating_Airline,flights,avg_dep_delay,id
AA,4,14.0,1
DL,3,34.0,2
UA,3,2.0,3

The order of the airlines and the number formatting can differ, but the values should match.

8. Run it from the command line​

The same workflow runs without the GUI through hop-run, which is how you run it from a script or a scheduler. From the Hop installation folder:

hop-run.bat -j flights -e flights-dev -r local -f C:\path\to\flights\flights-processing.hwf

-j selects the project, -e the environment, -r the run configuration and -f the file. The log appears in the terminal, and hop-run exits with code 0 when everything succeeded.

9. Run it in Docker​

In production, Hop usually runs in the apache/hop container image. Mount the project folder and tell the container what to run:

docker run -it --rm \
--env HOP_LOG_LEVEL=Basic \
--env HOP_PROJECT_FOLDER=/files \
--env HOP_PROJECT_NAME=flights \
--env HOP_FILE_PATH='${PROJECT_HOME}/flights-processing.hwf' \
--env HOP_RUN_CONFIG=local \
-v /path/to/flights:/files \
apache/hop:latest

The container registers the project, runs the workflow, prints the log and stops. The output files appear in your project's output folder, because that folder is mounted into the container.

10. Put the project under version control​

Everything you built is a plain file in the project folder, so it fits in Git like any other code. In Hop GUI, open the File Explorer perspective (Ctrl+Shift+E). If the project folder is a Git repository, Hop shows the status of every file in colour, and you can add, commit and push from the toolbar.

Leave the output folder out of the repository: it holds results, not project files.

What you learned​

  • A project holds your pipelines, workflows and metadata in one folder; an environment holds the settings for one stage of its life.
  • A pipeline transforms data; a workflow decides what runs, in which order, and what happens on failure.
  • The same workflow runs unchanged from Hop GUI, from hop-run and in Docker.

Next steps​