Check logs and execution information
Apache Hop tells you what a run did in three places: the log and metrics of the run you just started, a browsable history of past runs, and, if you set it up, run details written to a database of your choice. This guide covers each of them.
Check the run you just started
When you run a pipeline or workflow from Hop GUI, the Execution Results panel opens at the bottom of the window:
- The Metrics tab shows one line per transform, with the rows it read, wrote and rejected, and the time it took. It is the fastest way to see where rows disappeared or where time went.
- The Logging tab shows the log of the run. When something fails, the reason is here.
- On the canvas, transforms and actions that finished get a green check mark; the one that failed is marked red.
To look at the data itself, right-click a transform, choose Preview & debug output and click Quick Launch: Hop runs just enough of the pipeline to show the rows at that point. See Run, preview and debug a pipeline for breakpoints and conditional debugging.
Choose a log level
You pick the log level when you launch a run: in the run dialog in Hop GUI, with -l or
--level for hop-run, and with HOP_LOG_LEVEL in the apache/hop container.
| Level | Use it for |
|---|---|
Nothing, Error | Only when output must stay minimal; you lose the context of what happened before an error |
Minimal | Production runs where you only need start, end and errors |
Basic | The default. Production runs, and most development |
Detailed | Finding out what a transform or action is doing |
Debug | Troubleshooting a specific problem |
Rowlevel | Seeing individual rows pass through. Very verbose: use it on small test data only |
Logs from the command line and containers
hop-run writes its log to the terminal. When a scheduler runs it, redirect the output to a file
or let the scheduler capture it:
./hop-run.sh -j my-project -e prod -r local -f /data/my-project/main.hwf -l BASIC >> /var/log/hop/main.log 2>&1
The apache/hop container writes its log to standard output, so it ends up wherever your platform
collects container logs: docker logs, your Kubernetes logging stack or your CI job output.
HOP_LOG_PATH sets where the container also writes its log file.
Hop GUI keeps its own log in the audit folder of your Hop configuration (or the folder set in
HOP_AUDIT_FOLDER).
Keep a history of every run
The logs above are gone once you close the run or the terminal. To look back at yesterday's run, Hop records execution information: the status, log, metrics and optionally a sample of the data of every run, for the workflow and every pipeline it started.
- In the Metadata perspective, add an Execution Information Location. The simplest type is File location, which writes to a folder; other types store executions in Elasticsearch, Neo4j, a database or on a Hop Server. Choose the type that everyone who needs the history can reach.
- Open your pipeline and workflow run configurations and select this location under Execution information location.
- Optionally, add an Execution Data Profile and select it in the pipeline run configuration, to keep samples of the rows each transform produced: the first rows, the last rows or a random sample.
Every run with those run configurations is now recorded. Open the Execution Information perspective in Hop GUI to browse the history: start from a workflow run, drill down into the pipelines it started, and jump straight to the transform that failed.
Because the location can be a shared folder, cloud storage or a Hop Server, runs on a server, in a container or on a Spark cluster can be inspected from your own Hop GUI.
Pipelines and workflows run on a Hop Server, in containers or on a Spark cluster - not on your machine.
See Execution Information Location for every location type and its options.
Write run details to a database
To report on runs, or to raise an alert when one fails, send the logging to a pipeline of your own:
- Create a pipeline that starts with the Pipeline logging transform. What it does next is up to you: write to a database table, a file, a message queue or a chat channel.
- In the Metadata perspective, add a Pipeline Log item. Point it at that logging pipeline, and select the pipelines whose activity it should log.
- For workflows, do the same with a Workflow Log item and a pipeline that starts with the Workflow logging transform.
The samples project that ships with Apache Hop contains a working example, pipeline-log-example.
See Pipeline Log and
Workflow Log.
Monitor many pipelines at once
Everything above works per run or per pipeline. Once you run many pipelines in production, you need one place that shows all of them: which ran, which failed, how long they took and how that is trending. You can build that on the database from the previous section with any reporting tool. Putki provides it ready-made, with dashboards, central searchable logs and chat alerts; see Putki.
Related
- Run pipelines and workflows: run configurations, remote servers, containers and scheduling.
- Apache Hop components: what each tool is for.
- New to Apache Hop? Read What is Apache Hop: the complete guide: what it is, how it compares to other ETL tools, and how Putki runs it in production.