Skip to main content

How to calculate new fields in Apache Hop

What this shows​

Calculator builds new fields out of the ones a row already carries. It offers a fixed list of operations rather than a formula language, so there is no expression to get wrong and nothing to compile. If the calculation you want is in the list, this is the cheapest way to do it.

Reference documentationThe complete list of Calculator options is documented in the Apache Hop manual, which is the authoritative reference.hop.apache.org →

Sample pipeline​

The pipeline in the video is transforms/calculator-basic.hpl, in the Putki tutorial samples project.

Transcript​

Full transcript

Calculator builds new fields out of the ones a row already carries. It offers a fixed list of operations rather than a formula language, so there is no expression to get wrong and nothing to compile. If the calculation you want is in the list, this is the cheapest way to do it.

A Data Grid supplies pairs of numbers, and Calculator adds four more fields to every row.

Every line in this table is one new field. The first column names it, and the second picks the operation from a fixed list. That list is long, and reading through it is usually faster than deciding you need a script.

Field A, B and C are what the operation reads. Most operations use two, some use one, and a few use all three. The order matters more than it looks. Division here reads the second number as Field A and the first as Field B, which is the reverse of the three lines above it, so it divides the second number by the first.

Value type fixes what the result is, rather than leaving it to be inferred. Remove is worth knowing about: a line marked there is worked out and then dropped from the output, which is how you use a calculation as a stepping stone to another one without shipping it downstream.

Running the pipeline works out all four fields in a single pass over the rows.

The two input numbers are still there, now followed by four fields that were not in the incoming rows. Notice the last one. Subtraction, addition and multiplication all read the first number and then the second, but division was configured the other way round. On the first row that gives nought point one, which is ten divided by one hundred, and not the ten you would get from dividing the other way.