We are looking to improve our service on Prep Air so want to calculate NPS from a survey and see how we compare.
Step 1 - Combine Data
First we want to combine the data from both of the files. We could use the union tool for this, but I have used the Wildcard Union in the input tool. You don't need to add any matching patterns as we want to bring all of the files through:
We should now have both of our inputs in a single table:
Step 2 - Classify Customers
Next we want to classify the responses that we want to compare, so the first step is to count the total number of customers for each Airline. To calculate this we can use a fixed LOD:
Number of Customers
After counting the customers, we can then filter to only airlines with more than 50 customers. You can filter this by using a range filter on Number of Customers:
Finally, to classify the customers responses we can use the following IF statement:
Classification
IF [How likely are you to recommend this airline?]<7
THEN "Detractor"
ELSEIF[How likely are you to recommend this airline?]<9
THEN "Passive"
ELSE "Promoter"
END
After the classification our data should look like this:
Step 3 - Calculate NPS
We can now turn to the main part of the challenge, calculating the NPS for each Airline.
First we need to count the number of customers for each classification and airline by using an aggregate tool:
After the aggregation, we only want to focus on Detractor and Promoter so we can exclude the Passive classifications. Also, to make things easier we can rename Number of Customers to Total Customers and Customer ID to Number of Customers.
Now we can calculate the % Total in each Airline & Classification using this calculation:
% Total
100*[Number of Customers] / [Total Customers]
Then make this a whole number and remove the Total Customer and Number of Customers fields.
Finally we are ready to pivot our table so that we have the Detractor and Promoter scores on a single row. The rows to columns pivot setup looks like this:
Now we are ready to calculate the NPS:
NPS [Promoter]-[Detractor]
Our table looks like this:
Step 4 - Calculate Z-Score
The last step is to calculate the Z-Score for each airline. We calculate this with the following calculation
[NPS]-[Average] / [Standard Deviation]
Before we get to this stage we first need to calculate the overall Average and Standard Dev across the data set. We can use a fixed LOD for this:
Average
Standard Dev
Now we're ready to calculate the Z-Score:
Z-Score
ROUND(
([NPS]-[Average])
/
[Standard Deviation]
,2)
Our data should now look like this:
Then finally we can filter for just Prep Air (filter for selected values) so our output will look like this:
You can also post your solution on the Tableau Forum where we have a Preppin' Data community page. Post your solutions and ask questions if you need any help!
Challenge by: Jenny Martin As I've mentioned before in a previous challenge, I'm a big fan of a quiz show called Richard Osman's House of Games. However, I've often found the way that they decide the overall winner of the week a little troubling. Each day the player who has scored the most, will receive 4 points, 2nd place will receive 3 points, 3rd place will receive 2 points and last place will receive one point. These points will be added up across the week to determine the overall winner, but with a twist! Each Friday double points are awarded so 1st place receives 8 points and so on. This leads me to wondering: Would there be a different winner if there was no double points Friday? What about if participants weren't ranked at the end of each day and they had a running total score across the week instead, would that lead to a different winner? What about doubling the scores on the Friday, instead of the points awarded? Input Luckily I didn't have to collect ...
Free isn't always a good thing. In data, Free text is the example to state when proving that statements correct. However, lots of benefit can be gained from understanding data that has been entered in Free Text fields. What do we mean by Free Text? Free Text is the string based data that comes from allowing people to type answers in to systems and forms. The resulting data is normally stored within one column, with one answer per cell. As Free Text means the answer could be anything, this is what you get - absolutely anything. From expletives to slang, the words you will find in the data may be a challenge to interpret but the text is the closest way to collect the voice of your customer / employee. The Free Text field is likely to contain long, rambling sentences that can simply be analysed. If you count these fields, you are likely to have one of each entry each. Therefore, simply counting the entries will not provide anything meaningful to your analysis. The value is in ...
Challenge by: Robbin Vernooij Recently, one of the Data School Coaches, Robbin, set the following challenge. It seemed perfect for a Preppin' Data, so over to Robbin: We'd like to get historical data on the highest paid athletes so we can do temporal analysis. Lucky us, it turns out Wikipedia has been tracking the Forbes list of the world's highest-paid athletes. Unlucky us, it is in an HTML table format with human readable symbols and table by table basis. Now it's time for you to clean it up into one single dataset, so that it's ready for analysis. Inputs The data for this challenge comes from this Wikipedia page . There is a table for each year that looks like this (2024 example): As well as a source table: Requirements Input the data Bring all the year tables together into a single table Merge any mismatched fields (there should not be any Null values) Create a numeric Year field Clean up the fields with the monetary amounts One way of doing this could ...