bank <- read.table("https://ec242.netlify.app/data/bank.csv",
header = TRUE,
sep = ",") %>%
dplyr::select(-default)14: Applied Logistic Regression - Classification
This assignment is due on Monday, December 7th
All assignments are due on D2L by 11:59pm on the due date. Late work is not accepted. You do not need to submit your .rmd file - just the properly-knitted PDF. All assignments must be properly rendered to PDF using Latex. Make sure you start your assignment sufficiently early such that you have time to address rendering issues. Come to office hours or use the course Slack if you have issues. Using an Rstudio instance on posit.cloud is always a feasible alternative. Remember, if you use any AI for coding, you must comment each line with your own interpretation of what that line of code does.
Backstory and Set Up
You work in marketing analytics at a Portuguese bank. The call center phoned 4,521 customers to pitch a term deposit and talked 521 of them into it. Management has looked at that ratio and concluded, as management does, that the problem is analytics. You get the call logs and the job of working out who is worth calling.
The variable y records whether the customer opened the account. That’s your target. There is also a column called default – whether the customer is already in default on some other credit. That’s a predictor, not the thing you’re predicting, and with only 76 of the 4,521 in default it does approximately nothing, so the snippet below drops it.
This is some new data. The snippet below loads it.
There’s not going to be a whole lot of wind-up here. You should be well-versed in doing these sorts of things by now (if not, look back at the previous lab for sample code).
EXERCISE 1 of 1
Encode the outcome we’re trying to predict (
y) as a binary 0/1. (Thedefaultcolumn is already gone – the snippet above dropped it.)Check your data for any
NAs (as should be customary).Split the data into an 80/20 train vs. test split. Make sure you explicitly set the seed for replicability.
Before you fit anything, work out how accurate you would be if you skipped the modeling entirely and predicted “no” for every customer in your test set. It’s one line. Write the number down – it should land around 88%. That is the number every model below has to beat, and I want it reported next to every accuracy you produce.
Run a series of logistic regressions with between 1 and 4 predictors of your choice (you can use interactions).
Create eight total confusion matrices: four by applying your models to the training data, and four by applying your models to the test data. Briefly discuss your findings. How does the error rate, sensitivity, and specificity change as the number of predictors increases?
A few hints:
- If you are not getting a 2x2 confusion matrix, you might need to adjust your cutoff probability.
- It might be the case that your model perfectly predicts the outcome variable when the setup cutoff probability is too high.
- You need to make sure your predictions take the same possible values as the
actualdata (which, remember, you had to convert to a binary 0/1 variable)
Now make a decision with it. Your call center can make 200 calls next week. Not 201. Take your best model, rank everyone in the test set by predicted probability, call the top 200, and report how many of them actually opened a term deposit. Then report how many you’d have landed by picking 200 people at random. One sentence, two numbers.
Two rules. You may not use
duration– it’s the length of the call, and you’re deciding whom to call. You do not know how long a call lasted until you’ve made it. Put it in and your model will look fantastic and mean nothing. Second, make sure at least one predictor is numeric, or hundreds of customers tie at the same probability and no cutoff gives you exactly 200.Finally: what cutoff probability does “call the top 200” work out to, and why isn’t it 0.5?