13: Text as Data
This assignment is due on Monday, November 30th
All assignments are due on D2L by 11:59pm on the due date. Late work is not accepted. You do not need to submit your .rmd file - just the properly-knitted PDF. All assignments must be properly rendered to PDF using Latex. Make sure you start your assignment sufficiently early such that you have time to address rendering issues. Come to office hours or use the course Slack if you have issues. Using an Rstudio instance on posit.cloud is always a feasible alternative. Remember, if you use any AI for coding, you must comment each line with your own interpretation of what that line of code does.
For this lab, you will need to make sure you set echo = T in the knitr options that are set in your template’s first chunk. Before turning your assignment in, check to make sure your code is showing. Labs that do not show all code will earn a zero.
Your aunt ran the consignment auction at the county fairgrounds for thirty years and kept every record in one spreadsheet. She has retired to Florida and stopped answering email, and you have inherited the file. Every column in it is text, because at some point somebody typed a dollar sign into a number column and Excel quietly gave up. Below are four of those columns. You could fix these four rows by hand in about two minutes. The file has forty thousand.
This lab is deceptively short. Regular expressions can take a while to learn, and can be very frustrating. Leave yourself ample time.
In this lab, I set limits on lines of code you can use in order to force you to do this efficiently (and not, say, hand-code the results). The data entry lines don’t count as lines of code, nor does the start of the pipe e.g. df = c(1,2,3) is not a line, nor is df %>%.
- Column A, the sales figures for 2020. Create the following
Sales2020object. We want to convert the sales into a numeric format, but have to wrestle with the extra characters. Using no more than two lines of code, convert these to numeric.
Sales2020 = c('$1,420,142',
'$438,125.82',
'120,223.50',
'42,140')
- Column B, who to call about each lot. We have the following names:
Students = c('Ali',' Meza','McEvoy', 'Mc Evoy', '.Donaldson','Kirkpatrick ')
We would like to clean these student names and place them in alphabetical order. However, we note that “Mc Evoy” is a different name from “McEvoy” (the Irish are particular about that), so we need no more than two lines of code that can print these names in order (use sort to place the cleaned names in order).
- Column C, the price column, as typed. We want to check if each of these are valid, complete (including two decimal places for cents) prices. Write a line of code that returns TRUE if the price is in US dollars and includes cents to two digits.
Prices = c('$12.95',
'$\beta$',
'$1944.55',
'3.14',
'$CAN',
'$12.',
'$109',
'4,05',
'$200.00')
- Column D, which she pasted in from a land listings site and never touched again. We will use groups to extract two columns of information from it:
landSales = c('Sold 12 acres at $105 per acre',
'200 ac. at $58.90 each',
'.25 acre lot for $1,000.00 ea',
'Offered 50 acres for $5,000 per')
We want to calculate the total amount spent on each of these transactions (which are all in “per acre” prices). In no more than 5 lines of code (each pipe is a separate line), extract the acres offered, the price per acre, and then calculate the total transaction (acres x price per acre). You can print the data.frame / tibble, or use %>% knitr::kable() to output the table. Your final output should have four rows and three columns. Hint: use str_match with groups.
Then, in one sentence: the per-acre prices in this column differ by a factor of about eighty-five. Is that four different kinds of land, or four different kinds of typo? Which one would you go back to your aunt about, if she were answering email? That sentence is prose, not code -- it doesn't count against your five lines.