
3: Applying ggplot2 to Real Data
This assignment is due on Monday, September 21st
All assignments are due on D2L by 11:59pm on the due date. Late work is not accepted. You do not need to submit your .rmd file - just the properly-knitted PDF. All assignments must be properly rendered to PDF using Latex. Make sure you start your assignment sufficiently early such that you have time to address rendering issues. Come to office hours or use the course Slack if you have issues. Using an Rstudio instance on posit.cloud is always a feasible alternative. Remember, if you use any AI for coding, you must comment each line with your own interpretation of what that line of code does.
Preliminaries
As always, we will first have to load ggplot2. To do this, we will load the tidyverse by running this code:
library(tidyverse)
Background
When New York City shut down in spring 2020, construction stopped with it — unless the Department of Buildings decided your site was “essential.” Somebody in an office made that call one site at a time, and because this is a public agency, the list of what they decided is now a map and a CSV you can download.
You work in that office. So does the deputy commissioner, who has just been told a reporter has the same file you do and is running a piece Friday under the headline “Manhattan Never Stopped Building.” She wants to know whether the reporter has a point. She wants it on one page, in plots she can look at for four seconds and understand, because that is roughly how long she will look at them.
As you hopefully figured out by now, you’ll be doing all your R work in R Markdown. You can use an RStudio Project to keep your files well organized (either on your computer or on RStudio.cloud), but this is optional. If you decide to do so, either create a new project for this exercise only, or make a project for all your work in this class.
You’ll need to download one CSV file and put it somewhere on your computer (or upload it to RStudio.cloud if you’ve gone that direction)—preferably in a folder named data in your project folder. You can download the data from the DOB’s map, or use this link to get it directly:
R Markdown
Writing regular text with R Markdown follows the rules of Markdown. You can make lists; different-size headers, etc. This should be relatively straightfoward. We talked about a few Markdown features like bold and italics in class. See this resource for more formatting.
You’ll also need to insert your own code chunks where needed. Rather than typing them by hand (that’s tedious and you might miscount the number of backticks!), use the “Insert” button at the top of the editing window, or type ctrl + alt + i on Windows, or ⌘ + ⌥ + i on macOS.
Data Prep
Once you download the EssentialConstruction.csv file and save it in your project folder, you can open it and start cleaning. Loading in the basic data is straightforward:
library(tidyverse)
essential = read_csv('pathTo/EssentialConstruction.csv')Where the “pathTo” part is the path to your local folder. If you saved the data in the same folder as your template, then you can just use:
essential = read_csv('EssentialConstruction.csv')Once loaded, note that each row is an approved project (the JOB NUMBERS are approved projects, so each row is one approved project).
Two things to fix before you plot anything.
First: one of our columns has inconsistent capitalization — BRONX and Bronx are two different strings as far as R is concerned, and every borough in this file shows up both ways. Use case_when (or any other method) to make them consistent.
Second, the denominator. Counting projects tells you where the most projects were. It doesn’t tell you where the most projects were per person. Tuesday’s murder plot was this same trick, and I made you sit through it for a reason. Here are the 2020 Census populations. Type them in, it’s five rows.
boro_pop <- tribble(
~BOROUGH, ~population,
"Bronx", 1472654,
"Brooklyn", 2736074,
"Manhattan", 1694251,
"Queens", 2405464,
"Staten Island", 495747
)left_join that on. And note what happens if you skipped the capitalization fix: about a third of your rows come back with a population of NA. The bad news is that your plot still draws. It just draws the wrong answer, and nothing tells you.
Each row is an approved construction project.
A. Show approved projects per 100,000 residents by borough using a bar chart, ordered from most to least — use reorder(), because alphabetical order tells the reader nothing. Make sure all the elements of your plot (axes, legend, etc.) are labeled.
One warning: `geom_bar` counts rows, and you are no longer counting rows. Look up what `stat = "identity"` does, or use `geom_col`.
B. Show the count or proportion of approved projects by category using a lollipop chart. Not sure of what a lollipop chart is? Google R ggplot lollipop. A huge portion of knowing how to code is knowing how to google, find examples, and figure out where to put your variables from your data! Make sure all the elements of your plot (axes, legend, etc.) are labeled. Make sure ticks are well-placed, numbers are properly formatted, etc. Everything should look professional. Grading will be very picky on this!
Now use an appropriate facet function (wrap or grid) to re-make your borough bar chart from Exercise 1A across each CATEGORY. Again, make sure everything is labeled well and properly formatted.
Then answer the deputy commissioner in three sentences or fewer: does the reporter have a point? Point at the specific plot that makes your case.
One warning — Approved Work is more than half the rows and it is a residual bucket. It means “approved,” not a kind of building. Don’t build an argument on it without saying so.
Getting help
Use the EC242 Slack if you get stuck (click the Slack logo at the top right of this website header).
Turning everything in
When you’re all done, click on the “Knit” button at the top of the editing window and create a PDF. Upload the PDF file to D2L.