Unit 1 Exam - Sample Questions

Details

In this exam, you will be asked to answer questions about pre-written code; to fill in small blanks in code; or to sketch the output of a process or analysis. Your fill-in-the-blank answers need to be conceptually correct, not perfect. For example, if you fill in a blank with mutates instead of mutate, you will still get full credit.

You may assume throughout the exam that there are no syntax errors in the code - for example, you will not be asked to spot a missing comma, an unclosed parentheses, or a typo in an object name. There are also no “trick” questions - the code you are shown can be expected to run without error, unless specifically stated in the question.

Part A - Complex Data Pipelines Sample Question

Consider the following two datasets:

A list of all Pixar films, with some information about them:

pixar_films |> head()
# A tibble: 6 × 5
  number title           release_date run_time film_rating
   <dbl> <chr>           <date>          <dbl> <chr>      
1      1 Toy Story       1995-11-22         81 G          
2      2 A Bug's Life    1998-11-25         95 G          
3      3 Toy Story 2     1999-11-24         92 G          
4      4 Monsters, Inc.  2001-11-02         92 G          
5      5 Finding Nemo    2003-05-30        100 G          
6      6 The Incredibles 2004-11-05        115 PG         

Ratings from various film rating sources for Pixar movies:

public_response |> head()
# A tibble: 6 × 5
  film            rotten_tomatoes metacritic cinema_score critics_choice
  <chr>                     <dbl>      <dbl> <chr>                 <dbl>
1 Toy Story                   100         95 A                        NA
2 A Bug's Life                 92         77 A                        NA
3 Toy Story 2                 100         88 A+                      100
4 Monsters, Inc.               96         79 A+                       92
5 Finding Nemo                 99         90 A+                       97
6 The Incredibles              97         90 A+                       88

Question 1

First, we want to add the ratings data from the imdb dataset to the Pixar movie list.

Fill in the blank to make this work.

pixar_new <- pixar_films |>
  left_join(public_response,
            by = join_by(title == film))

Question 2

Notice that the original pixar dataset has the same number of rows as the new, joined dataset:

nrow(pixar_films)
[1] 27
nrow(pixar_new)
[1] 27

Does this mean that all 27 of Pixar’s movies were present in the public_response dataset?

Why or why not?

Question 3

Fill in the blanks in the following pipeline, which computes the number of Pixar movies released each decade receiving particular cinema scores.

pixar_new |>
  mutate(
    year = lubridate::year(release_date),
    decade = floor(year/10),
    decade = glue::glue("{decade}0s")
  ) |>
  count(decade, cinema_score) |>
  pivot_wider(
    names_from = cinema_score,
    values_from = n       
  )
# A tibble: 4 × 5
  decade     A  `A+`  `A-`  `NA`
  <glue> <int> <int> <int> <int>
1 1990s      2     1    NA    NA
2 2000s      3     4    NA    NA
3 2010s      8     2     1    NA
4 2020s     NA    NA     1     5

Part B - Webscraping Sample Question

Wikipedia sites tend to contain consistent information across similar articles, usually in a box on the right. Entries about people typically include a picture of the person and their date of birth at the top. For example, here is Hadley Wickham’s Wikipedia page:

A screenshot of the Wikipedia page for Hadley Wickham. On the right side of the page, there is a gray box with a photo of Hadley, followed by a list of information about Hadley (e.g., birthday, alma matter, what he is known for, what awards he has won).

Consider the following function, designed to scrape the birthdate of a Wikipedia entry:

get_age <- function(url) {
  
  my_html <- read_html(url)
  
  my_html |>
    html_element("table") |>
    html_elements("span.noprint.ForceAgeToShow") |>
    html_text() |>
    parse_number()
  
}
get_age("https://en.wikipedia.org/wiki/Hadley_Wickham")
[1] 46

Question 4

You notice that Dr. Bodwin is not famous enough for a Wikipedia entry, so this code breaks:

get_age("https://en.wikipedia.org/wiki/Kelly_Bodwin")
Error in `open.connection()`:
! cannot open the connection

You’ll also notice that some famous people share a name, such as Robert Gentleman, one of the creators of R. This also breaks your function.

get_age("https://en.wikipedia.org/wiki/Robert_Gentleman")
numeric(0)

What would you change in the function to deal with these types of issues?

Question 5

How would you make use of this function to add a column named age to the following dataset, named r_people? Specifically, what functions would you use and how would you use them?

# A tibble: 6 × 2
  name             position      
  <chr>            <chr>         
1 Robert Gentleman R Creator     
2 Ross Ihaka       R Creator     
3 Hadley Wickham   Posit Employee
4 Jenny Bryan      Posit Employee
5 Allison Theobold Cal Poly Prof 
6 Kelly Bodwin     Cal Poly Prof 

Part C - More Advanced Visualizations Sample Question

The following outputs show a summary table or visualization from the dataset pixar_new. For each, provide:

  1. what (if any) data wrangling steps are needed to prepare the data for this visualization,

  2. the aesthetic(s) of the visualization or the gt formatting used, and

  3. one major improvement you would make to the analysis.

Note that a major improvement is something that changes the structure of the output, not just a different color, label, or stylistic choice. At least one of your (iii) answers must suggest a different geometry for the plot.

For these answers, you may respond in either words or code. For example, these two answers would both be acceptable:

  • mutate(z = x + y)
  • “Add a new column named z that is the sum of x and y.”

Question 6

  • Data Wrangling:

  • Aesthetic(s):

  • Major Improvement:

Question 7

Original Movies and Their Spinoffs
Are the spinoff ratings always worse?
Movie
Film Score
Rotten Tomatoes Metacritic
Cars 74 73
Cars 2 40 57
Cars 3 69 59
Toy Story 100 95
Toy Story 2 100 88
Toy Story 3 98 92
Toy Story 4 97 84
  • Data Wrangling:

  • Aesthetic(s):

  • Major Improvement: