Project Checkpoint 3

Setup

For simplicity here, I will study only the top 10 countries in terms of female medals. This code is from last week’s project step.

# Function to count medals but not overcount team sports

num_female_medals <- function(country) {
  
  female_medals <- olympics |>
    filter(team == country,
           sex == "F") |>
    distinct(games, event, medal) |>
    pull(medal)
  
  
  sum(!is.na(female_medals), na.rm = TRUE)
  
}
top_countries <- olympics |>
  distinct(team) |>
  mutate(
    n_f_medals = map_dbl(team, num_female_medals)
  ) |>
  slice_max(n_f_medals, n = 10) |>
  pull(team)

top_countries

New data from API

I’m choosing to use the API-Ninjas API for country populations. This API requires an account and API key. In the code that you see below, I have replaced hidden the line that stores my API key, so that it is not publicly shared.

get_population <- function(country, key = my_key)
  {
  
  url <- glue("https://api.api-ninjas.com/v1/population?country={country};X-Api-Key={key}")
  my_data <- safely(fromJSON)(url)
  
  if (is.null(my_data$result)) {
    return(NA)
  } else {
    return(my_data$result$historical_population)
  }
}

get_population("Japan")
get_population("United States")

Since “United States” has a space in it, it isn’t being searched properly by the API URL. After some investigation on the API site and experimentation, I see that I can use country abbreviations instead.

top_countries <- olympics |>
  distinct(team, noc) |>
  mutate(
    n_f_medals = map_dbl(team, num_female_medals)
  ) |>
  slice_max(n_f_medals, n = 10) |>
  pull(noc)

top_countries
data_list <- top_countries |>
  map(get_population)

# Check for ones that didn't work
data_list |> is.na() |> sum()

Now let’s make this a tibble so we can join it to our original data:

pop_df <- data_list |>
  set_names(top_countries) |>
  bind_rows(
    .id = "noc"
  ) 

pop_df |> head()

Now, we join the datasets:

olymp_top <- olympics |>
  filter(noc %in% top_countries) |>
  left_join(pop_df)

olymp_top |> head()

Note: There are a lot of NAs because the population data doesn’t go as far back in time as the olympic data!

Finally, we use our new information to produce an interesting summary:

olymp_top |>
  summarize(
    n_gold = sum(medal == "Gold", na.rm = TRUE),
    population = population[1],
    .by = c(noc, sex, year)
  ) |>
  mutate(
    n_gold_pp = n_gold/population
  ) |>
  filter(sex == "F") |>
  group_by(noc) |>
  summarize(
    n_gold_pp = median(n_gold_pp, na.rm = TRUE)
  ) |>
  arrange(desc(n_gold_pp))