A cross-section has one row per country. A panel has one row per country and year, and almost every country fact is a fact about a date: who belonged to the EU, which income group a country was in, which region the World Bank put it in. Treating those as fixed attributes of a country quietly paints today’s world over every year of the panel. This vignette walks through the verbs that keep the dates attached. Everything here runs offline.
A small panel
pan <- expand.grid(iso3c = c("GBR", "FRA", "VNM", "PAK", "AFG"),
year = 2019:2026, stringsAsFactors = FALSE)
pan$gdp <- round(1000 * (1 + 0.03)^(pan$year - 2019) *
c(GBR = 42, FRA = 39, VNM = 3.5, PAK = 1.5, AFG = 0.5)[pan$iso3c])
head(pan)
#> iso3c year gdp
#> 1 GBR 2019 42000
#> 2 FRA 2019 39000
#> 3 VNM 2019 3500
#> 4 PAK 2019 1500
#> 5 AFG 2019 500
#> 6 GBR 2020 43260Joining two panels
country_join() reconciles the country keys on both
sides. When both tables carry a year, joining on the
country alone would pair every year of one with every year of the other,
so it joins on the year too, and says so:
pop <- expand.grid(iso3c = c("United Kingdom", "France", "Viet Nam"),
year = 2019:2026, stringsAsFactors = FALSE)
pop$population <- 1e6 * c(67, 68, 98)[match(pop$iso3c, unique(pop$iso3c))]
joined <- country_join(pan, pop, iso3c, iso3c, origin_x = "iso3c")
#> Joining on year as well as iso3c.
#> ℹ Pass `also_by = character()` to join on the country alone.
nrow(joined) == nrow(pan)
#> [1] TRUEPass also_by = character() to join on the country alone
when one table really is a cross-section to broadcast, and a named
vector (also_by = c(year = "yr")) when the year columns are
named differently.
Membership as of each row
in_group() takes one date per row, so each row asks
about its own year. The United Kingdom left the EU on 31 January
2020:
uk <- pan[pan$iso3c == "GBR", ]
uk$eu <- in_group(uk$iso3c, "EU", origin = "iso3c", as_of = uk$year)
uk[, c("year", "eu")]
#> year eu
#> 1 2019 TRUE
#> 6 2020 TRUE
#> 11 2021 FALSE
#> 16 2022 FALSE
#> 21 2023 FALSE
#> 26 2024 FALSE
#> 31 2025 FALSE
#> 36 2026 FALSEA bare year means 1 January of that year, so the UK is a member on 1 January 2020 and not on 1 January 2021. Suspensions are spells of their own: Syria is not counted in the Arab League from 2011 to 2023.
Classifications as they were
The World Bank classifies every economy once a year, on 1 July, from
its gross national income two years earlier.
classify_countries() reads the dated table, so each row
gets the class in force at its date:
classify_countries(pan[pan$iso3c == "VNM" & pan$year >= 2024, ], "income")
#> iso3c year gdp income
#> 28 VNM 2024 4057 Lower middle income
#> 33 VNM 2025 4179 Lower middle income
#> 38 VNM 2026 4305 Lower middle incomeViet Nam was lower middle income until 30 June 2026. The class computed from a year’s income is published two fiscal years later: Viet Nam’s 2025 income put it in the upper middle group, in force from 1 July 2026.
classify_countries(data.frame(iso3c = "VNM", year = 2025), "income",
basis = "data_year")
#> iso3c year income
#> 1 VNM 2025 Upper middle incomeThe region that moved
On 1 July 2025 the World Bank moved Afghanistan and Pakistan from South Asia into its Middle East and North Africa region, renamed “Middle East, North Africa, Afghanistan & Pakistan”. A panel classified with today’s list puts them there in 2019 too:
reg <- classify_countries(pan[pan$iso3c %in% c("PAK", "AFG") &
pan$year %in% c(2025, 2026), ], "region")
reg[, c("iso3c", "year", "region")]
#> iso3c year region
#> 34 PAK 2025 South Asia
#> 35 AFG 2025 South Asia
#> 39 PAK 2026 Middle East, North Africa, Afghanistan & Pakistan
#> 40 AFG 2026 Middle East, North Africa, Afghanistan & PakistanLags keyed on the year
lag_by_country(), diff_by_country() and
growth_rate() look up the value a year earlier, not the
previous row. On a gapped panel the change across the gap is
NA rather than a two-year change under a one-year
label:
gappy <- pan[pan$iso3c == "FRA" & pan$year != 2022, c("iso3c", "year", "gdp")]
growth_rate(gappy, gdp)[, c("year", "gdp_growth")]
#> # A tibble: 7 × 2
#> year gdp_growth
#> <int> <dbl>
#> 1 2019 NA
#> 2 2020 0.0300
#> 3 2021 0.0300
#> 4 2023 NA
#> 5 2024 0.0300
#> 6 2025 0.0300
#> 7 2026 0.0300by = "row" keeps the old behaviour, for a panel that is
irregular by design, and warns about the gaps.
Aggregates that know what they are missing
A regional average over the countries that happen to report is a
different number from the region’s average.
aggregate_regions() and aggregate_groups()
report how much of each group is behind each value, and leave a group
out when too little of it is (min_coverage, two thirds by
default):
aggregate_groups(world_snapshot$countries, gdp_per_capita,
groups = c("EU", "ASEAN", "SADC"), fun = "weighted_mean",
weight = population)
#> # A tibble: 3 × 6
#> group gdp_per_capita n_countries n_reporting coverage coverage_weighted
#> <chr> <dbl> <int> <int> <dbl> <dbl>
#> 1 ASEAN 5105. 10 10 1 1
#> 2 EU 34927. 27 27 1 1
#> 3 SADC 1831. 16 16 1 1The coverage columns say which share of the group’s members, and of its population when a weight is given, the number rests on. The small panel above has two of the EU’s 28 members in 2019, so its EU figure is refused rather than passed off as the EU’s:
aggregate_groups(pan, gdp, groups = "EU", as_of = 2019)
#> Warning in aggregate_groups(pan, gdp, groups = "EU", as_of = 2019): 1 group falls below the 67% coverage rule, so its gdp is NA:
#> • "EU (7%)"
#> ℹ An aggregate over part of a group is not the group's figure. Pass
#> `min_coverage = 0` to compute it anyway; the coverage columns say how much
#> stands behind each row.
#> # A tibble: 1 × 5
#> group gdp n_countries n_reporting coverage
#> <chr> <dbl> <int> <int> <dbl>
#> 1 EU NA 28 2 0.0714