Socioeconomic variables
Education, employment and income following the SEPLINE guideline
This page gives you three stand-alone recipes, one per socioeconomic dimension. Each one takes your cohort, reads one or two registers, and ends with a single table holding one category per cohort row: education_cat, occupation_cat or income_cat. Run the setup once, then any recipe on its own, in any order.
Every recipe follows the same five steps: open the register → fetch → attach to the cohort → categorise → check.
SEPLINE in brief
SEPLINE is a 2025 Danish national guideline for measuring socioeconomic position (SEP) in register-based research: which register variables to use, when to measure them, and how to categorise them, so that studies saying “adjusted for socioeconomic position” mean comparable things. Hjorth et al. Clinical Epidemiology 2025, doi:10.2147/CLEP.S520772, with a corrigendum (Hjorth et al. 2026). The table numbers on this page refer to that article.
The recipes show how to implement SEPLINE’s recommendations. They are not validated, and where SEPLINE leaves a choice open, the page says which choice it made. Review them against your own analysis plan before you use them. Have working code, or input on the categorisations? Get in touch.
What SEPLINE covers, and what this page leaves out
- Four SEP dimensions, three here. SEPLINE covers educational level, labour market affiliation, income and wealth. Wealth (Table 7:
FAMFORMREST_NY05,FGNF_2020) is not built out on this page yet. - Related factors are not SEP. SEPLINE also covers cohabitation or marital status, ethnicity and area-level characteristics, as sociodemographic factors that sit alongside SEP. For Denmark it recommends cohabitation over marital status, because many couples live together unmarried. None of these are implemented here.
- Separate indicators, not one score. SEPLINE stops at the individual indicators and says that composite indices “are not addressed in this guideline”. Expect three separate covariates at the end of this page, not one SEP score.
- The indicators are not interchangeable. SEPLINE stresses that education, labour market affiliation and income capture different things, so the choice should follow the research question.
When each dimension is measured
All three registers are annual. Taking the value from the index year itself risks measuring status after your exposure (a job lost because of the illness you study), which is reverse causation. SEPLINE sets the timing per dimension:
| Dimension | SEPLINE’s rule | This page |
|---|---|---|
| Education | “The latest available educational level at index date”; if a year is missing, take it from other years (Table 2) | the most recent record in or before the year before index |
| Employment | “The latest available status at index date”; with annual AKM, the year prior to index (Table 5 and text) | the year before index |
| Income | The mean of “the three calendar years preceding the year of index” (Table 8) | the year before index and the two years before that |
For education the year before index is not a stricter choice than SEPLINE’s: UDDA holds the status on 1 October each year (Table 1), so for an index date before 1 October the index year’s value does not yet exist at index. If your index dates fall on or after 1 October, you may include the index year (see the education recipe).
Setup (run once)
The cohort: a row id and a baseline year
Every recipe starts from cohort, one row per person and index date, with pnr and index_date. Two columns are added once and used everywhere:
library(dplyr) # mutate, left_join, group_by, summarise
library(lubridate) # year()
cohort <- cohort %>%
mutate(
row_id = row_number(), # one id per cohort row - stays unique even if a pnr repeats
year_baseline = year(index_date) - 1 # the year before index, see the timing table above
)Every join uses row_id (to attach a finished category back to the cohort) or pnr + a year (to pull a value out of a register), never pnr alone. If your cohort has one row per person, this changes nothing. If a pnr can appear more than once (for example controls matched with replacement, see Comparison cohort), row_id keeps each row’s own index date apart.
Open the registers
The recipes read the registers through one DuckDB connection instead of read_register().
Only the income recipe requires it. Education and employment filter to your cohort before collect(), so only a cohort-sized extract reaches R, and they would work just as well with read_register(), as on the guide’s other pages. They use the same connection here so that all three recipes are opened the same way.
Income needs it so R does not run out of memory and crash. Its cutpoints come from the whole Danish population: step 7 joins BEF and FAIK for everyone and computes the quantiles before anything can be reduced. Pulled into R, that is one row per person per reference year (about six million), which is more than R’s memory holds on DST: R stops with a memory error, or the whole session freezes and you lose what you had not saved. With both registers on one DuckDB connection the work happens inside the database instead, only the small table of cutpoints comes back to R, the quantiles are exact, and DuckDB writes to disk rather than filling the memory. How the pattern works: Several registers on one DuckDB connection.
If you cannot use the connection, the fallback is to compute the cutpoints in R on a random, representative sample of the population. Sample people, not rows: draw a random set of pnr first and keep all of their years, because sampling rows leaves most people with only part of their 3-year window and silently weakens the mean. Keep the sample large (for example 10% of the population), since every sex x age band only gets its share and the oldest bands are small. Cutpoints from a sample are noisier and are a deviation from SEPLINE’s “general population”, so report it.
library(DBI) # dbConnect, dbExecute
library(duckdb) # the database engine every query below runs in
con <- dbConnect(duckdb())
# Optional, and worth it for the income recipe: cap DuckDB's RAM (its default
# is 80% of the machine) and give it a folder to spill to instead of running
# out of memory. Remove the # and set a folder in your own project.
# dbExecute(con, "SET memory_limit = '8GB'")
# dbExecute(con, "SET temp_directory = 'E:/workdata/[projectnumber]/duckdb_tmp'")
# The folder that holds one subfolder per register - edit to your project
parquet_dir <- "E:/workdata/[projectnumber]/cleaned-data/parquet-registers/"
# open_register("akm") makes a view of that register's parquet files and
# returns it as a lazy table: nothing is read until a query needs it.
# ** = the register's folder and every subfolder
# hive_partitioning = true reads the year from folder names like year=2015/
open_register <- function(name) {
dbExecute(con, paste0(
"CREATE OR REPLACE VIEW ", name, " AS SELECT * FROM read_parquet('",
parquet_dir, name, "/**/*.parquet', hive_partitioning = true)"
))
tbl(con, name) %>% rename_with(tolower) # lower-case column names
}When you are done with all recipes, close the connection with dbDisconnect(con, shutdown = TRUE).
year is the folder a row came from. It comes from the parquet folder names, not from DST, so check the name in your own files (Phase 4). Never assign to it (mutate(year = ...)): every recipe joins on person plus year, and overwriting it breaks those joins without an error.
Missing is a category, not a dropped row
A register extract only holds people who have a record. Someone with no record is absent from it, and after a left join onto the cohort they have NA. Most model functions, including survival::coxph(), then drop that person silently, and your model’s N no longer matches your cohort’s.
Each recipe therefore joins onto the full cohort by row_id before it categorises, turns NA into an explicit “Missing” category, and checks with stopifnot(nrow(...) == nrow(cohort)) that nobody has been lost. SEPLINE recommends exactly this: people with missing values “may constitute a vulnerable group”, so keep them as a category and look at them separately in relation to the outcome.
Education - UDDA (hfaudd)
hfaudd is the højest fuldførte AUDD: a four-digit DISCED-15 code for the specific education a person completed (4272 is an electrician, installation technology). It is an identifier, not a scale, so the level has to be looked up. SEPLINE names DST’s own format for it (Table 3), always the newest version: AUDD..._HOVED_L1L5 for the main groups, which turns the four-digit code into a two-digit main area (10 Grundskole, 20 Gymnasiale uddannelser, … 80 Ph.d.), or AUDD..._HOVED_L1L2 for a more detailed grouping. This recipe uses the main groups. How to read those file names on the server is in Format tables.
Do not take the level from the first two digits (substr(hfaudd, 1, 2)). They look like the main-area codes but are wrong for roughly 4 codes in 10, without any error. Why is in Overview of registers.
The categories (SEPLINE Table 3):
| Category | Main areas |
|---|---|
| Short | 10 primary and lower secondary, 15 preparatory education |
| Medium | 20 general upper secondary, 30 vocational, 35 access programmes |
| Long | 40 academy profession, 50 professional bachelor, 60 bachelor, 70 master, 80 PhD |
| Missing | 90 unknown, no record, any main area not listed above, and imputed records |
SEPLINE counts 40 as long education on purpose, to match the international classification (ISCED 5, short-cycle tertiary).
# 1. Open the register
udda <- open_register("udda")
# 2. Fetch every education record up to the latest baseline year in the
# cohort. All years, not just the baseline year: step 3 needs each row's
# history to fall back on when its baseline year has no record.
udda_data <- udda %>%
semi_join(tibble(pnr = unique(cohort$pnr)), by = "pnr", copy = TRUE) %>% # only the cohort; copy = TRUE sends the pnr list into DuckDB
filter(year <= !!max(cohort$year_baseline)) %>%
select(pnr, year, hfaudd, hf_kilde) %>% # hf_kilde: where the record came from (step 5)
collect()
# 3. Attach: for each cohort row, its own most recent record in or before its
# baseline year. A completed education does not go away, so an older
# record is the right fallback when the baseline year is missing (Table 2).
cohort_udda <- cohort %>%
select(row_id, pnr, year_baseline) %>%
left_join(udda_data, by = "pnr", relationship = "many-to-many") %>% # every row x every record of that pnr, on purpose
filter(is.na(year) | year <= year_baseline) %>% # keep only records up to this row's baseline
group_by(row_id) %>%
slice_max(year, n = 1, with_ties = FALSE) %>% # this row's most recent record
ungroup()
stopifnot(nrow(cohort_udda) == nrow(cohort))Step 4: the lookup. Use DST’s format on the server, which is what SEPLINE names. The heaven tab gives the same kind of table outside DST, for testing the code. Both make edu_lookup with the columns hfaudd and main_area.
library(haven) # read_sas()
# The file names below have not been checked on the server: list the folder
# and pick the newest completed-education (audd), main-area (hoved), level
# 1 -> 5 (l1l5), character (c_), code-only (_k) file.
fmt_dir <- "E:/Formater/SAS formater i Danmarks Statistik/SAS_datasaet/Disced/"
list.files(fmt_dir, pattern = "hoved", ignore.case = TRUE)
edu_format <- read_sas(paste0(fmt_dir, "c_audd2023_hoved_l1l5_k.sas7bdat")) # (example) EDIT to the newest file
names(edu_format) # expect START (the four-digit code) and one label column
edu_lookup <- edu_format %>%
select(hfaudd = 1, main_area = 2) %>% # 1st column: the code; 2nd: its main area
mutate(hfaudd = as.character(hfaudd), main_area = as.numeric(main_area))library(heaven) # pre-installed on DST; elsewhere pak::pak("tagteam/heaven")
data(edu_code) # one row per education code
edu_lookup <- edu_code %>%
as_tibble() %>% # edu_code is a data.table
transmute(hfaudd = as.character(hfaudd), main_area = as.numeric(number)) # number = the two-digit main areaedu_code is a copy of the classification made at one point in time. SEPLINE asks for the latest format because education levels change: nurses, for example, moved from 40 to 50. Use it to test the code, and DST’s format for the analysis.
# 5. Categorise per SEPLINE Table 3
cohort_udda <- cohort_udda %>%
mutate(hfaudd = as.character(hfaudd)) %>% # the lookup is character
left_join(edu_lookup, by = "hfaudd") %>%
mutate(
# Imputed by DST from sparse data (mostly immigrants). SEPLINE groups
# these as missing for individual-level analyses (Table 2 and 3); delete
# this condition below if you want to keep them.
imputed = as.numeric(hf_kilde) %in% c(9, 10, 18),
education_cat = case_when(
imputed ~ "Missing",
main_area %in% c(10, 15) ~ "Short",
main_area %in% c(20, 30, 35) ~ "Medium",
main_area %in% c(40, 50, 60, 70, 80) ~ "Long",
TRUE ~ "Missing" # 90, no record, no match in the lookup, or a main area SEPLINE does not list
)
)
# 6. Check: how many had a code that the lookup did not know? A high number
# means a type mismatch on hfaudd or the wrong lookup file.
cohort_udda %>%
summarise(no_match = sum(!is.na(hfaudd) & is.na(main_area)), total = n())
education_ses <- cohort_udda %>% select(row_id, education_cat)
stopifnot(nrow(education_ses) == nrow(cohort))Choices SEPLINE leaves open for education
- Realskole (middle school). People whose highest education is realskole fall in main area
10and are therefore Short. SEPLINE suggests considering them Medium, depending on your population (Table 3, note c). Their codes are1021,1022,1023,1121,1122,1123,1423,1522,1523,1721,1722and1723; addhfaudd %in% realskole_codes ~ "Medium",before the Short line if you do. - Born 1920 or earlier. Education is missing or poorly recorded for most of them, and SEPLINE groups them as missing if they are not excluded.
- Under 30. Many have not finished their education. SEPLINE suggests ongoing education from KOTRE, or the parents’ education, for them (Table 2).
- Index date on or after 1 October. Then the index year’s UDDA status already exists at index, and you may use
year(index_date)instead ofyear_baselinein step 3 (and fetch one more year in step 2). - Categories. SEPLINE’s three levels plus Missing are its recommendation, and it notes that the meaning of a given length of education differs between birth cohorts.
Employment - AKM (socio13)
socio13 is DST’s code for a person’s main labour market status in the year, based on the activity that brought in the most income. The full code list is at SOCIO →; to attach DST’s labels in R, see Format tables. DST’s words are narrower than everyday ones (“beskæftiget” has thresholds attached), so look a category up at Hvad betyder → before you describe it in a methods section.
The categories (SEPLINE Table 6, with the corrigendum):
| Main group | socio13 codes |
Subgroups |
|---|---|---|
| A: Working | 110-114, 120, 131-135, 139, 310 |
self-employed (110-114, 120), managers (131, 132), employees (133-135, 139), students (310) |
| B: Unemployed | 210, 410 |
at least half the year unemployed (210), other (410) |
| C: Outside workforce | 220, 321, 330 |
sick or other leave (220), disability pension (321), social security incl. flex job (330) |
| D: Retired | 322, 323 |
age retirement pension (322), post-employment pension, efterløn (323) |
| Missing/other | 0, 420, no record, any other code |
420 is children under 15 |
Students belong to A: Working in SEPLINE, not to a group of their own. The original article also listed 410 under Missing/other; the corrigendum leaves it only under B.
# 1. Open the register
akm <- open_register("akm")
# 2. Fetch the status for the baseline years the cohort needs. unique() only
# avoids asking for the same year twice; step 3 pairs each row with its
# own year.
baseline_years <- unique(cohort$year_baseline)
akm_data <- akm %>%
semi_join(tibble(pnr = unique(cohort$pnr)), by = "pnr", copy = TRUE) %>% # only the cohort
filter(year %in% !!baseline_years) %>%
select(pnr, year, socio13) %>%
collect()
# 3. Attach: each cohort row gets the status for its own baseline year
cohort_akm <- cohort %>%
select(row_id, pnr, year_baseline) %>%
left_join(akm_data, by = c("pnr", "year_baseline" = "year"))
stopifnot(nrow(cohort_akm) == nrow(cohort))
# 4. Categorise per SEPLINE Table 6: the main group, and the subgroup for
# analyses that need it (SEPLINE notes the subgroups are not validated)
cohort_akm <- cohort_akm %>%
mutate(
occupation_cat = case_when(
socio13 %in% c(110:114, 120, 131:135, 139, 310) ~ "A: Working",
socio13 %in% c(210, 410) ~ "B: Unemployed",
socio13 %in% c(220, 321, 330) ~ "C: Outside workforce",
socio13 %in% c(322, 323) ~ "D: Retired",
TRUE ~ "Missing/other" # 0, 420 (children under 15), no record, any other code
),
occupation_sub = case_when(
socio13 %in% c(110:114, 120) ~ "A1.1 Self-employed",
socio13 %in% c(131, 132) ~ "A1.2 Manager or high-level employee",
socio13 %in% c(133:135, 139) ~ "A1.3 Employee",
socio13 == 310 ~ "A2 Student",
socio13 %in% c(210, 410) ~ "B1 Unemployed",
socio13 == 330 ~ "C1 Social security",
socio13 == 220 ~ "C2 Sick leave or other leave",
socio13 == 321 ~ "C3 Disability pension",
socio13 == 323 ~ "D1.1 Post-employment pension",
socio13 == 322 ~ "D1.2 Age retirement pension",
TRUE ~ "Missing/other"
)
)
# 5. Check: 420 means a child under 15. In an adult cohort that points to a
# wrong index date or baseline year, not to incomplete data.
cohort_akm %>% count(occupation_cat)
employment_ses <- cohort_akm %>% select(row_id, occupation_cat, occupation_sub)
stopifnot(nrow(employment_ses) == nrow(cohort))AKM has a data gap in 2014 (SEPLINE Table 4, which refers to DST’s documentation for the details). If any baseline year is 2014, compare that year’s counts with a neighbouring year before you trust it: akm %>% filter(year %in% c(2013, 2014)) %>% count(year, socio13) %>% collect().
Choices SEPLINE leaves open for employment
- Timing. SEPLINE suggests going further back than the year before index when reverse causation is a strong concern for your outcome, or measuring before retirement. DREAM gives weekly status if you need to measure closer to index, with its own caveats (Table 5).
- Coarser groups. Among people of working age, the main groups can be merged into able-bodied (A + B) versus not (C), with some misclassification because the reasons for leaving the labour market are unknown.
- Flex jobs are inside
330(C1) insocio13. IDA and DREAM classify them separately, if you need them as “working under special conditions”. - Students before 2012. In the older classification (socio02), students earning more than 60,000 DKK were counted as employees;
socio13counts them as students, so the share of students rises from 2012 (Table 4).
Check your AKM: one row per person per year
The recipe assumes one row per pnr per year. DST resolves several jobs into one main status before delivery, so that should hold. Check it:
akm_data %>% count(pnr, year) %>% filter(n > 1)If rows repeat with the same socio13, remove the copies with distinct(pnr, year, socio13) and run the check again. If the same person-year has different codes, neither SEPLINE nor DST says which to keep. Pick a documented priority (for example A > B > C > D, by attachment to the labour market) and report it as your own choice.
Income - FAIK via BEF (famaekvivadisp_13)
What SEPLINE recommends (Tables 7-9):
- Measure: equivalised disposable household income,
famaekvivadisp_13, as the default. Personal income (perindkialt_13) is an alternative when the research question needs it; this recipe implements the household measure. - Time: the mean of the three calendar years before the index year. If one year is missing, the mean of the other two; if two, the one that exists; if all three, missing.
- Standardisation: quintiles of the general population with the same sex and 5-year age group, not of your own cohort. The cutpoints are also computed per calendar year: SEPLINE’s text describes income “calculated in income quintiles by calendar year”, and incomes rise every year.
- Categories: quintile 1 = Low, quintiles 2-4 = Medium, quintile 5 = High (Table 9).
How the data fits together. FAIK holds income per family (familie_id) and year. A person’s family comes from BEF, so the recipe goes person → BEF → family → FAIK. Use the _13 version of the variable for the whole period: the older famaekvivadisp has a different definition and stops in 2012, while famaekvivadisp_13 covers 1987 onward on one definition.
It works on the whole population, inside the database. Step 7 joins BEF and FAIK for everyone in Denmark and computes the cutpoints in DuckDB; only one row per sex x age group x year comes back to R. This has been run on DST for two reference years, where it took seconds. (That it gives the same cutpoints as computing them in R has been checked on synthetic data only.)
# 1. Open the registers. FAIK is reduced to one row per family per year:
# that changes nothing when FAIK already is that, and removes repeated
# copies when it is not (see "Check your FAIK" below).
bef <- open_register("bef")
faik <- open_register("faik") %>%
distinct(familie_id, year, famaekvivadisp_13)
# 2. The years needed: each baseline year and the two years before it
income_years <- unique(c(cohort$year_baseline, cohort$year_baseline - 1, cohort$year_baseline - 2))
# 3. One row per cohort row and year in its window, so each row keeps its
# own three years even when a pnr repeats with different index dates
cohort_income_years <- bind_rows(
cohort %>% transmute(row_id, pnr, year = year_baseline),
cohort %>% transmute(row_id, pnr, year = year_baseline - 1),
cohort %>% transmute(row_id, pnr, year = year_baseline - 2)
)
# 4. Fetch the cohort's families and their income for those years
bef_cohort <- bef %>%
semi_join(tibble(pnr = unique(cohort$pnr)), by = "pnr", copy = TRUE) %>% # only the cohort
filter(year %in% !!income_years) %>%
select(pnr, year, familie_id, koen, foed_dag) %>% # koen + foed_dag for the sex x age strata in step 8
collect()
faik_cohort <- faik %>%
filter(year %in% !!income_years) %>%
collect()
# 5. Each cohort row's 3-year mean. na.rm = TRUE gives exactly SEPLINE's
# missing-year rule; n_years records how many years were used. A missing
# year is usually migration: FAIK only has income for families with a
# member living in Denmark the whole calendar year (Table 7).
income_mean <- cohort_income_years %>%
left_join(bef_cohort, by = c("pnr", "year")) %>% # this row's family in that year
left_join(faik_cohort, by = c("familie_id", "year")) %>% # that family's income
group_by(row_id) %>%
summarise(
n_years = sum(!is.na(famaekvivadisp_13)),
mean_income = mean(famaekvivadisp_13, na.rm = TRUE),
.groups = "drop"
) %>%
mutate(mean_income = if_else(n_years == 0L, NA_real_, mean_income)) # mean() of nothing is NaN, not NA
# At most 3 years per row. More means a join repeated rows: one of the
# registers is not one row per person or family per year in your delivery.
stopifnot(max(income_mean$n_years) <= 3)
# 6. Negative means (Table 8): mostly the self-employed, and often wealthy.
# SEPLINE suggests excluding them, or two sensitivity analyses with them
# counted as Low and as High. Count them, then decide - do not ignore them.
income_mean %>% summarise(n_negative = sum(mean_income < 0, na.rm = TRUE), total = n())
# 7. Cutpoints from the whole population, computed inside DuckDB.
# A reference year is a cohort baseline year; its window is that year and
# the two before it, as in step 5. Each income row is stacked three times,
# labelled with the three reference years it belongs to (the 2014 row
# counts for 2014, 2015 and 2016), so one grouped mean gives every
# person's 3-year mean for every reference year at once.
# Reduce(union_all, list(...)) does the stacking. Do not write
# union_all(a, b, c): union_all() takes two tables and silently ignores
# a third, which gave cutpoints up to 13% wrong in testing.
# Age is fixed at the reference year (the age reached that year), so a
# birthday inside the window cannot split one person across two age bands.
# SEPLINE does not say when age is measured; step 8 must use the same rule.
population <- bef %>%
filter(year %in% !!income_years) %>% # everyone, not just the cohort
select(pnr, year, familie_id, koen, foed_dag) %>%
left_join(faik %>% filter(year %in% !!income_years), by = c("familie_id", "year"))
ref_years <- tibble(ref_year = unique(cohort$year_baseline))
cutpoints <- Reduce(union_all, list(
population %>% mutate(ref_year = year),
population %>% mutate(ref_year = year + 1L),
population %>% mutate(ref_year = year + 2L)
)) %>%
semi_join(ref_years, by = "ref_year", copy = TRUE) %>% # only the reference years the cohort uses
group_by(ref_year, pnr, koen, foed_dag) %>%
summarise(mean_income = mean(famaekvivadisp_13, na.rm = TRUE), .groups = "drop") %>% # each person's 3-year mean
mutate(age_group = floor((ref_year - year(foed_dag)) / 5) * 5) %>% # 5-year bands: 0, 5, 10, ...
group_by(koen, age_group, ref_year) %>%
summarise(
q20 = quantile(mean_income, 0.2, na.rm = TRUE),
q80 = quantile(mean_income, 0.8, na.rm = TRUE),
.groups = "drop"
) %>%
collect() # only the cutpoints reach R
# 8. Classify each cohort row against the cutpoints for its own sex, age band
# and baseline year - not against the cohort's own distribution
cohort_income <- cohort %>%
select(row_id, pnr, year_baseline) %>%
left_join(bef_cohort %>% select(pnr, year, koen, foed_dag), by = c("pnr", "year_baseline" = "year")) %>%
mutate(age_group = floor((year_baseline - year(foed_dag)) / 5) * 5) %>% # same rule as step 7
left_join(income_mean, by = "row_id") %>%
left_join(cutpoints, by = c("koen", "age_group", "year_baseline" = "ref_year")) %>%
mutate(
income_cat = case_when(
is.na(mean_income) ~ "Missing", # no income in any of the three years
mean_income <= q20 ~ "Low", # quintile 1
mean_income > q80 ~ "High", # quintile 5
TRUE ~ "Medium" # quintiles 2-4
)
)
income_ses <- cohort_income %>% select(row_id, mean_income, income_cat)
stopifnot(nrow(income_ses) == nrow(cohort))Only Q20 and Q80 are needed for Low/Medium/High. If you want all five quintiles, add q40 and q60 in step 7 the same way.
Check your FAIK: one row per family per year
DST documents FAIK as one row per family per year, with no pnr. In at least one delivery (DARTER) FAIK from 2022 onward also carries pnr and repeats each family’s row once per family member. The distinct() in step 1 removes those copies, but it is only safe if every copy carries the same income. Check both, and expect 0 from each:
faik_raw <- open_register("faik") # before distinct()
# 1. Every copy of a family-year carries the same income
faik_raw %>%
group_by(familie_id, year) %>%
summarise(n_values = n_distinct(famaekvivadisp_13), .groups = "drop") %>%
filter(n_values > 1) %>%
count() %>%
collect()
# 2. After distinct(), one row per family-year is left
faik_raw %>%
distinct(familie_id, year, famaekvivadisp_13) %>%
count(familie_id, year) %>%
filter(n > 1) %>%
count() %>%
collect()Without the distinct(), each person would still be averaged on their own, but a duplicated year would count once per family member: n_years could exceed 3, and years with more copies would weigh more than the others.
Choices and caveats for income
- Match the population’s ages to your cohort’s. An age limit only changes the youngest bands, and the rule is to use the same ages on both sides. Band
15covers 15-19: if your cohort starts at 18, addfilter(ref_year - year(foed_dag) >= 18)in step 7 before the age band is computed, and the same limit in step 8. - Young adults often carry their parents’ income. DST counts children under 25 living at home as part of the family.
- Inflation. SEPLINE says income should account for inflation, and names quintiles by calendar year as one way. That is what the recipe does: the cohort and the population are compared within the same year and window, so a price index would change nothing. It matters only if you use income as a continuous value across years.
- Continuous income. SEPLINE warns that a linear term is biased and recommends splines or fractional polynomials.
- A different reference population (for example a comparator arm with your own eligibility criteria) is a legitimate choice but a deviation from SEPLINE’s general population: report it. A random sample of the population is a fallback only if you cannot use the DuckDB connection; see Open the registers for how to draw it.
- Why not DST’s published deciles? StatBank’s decile tables (IFOR25 and IFOR20-22) are not split by sex and age, and they are for one year’s income, while SEPLINE classifies a 3-year mean. Averaging pulls incomes together, so annual bounds would put fewer than 20% in Low and in High.
- 2014-2016 have a small known error. DST’s documentation of
FAMAEKVIVADISP_13says tax-free cash benefits were missing or wrong in some municipalities in those years, corrected only on DST’s datasets for income distribution and the SDG indicators (about 0.04 Gini points, or about 4,000 people in the low-income groups). - A ready-made function.
heaven::averageIncome()computes the multi-year mean if you have already linked income topnr. It does not compute population cutpoints.
Assemble all SES variables onto the cohort
Each recipe ended with one table keyed on row_id. Join the ones you ran:
cohort_ses <- cohort %>%
left_join(education_ses, by = "row_id") %>% # from the education recipe
left_join(employment_ses, by = "row_id") %>% # from the employment recipe
left_join(income_ses, by = "row_id") # from the income recipe
stopifnot(nrow(cohort_ses) == nrow(cohort))
# The Missing categories are people, not errors. Look at how many there are
# in each, and whether they differ on your outcome (SEPLINE recommends it).
cohort_ses %>% count(education_cat)
cohort_ses %>% count(occupation_cat)
cohort_ses %>% count(income_cat)See also
- SEPLINE article (Hjorth et al. 2025): the full reasoning behind every rule on this page
- Several registers on one DuckDB connection: why the recipes do not use
read_register() - Format tables: the education lookup files and how to read their names
- Overview of registers: column names for AKM, FAIK, UDDA and BEF
Next steps
You now have SES covariates. Next steps depend on what you still need:
- Specialist packages (OSDC for diabetes classification, NMI as a mortality risk score, analysis tools)? → Algorithms & special packages
- Ready to export results? → Phase 14 - Export and repatriation