Construct either a full (single-year age intervals) or an abridged (wider age intervals) life table from a variety of input data types. The function accepts:
Death counts and mid-interval population estimates (
Dx, Ex)Age-specific death rates (
mx)Death probabilities (
qx)Survivorship curve (
lx)Distribution of deaths (
dx)Remaining life expectancy (
ex)
Exactly one of these input options must be provided; supplying more
than one is an error. The input can be a numeric vector,
matrix, or data.frame. When a matrix or
data.frame with multiple columns is supplied, the function
computes one life table per column.
Usage
LifeTable(x, Dx = NULL, Ex = NULL,
mx = NULL,
qx = NULL,
lx = NULL,
dx = NULL,
ex = NULL,
sex = NULL,
lx0 = 1e5,
ax = "andreev_kingkade",
close = NULL,
omega = NULL,
fit_from = NULL)Arguments
- x
Numeric vector of ages at the beginning of each age interval. For a full life table, use single-year ages (e.g.,
0:110). For an abridged life table, use the lower bound of each interval (e.g.,c(0, 1, 5, 10, ..., 110)).- Dx
Death counts. Each element represents the total number of deaths during the calendar year to persons aged
xtox + n(wherenis the length of the age interval). Must be provided together withEx.- Ex
Exposure-to-risk in the period. This is usually approximated by the mid-year population aged
xtox + n. Must be provided together withDx.- mx
Age-specific death rate in the age interval
[x, x+n). Defined asDx / Ex.- qx
Probability of dying within the age interval
[x, x+n).- lx
Probability of surviving to exact age
x(iflx0 = 1), or the number of survivors at exact agex(iflx0 > 1). Whenlxis the sole input, the values are re-scaled to the chosen radixlx0.- dx
Number of deaths in the life-table population occurring in the age interval
[x, x+n). Whendxis the sole input, the values are re-scaled to sum tolx0.- ex
Remaining life expectancy at age
x, in years. Whenexis the sole input the function builds the life table that reproduces the supplied curve: the survivorship, death probabilities and death rates are recovered fromexand theaxconvention (seeDetails). A curve that rises with age at the youngest ages is allowed; a missing value anywhere in the curve, or a curve that is not a feasible life table (falling faster than the interval width, or below its ownax), is an error naming the affected ages. A missingexis never repaired the way a missing rate is.- sex
Sex of the population. Options are
NULL(default),"male","female", or"total". When specified, the first two entries of theaxcolumn are adjusted using Coale-Demeny coefficients, producing more accurate life-table values at the youngest ages. The adjustment differs slightly between males and females.- lx0
Radix, the starting population (or probability scale) at age 0. Default is
100,000. All subsequent life-table columns (lx,dx,Lx,Tx) are scaled accordingly.- ax
The average number of person-years lived in each age interval by those who die in it, given either as values or as the method that produces them. Accepts two forms:
a numeric scalar or vector: a scalar is applied to all intervals, a vector must have the same length as
x. A common assumption isax = 0.5, which places deaths at the midpoint of each interval;a method name, one of
"andreev_kingkade"(the default),"cfm","preston"or"coale_demeny"(see below).
"andreev_kingkade"follows the rule the Human Mortality Database applies to its period life tables (Methods Protocol, version 6, section 7.1): the Andreev-Kingkade (2015) formula sets the value for the first year of life fromm0, every other closed interval uses half its length (n/2), and the open interval uses1/mx. It is the most accurate of the four methods and the one that reproduces the published HMD life tables most closely. Because it produces a numericaxbefore the rates are converted, themxtoqxstep uses the exact identityqx = n*mx/(1 + (n - ax)*mx), which is the identity the protocol uses (equation 74). The Andreev-Kingkade value is a property of a one-year first interval that starts at birth, so when the first interval is wider or the table starts above age 0 the interval keeps the ordinary midpoint value instead.The other three methods share the same basis, the standard lifetable identity
ax = n + 1/m - n/qunder a constant force of mortality (Preston, Heuveline and Guillot 2001, eq. 3.15), and differ in how they treat the first two intervals:"cfm": the identity alone, for every interval;"preston": the identity, with the first two intervals replaced by the Coale-Demeny West separation factors published by Preston et al. (2001), table 3.3, whensexis given;"coale_demeny": the identity, with the first two intervals replaced by the original 1983 Coale-Demeny rule expressed inqx, reproduced by the PAS software, whensexis given.
"preston"and"coale_demeny"agree exactly oncemx[1]reaches 0.107, and differ by at most a few thousandths of a year below it. Both coincide with"cfm"whensexisNULL."cfm"and the Coale-Demeny variants adjustaxafter the rates have been converted, so they leave the constant force of mortality conversion in place; only"andreev_kingkade"changes the conversion itself.A value supplied for the open age interval is kept as given. Everybody alive there dies in that interval, so its own rate implies what the average time lived in it should be (
1/mx); a value that differs is reported and left alone, which leaves theaxandmxcolumns of the table disagreeing at that age, or the model-implied value whencloseis set. The one exception is a table entered fromex, where the curve being inverted fixes that interval. SeeDetails.- close
The method used to close the open age interval, named by the mortality-law code that implements it.
NULL(the default) keeps the standard reciprocal close,mx[N] = (the observed rate). A code fromavailableLaws(e.g."kannisto") closes the table instead with the model-implied average force of mortality over the open interval: the law is fitted to the closed intervals and the observedmx[N]is replaced by1/ex[N]. This corrects the bias of the reciprocal close when the open interval begins at a young age, and it does not change the age grid. The standard row identities (qx[N] = 1,ax[N] = ex[N] = 1/mx[N]) still hold on the corrected rate. The same code is used whenomegaextends the table; whenomegais set andcloseisNULL, the extrapolation defaults to"kannisto".- omega
The age at which to close the table when it should be extended beyond the input's open age.
NULL(the default) keeps the table closed at the input's own open age. Whenomegais greater than the last age inx, the death rates from the open age up toomegaare obtained by extrapolating thecloselaw (see alsofit_from) and the table is closed atomega. The input's own open interval is replaced by the extrapolated values, so a user-suppliedaxis re-estimated on the extended grid. Anomegathat does not exceed the last age inxleaves the table unchanged, with a warning. Ignored by the in-place close, which operates on the input's own open age.- fit_from
The age from which the closing law is fitted.
NULL(the default) uses 60 when the input reaches age 85, and the last 20 years of the input otherwise. Ignored when neitherclosenoromegais set.
Value
An object of class "LifeTable" containing the following
components:
- lt
A
data.framewith the complete life table, including columns for age interval (x.int), exact age (x), death rate (mx), death probability (qx), person-years lived by decedents (ax), survivorship (lx), death distribution (dx), person-years lived (Lx), total person-years remaining (Tx), and life expectancy (ex).- call
The matched function call.
- process_date
Timestamp of when the life table was computed.
Details
A life table (also called a mortality table or actuarial table) summarises the mortality experience of a population. For each age (or age interval) it reports:
Death rates (
mx) and death probabilities (qx)Survivorship (
lx)Distribution of deaths (
dx)Person-years lived (
Lx) and total person-years remaining (Tx)Life expectancy (
ex)
The life table is constructed sequentially: from the input data the
function derives mx, then qx, then lx, dx,
Lx, Tx, and finally ex. The conversion between
mx and qx follows the ax method. The default,
ax = "andreev_kingkade", is the rule the Human Mortality Database
applies to its period life tables (Methods Protocol, version 6): the
Andreev-Kingkade (2015) formula for the first year of life and half the
interval width elsewhere, converted through the exact
interval identity so that the mx, qx and ax columns
are mutually consistent. The constant-force-of-mortality (CFM)
assumption, ax = n + 1/m - n/q, is available with ax =
"cfm". When sex is given, the first two values of the ax
column can instead be adjusted using the Coale-Demeny method, which
accounts for the different infant mortality patterns between males and
females; two published parameterisations of that adjustment are available
through ax = "preston", expressed in terms of mx, and
ax = "coale_demeny", expressed in terms of qx and retrieved
from mx by the PAS inversion. The latter reproduces the
coefficients used by the Coale-Demeny 1983 regional model tables and by
the PAS software.
The ex input runs the table backwards. The survivorship ratio of
each closed interval follows from the identity
$$e_x l_x - e_{x+n} l_{x+n} = nL_x = a_x l_x + (n - a_x) l_{x+n}$$, so
that \(r_x = l_{x+n}/l_x = (e_x - a_x)/(e_{x+n} + n - a_x)\). When
ax is supplied as a numeric vector the recovery is a single sweep.
When it is not, each interval is solved so that the ax the method in force
would assign to the recovered interval reproduces the ratio it came from;
the recovered table is then identical to the one the same ex curve
describes under that method. The open interval closes by the standard rule,
\(a_N = e_N\) and \(m_N = 1/e_N\). A curve of life expectancy is
allowed to rise with age at the youngest ages (life expectancy at birth is
lower than at age one when infant mortality is high), so a rise is not an
error; a curve that falls by more than the interval width, or below its own
\(a_x\), is not a feasible life table and is reported by age. Because
\(e_x\) alone does not pin \(a_x\), pass ax alongside ex
when the ax convention matters. One nuance follows from the forward build:
when ax is "preston" or "coale_demeny" and a sex
is given, the forward table adjusts ax in the first two intervals
after converting the rates, so its stored mx there does not satisfy
the exact identity on its own ax. The inverse returns the
identity-consistent mx, so in that one family the recovered
mx in the first two rows can differ from the forward table's at the
third decimal while ex, qx and ax still reproduce
exactly.
References
Coale, A. J., Demeny, P., and Vaughan, B. (1983). Regional Model Life Tables and Stable Populations. 2nd ed. New York: Academic Press.
Preston, S. H., Heuveline, P., and Guillot, M. (2001). Demography: Measuring and Modeling Population Processes. Oxford: Blackwell Publishers.
When ax is supplied by the user as a numeric vector the conversion
between mx and qx uses the exact interval identity
qx = nx * mx / (1 + (nx - ax) * mx) (and its inverse) instead
of the CFM approximation, so that the mx, qx and
ax columns of the result are mutually consistent. The
"andreev_kingkade" method does the same, because it builds a
numeric ax before the conversion. An ax
that an interval cannot support (ax * mx > 1) is replaced with
the implied average, 1/mx, and the affected ages are reported.
The open (closing) age interval follows its own rule:
ax[N] = 1/mx[N], ex[N] = 1/mx[N] and
Lx[N] = lx[N]/mx[N], which keeps the closed table consistent
with qx[N] = 1. A user-supplied ax[N] is therefore
replaced, with a warning.
That rule assumes the force of mortality is roughly constant above the
open age, which holds when the open interval is old (85+ or 90+) but
not when the input stops early (70+ or 75+). The close argument
addresses that directly: given a mortality-law code, the observed
mx[N] is replaced by the model-implied average over the open
interval, 1/ex[N], without changing the age grid. The
omega argument is a different response to the same problem: it
extends the table to omega by extrapolating the rates and closes
there. Either argument alone leaves the other's default in place, and
the default (close = NULL, omega = NULL) is the standard
reciprocal close, so results are unchanged.
Missing values in mx or qx are never masked. They
localise to the interval quantities of the affected row (which are
returned as NA) and to the cumulative Tx and ex
columns at and before that row, while the survivorship chain
lx is bridged across the gap, so the ages above it remain
computable. A warning names the affected ages.
Examples
# Example 1 --- Full life tables with different inputs ------------
y <- 1900
x <- as.numeric(rownames(ahmd$mx))
Dx <- ahmd$Dx[, paste(y)]
Ex <- ahmd$Ex[, paste(y)]
LT1 <- LifeTable(x, Dx = Dx, Ex = Ex)
LT2 <- LifeTable(x, mx = LT1$lt$mx)
LT3 <- LifeTable(x, qx = LT1$lt$qx)
LT4 <- LifeTable(x, lx = LT1$lt$lx)
LT5 <- LifeTable(x, dx = LT1$lt$dx)
LT6 <- LifeTable(x, ex = LT1$lt$ex)
LT1
#>
#> Full Life Table
#>
#> Number of life tables: 1
#> Dimension: 111 x 10
#> Age intervals: [0,1) [1,2) [2,3) ... ... [108,109) [109,110) [110,+)
#>
#> x.int x mx qx ax lx dx Lx Tx ex
#> [0,1) 0 0.1513 0.1369 0.31 1e+05 13694 90505 4825279 48.25
#> [1,2) 1 0.0479 0.0468 0.5 86306 4040 84286 4734774 54.86
#> [2,3) 2 0.0191 0.0189 0.5 82266 1559 81486 4650488 56.53
#> [3,4) 3 0.013 0.0129 0.5 80707 1045 80185 4569001 56.61
#> [4,5) 4 0.0094 0.0093 0.5 79662 743 79290 4488817 56.35
#> [5,6) 5 0.0067 0.0067 0.5 78919 525 78656 4409526 55.87
#> <NA> ... ... ... ... ... ... ... ... ...
#> [108,109) 108 6.75 1 0.15 0 0 0 0 0
#> [109,110) 109 6.75 1 0.15 0 0 0 0 0
#> [110,+) 110 6.75 1 0.15 0 0 0 0 0
LT5
#>
#> Full Life Table
#>
#> Number of life tables: 1
#> Dimension: 111 x 10
#> Age intervals: [0,1) [1,2) [2,3) ... ... [108,109) [109,110) [110,+)
#>
#> x.int x mx qx ax lx dx Lx Tx ex
#> [0,1) 0 0.1513 0.1369 0.31 1e+05 13694 90505 4825279 48.25
#> [1,2) 1 0.0479 0.0468 0.5 86306 4040 84286 4734774 54.86
#> [2,3) 2 0.0191 0.0189 0.5 82266 1559 81486 4650488 56.53
#> [3,4) 3 0.013 0.0129 0.5 80707 1045 80185 4569001 56.61
#> [4,5) 4 0.0094 0.0093 0.5 79662 743 79290 4488817 56.35
#> [5,6) 5 0.0067 0.0067 0.5 78919 525 78656 4409526 55.87
#> <NA> ... ... ... ... ... ... ... ... ...
#> [108,109) 108 34.4251 1 0.03 0 0 0 0 0
#> [109,110) 109 83.186 1 0.01 0 0 0 0 0
#> [110,+) 110 201.0132 1 0 0 0 0 0 0
LT6
#>
#> Full Life Table
#>
#> Number of life tables: 1
#> Dimension: 111 x 10
#> Age intervals: [0,1) [1,2) [2,3) ... ... [108,109) [109,110) [110,+)
#>
#> x.int x mx qx ax lx dx Lx Tx ex
#> [0,1) 0 0.1513 0.1369 0.31 1e+05 13694 90505 4825279 48.25
#> [1,2) 1 0.0479 0.0468 0.5 86306 4040 84286 4734773 54.86
#> [2,3) 2 0.0191 0.0189 0.5 82266 1559 81486 4650488 56.53
#> [3,4) 3 0.013 0.0129 0.5 80707 1045 80185 4569001 56.61
#> [4,5) 4 0.0094 0.0093 0.5 79662 743 79290 4488816 56.35
#> [5,6) 5 0.0067 0.0067 0.5 78919 525 78656 4409526 55.87
#> <NA> ... ... ... ... ... ... ... ... ...
#> [108,109) 108 6.75 1 0.15 0 0 0 0 0
#> [109,110) 109 6.75 1 0.15 0 0 0 0 0
#> [110,+) 110 6.75 1 0.15 0 0 0 0 0
ls(LT5)
#> [1] "call" "lt" "process_date"
# Example 2 --- Compute multiple life tables at once ------------
LTs <- LifeTable(x, mx = ahmd$mx)
#> 'mx' contains missing or non-finite values at age(s) 107, 108, 109, 110. Rows below age 100 are returned as NA with the survivorship bridged across the gap; rows from age 100 are replaced with the highest rate observed there.
LTs
#>
#> Full Life Tables
#>
#> Number of life tables: 4
#> Dimension: 444 x 11
#> Age intervals: [0,1) [1,2) [2,3) ... ... [108,109) [109,110) [110,+)
#>
#> LT x.int x mx qx ax lx dx Lx Tx ex
#> 1850 [0,1) 0 0.1418 0.1291 0.31 1e+05 12907 91051 4366316 43.66
#> 1850 [1,2) 1 0.0586 0.0569 0.5 87093 4960 84613 4275266 49.09
#> 1850 [2,3) 2 0.0312 0.0307 0.5 82133 2523 80872 4190653 51.02
#> 1850 [3,4) 3 0.0217 0.0215 0.5 79610 1711 78754 4109781 51.62
#> 1850 [4,5) 4 0.0169 0.0167 0.5 77899 1304 77247 4031026 51.75
#> 1850 [5,6) 5 0.0128 0.0127 0.5 76594 974 76107 3953780 51.62
#> <NA> <NA> ... ... ... ... ... ... ... ... ...
#> 2010 [108,109) 108 0.6716 0.5028 0.5 40 20 30 60 1.48
#> 2010 [109,110) 109 0.4109 0.3408 0.5 20 7 17 30 1.48
#> 2010 [110,+) 110 1.0209 1 0.98 13 13 13 13 0.98
# A warning is printed if the input contains missing values.
# Some of the missing values can be handled automatically.
# Example 3 --- Abridged life table -----------------------------
x <- c(0, 1, seq(5, 110, by = 5))
mx <- c(.053, .005, .001, .0012, .0018, .002, .003, .004,
.004, .005, .006, .0093, .0129, .019, .031, .049,
.084, .129, .180, .2354, .3085, .390, .478, .551)
LT7 <- LifeTable(x, mx = mx, sex = "female")
LT7
#>
#> Abridged Life Table
#>
#> Number of life tables: 1
#> Dimension: 24 x 10
#> Age intervals: [0,1) [1,5) [5,10) ... ... [100,105) [105,110) [110,+)
#>
#> x.int x mx qx ax lx dx Lx Tx ex
#> [0,1) 0 0.053 0.051 0.25 1e+05 5098 96189 6491129 64.91
#> [1,5) 1 0.005 0.0198 1.99 94902 1879 375837 6394940 67.38
#> [5,10) 5 0.001 0.005 2.5 93023 464 463953 6019103 64.71
#> [10,15) 10 0.0012 0.006 2.5 92559 554 461409 5555150 60.02
#> [15,20) 15 0.0018 0.009 2.5 92005 824 457962 5093741 55.36
#> [20,25) 20 0.002 0.01 2.5 91181 907 453632 4635779 50.84
#> <NA> ... ... ... ... ... ... ... ... ...
#> [100,105) 100 0.39 0.8577 1.73 408 350 896 1016 2.49
#> [105,110) 105 0.478 0.9084 1.59 58 53 110 120 2.07
#> [110,+) 110 0.551 1 1.81 5 5 10 10 1.81
# Example 4 --- Abridged life table using a custom 'ax' --------
# This example reuses the ages (x) and death rates (mx) from Example 3.
# Note that 'ax' must have the same length as 'x', otherwise an error
# will be returned.
my_ax <- c(0.1, 1.5, rep(2, 19), 1, 1, 1)
LT8 <- LifeTable(x = x, mx = mx, ax = my_ax)
#> 'ax' at the open age interval (age 110) is used as supplied, 1. Everyone alive at 110 dies in that interval, so the average time lived in it is implied by its own death rate, 1/mx = 1.8149. Keeping the supplied value leaves the table's 'ax' and 'mx' columns disagreeing at that age.
# Example 5 --- The ax methods ------------------------------
# The default 'andreev_kingkade' follows the HMD Methods Protocol v6
# (Andreev-Kingkade a0, half-interval elsewhere). 'cfm' is the plain
# constant-force identity; 'preston' and 'coale_demeny' are the two
# Coale-Demeny conventions for the first two intervals (identical above
# m0 = 0.107).
LT9 <- LifeTable(x, mx = mx, sex = "female", ax = "andreev_kingkade")
LT10 <- LifeTable(x, mx = mx, sex = "female", ax = "cfm")
LT11 <- LifeTable(x, mx = mx, sex = "female", ax = "preston")
LT12 <- LifeTable(x, mx = mx, sex = "female", ax = "coale_demeny")
rbind(
andreev_kingkade = LT9$lt$ax[1:2],
cfm = LT10$lt$ax[1:2],
preston = LT11$lt$ax[1:2],
coale_demeny = LT12$lt$ax[1:2]
)
#> [,1] [,2]
#> andreev_kingkade 0.2523572 1.993333
#> cfm 0.4955835 1.993333
#> preston 0.2014000 1.441546
#> coale_demeny 0.2025524 1.441266
# Example 6 --- Closing the open interval accurately -----------
# The data stop at 75+; closing there assumes a constant hazard above 75.
# 'close' argument corrects the open-interval rate with a fitted law, on the same
# age grid;
# 'omega' argument offers the option to extend the table to 110 before closing it.
x5 <- c(0, 1, seq(5, 75, by = 5))
mx5 <- c(.053, .005, .001, .0012, .0018, .002, .003, .004,
.004, .005, .006, .0093, .0129, .019, .031, .049, .084)
LT13 <- LifeTable(x5, mx = mx5)
LT14 <- LifeTable(x5, mx = mx5, close = "kannisto")
LT15 <- LifeTable(x5, mx = mx5, close = "kannisto", omega = 110)
c(default = LT13$lt$ex[1], close = LT14$lt$ex[1], omega = LT15$lt$ex[1])
#> default close omega
#> 66.55551 64.71610 65.16495