Skip to contents

This function calculates a univariate correlated fitness function using a GAM smooth term. The formula is w ~ f(z). Traits should be standardized BEFORE calling this function.

Usage

univariate_spline(
  data,
  fitness_col,
  trait_col,
  fitness_type = c("auto", "binary", "count", "continuous"),
  group = NULL,
  relative_col = NULL,
  k = 10,
  bs = c("cr", "tp", "ps"),
  smoothing = c("GCV.Cp", "REML", "ML"),
  bootstrap = FALSE,
  n_boot = 1000,
  by_group = FALSE,
  count_family = c("poisson", "quasipoisson", "nb")
)

Arguments

data

A data frame containing fitness and trait measurements.

fitness_col

A string specifying the name of the fitness column.

trait_col

A string specifying the name of the trait column (must be numeric and standardized).

fitness_type

A string indicating the fitness type: "auto" (detect from the data, the default), "binary", "count", or "continuous". Binary and count fitness are fitted on the raw values with a binomial or Poisson family; continuous fitness on relative fitness with a Gaussian one.

group

Optional string specifying a grouping variable. If provided, group fixed effects are included (none when it has only one level); rows with no group label are one more level.

relative_col

Optional string naming a pre-computed relative fitness column to use for continuous fitness (e.g. one produced within groups by prepare_selection_data).

k

Integer specifying the basis dimension for the smooth term. Default is 10.

bs

Spline basis: "cr" (cubic regression spline, the default), "tp" (thin plate) or "ps" (P-spline).

smoothing

How the smoothing parameter is chosen: "GCV.Cp" (generalised cross-validation, the default, which mgcv applies in its UBRE form to binary and Poisson fitness and as plain GCV to quasi-Poisson; a negative binomial fit uses REML in its place), "REML" or "ML". The result's spline_type names the criterion mgcv used.

bootstrap

Logical; if TRUE the 95% ribbon is obtained by resampling individuals and refitting (Schluter 1988). The default FALSE uses the parametric Wald interval, which is fast; the bootstrap refits the spline n_boot times.

n_boot

Integer number of bootstrap resamples used when bootstrap = TRUE. Default is 1000.

by_group

Logical; if TRUE fit each level of group on its own rows and return a named list of fits, one per group. The default FALSE fits one curve with the group as a fixed effect. Prepare the data with the same group first, so that each group is standardised on its own.

count_family

Family for count fitness: "poisson" (the default), "quasipoisson" or "nb", as in correlated_fitness_surface(); the smoothing parameter is chosen under whichever is used. The result's dispersion is the Pearson dispersion of the fit, and a Poisson fit warns when it is above 1.5.

Value

A list of class "univariate_fitness" containing the fitted GAM model, a prediction grid, and metadata.

Details

By default the fitness function is a penalised cubic regression spline with the smoothing parameter chosen by generalised cross-validation (UBRE for binary and Poisson fitness), following Schluter (1988). bs and smoothing set the family of curves and the criterion, to match the smoother of another study; the amount of smoothing is still estimated from the data, though the criterion can change it a good deal. If mgcv's check suggests the basis dimension was too small a warning says so; the check works on the residuals, which for survival carry too little to test, so it is skipped there. With bootstrap = TRUE the result depends on the random seed; call set.seed() first for a reproducible ribbon.

Examples

prep <- prepare_selection_data(bumpus, "survival", "total_length")
uni <- univariate_spline(prep, "survival", "total_length")
#> Fitness type detected: binary
plot_univariate_fitness(uni, "total_length")


# a thin-plate basis with REML, to match another study's smoother
univariate_spline(prep, "survival", "total_length", bs = "tp", smoothing = "REML")$spline_type
#> Fitness type detected: binary
#> [1] "thin-plate spline (REML)"