Simulates a full data set for a given set of parameters and sampling design.

simulate_data(
  par_tab,
  group = 1,
  n_indiv = 100,
  antigenic_map = NULL,
  possible_exposure_times = NULL,
  measured_biomarker_ids = NULL,
  sampling_times,
  nsamps = 2,
  missing_data = 0,
  age_min = 5,
  age_max = 80,
  age_group_bounds = NULL,
  attack_rates,
  repeats = 1,
  measurement_bias = NULL,
  data_type = NULL,
  demographics = NULL,
  verbose = FALSE,
  starting_levels = NULL,
  exponential_waning = FALSE,
  coefficient_values = NULL
)

Arguments

par_tab

the parameter table controlling parameter ranges and values. When stratification is requested, coefficient rows are added automatically; use `coefficient_values` to set selected coefficient values for the simulated truth.

group

which group index to give this simulated data

n_indiv

number of individuals to simulate

antigenic_map

(optional) A data frame of antigenic x and y coordinates. Must have column names: x_coord; y_coord; inf_times. See example_antigenic_map.

possible_exposure_times

(optional) If no antigenic map is specified, this argument gives the vector of times at which individuals can be infected

measured_biomarker_ids

vector of biomarker IDs that have measurements, matching entries in possible_exposure_times

sampling_times

possible sample times for the individuals, matching the model time scale

nsamps

the number of samples each individual has (for example, `nsamps = 2` gives each individual two random sample times from `sampling_times`)

missing_data

numeric between 0 and 1, used to censor a proportion of observations at random (MAR)

age_min

simulated age minimum

age_max

simulated age maximum

age_group_bounds

optional age-group boundaries used when creating demographic groups

attack_rates

a vector or table of attack rates for each entry in possible_exposure_times to be used in the simulation (between 0 and 1). See simulate_attack_rates.

repeats

number of repeat observations for each year

measurement_bias

default NULL, optional vector of measurement shifts used when generating the simulated antibody levels

data_type

numeric or text observation-model types to use for each `biomarker_group`: `1` or `"discrete"` for discrete, bounded observations; `2` or `"continuous"` for continuous, bounded observations; or `3` or `"false_positive"` for continuous observations with the false-positive model. A single value is used for all biomarker groups.

demographics

if not NULL, a data frame giving demographic variables for each individual (1:n_indiv). It must include `birth` and can include `population_group` or variables used for stratification in `par_tab`.

verbose

if TRUE, prints additional messages

starting_levels

a data frame or function giving the starting biomarker level for each individual, `biomarker_group`, and `biomarker_id` combination. If NULL, starting levels are assumed to be 0.

exponential_waning

Deprecated compatibility argument. If TRUE, uses exponential waning rather than linear waning. Prefer a fixed `exponential_waning` row in `par_tab`, with `values = 1` and `par_type = 0`.

coefficient_values

optional data frame specifying coefficient values used when simulating stratified parameters. It must contain `parameter`, `stratification`, `stratification_level`, `biomarker_group`, and `value` columns. `parameter` is the base parameter name in `par_tab`, and `value` is the coefficient for the specified stratification level. If NULL, generated coefficients retain their existing default values.

Value

A list containing `antibody_data`, `infection_histories`, `attack_rates`, `phis`, `par_tab`, `population_groups`, `demographic_groups`, and `start_levels`.

See also

Examples

data(example_par_tab)
data(example_antigenic_map)

## Times at which individuals can be infected
possible_exposure_times <- example_antigenic_map$inf_times
## Simulate some random attack rates between 0 and 0.2
attack_rates <- simulate_attack_rates(possible_exposure_times,
                                      mean_par = 0.1)
## Vector giving the circulation times of measured antigens
sampled_antigens <- seq(min(possible_exposure_times), max(possible_exposure_times), by=2)
all_simulated_data <- simulate_data(par_tab=example_par_tab, group=1, n_indiv=50,    
                                   possible_exposure_times=possible_exposure_times,
                                   measured_biomarker_ids=sampled_antigens,
                                   sampling_times=2010:2015, nsamps=2, antigenic_map=example_antigenic_map, 
                                   age_min=10,age_max=75,
                                   attack_rates=attack_rates, repeats=2)
antibody_data <- all_simulated_data$antibody_data