
| # Seeders per group: | |
| # Contacts per group: | |
| # Replicates: | |
| Shape parameter k: | |
| % Infected ϕ: | |
| Dominance: | AB |
| Ntotal = | Individuals |
| SD in ag = | (Susceptibility) |
| SD in af = | (Infectivity) |
| SD in ar = | (Recovery) |
| SD in Δg = | (Dom. in Susc.) |
| SD in Δf = | (Dom. in Inf.) |
| SD in Δr = | (Dom. in Rec.) |
| NAA | = | ||
| NBB | = | ||
| NAB | = |
This tool assumes that a disease transmission experiment is going to be undertaken to establish how a particular SNP affects the susceptibility, infectivity and recovery rates of individuals. The SNP under study is assumed to take two alleles: A and B. One key question to ask is how the experiment should be designed in order to maximised the amount of available information. (Note, here it is assumed that the exact timings of infection and recovery events from the experiment are known.)
The tool works in the following way:
The first option to select is the type of experimental design to be carried out. These include the five potential designs outlined in the paper: "Single group (without dominance)", "Pure design (without dominance)", "Pure design (with dominance)", "Mixed design (without dominance)" and "Mixed design (with dominance)". Selecting "General" allows the user to set an arbitrary design.
The number of seeders and contacts refers to the number of individuals initially infected and susceptible in each contact group. The genetic compositions of these subpopulations is shown by the panels on the right (one for each contact group). Here NAA, NBB and NAB refer to the number of individuals in each of the three genotypes. These genetic compositions may be changed by either clicking on the numbers themselves, or by dragging the purple circles within the triangle plots.
Often the same basic design is repeated multiple times as a way of increasing statistical power. This duplication is represented by the number of "replicates".
The shape parameter k governs the gamma distributed recovery profile (note this quantity only affects estimates for recovery rate parameters).
The % infected slider sets ϕ, the expected fraction of contacts that become infected during the course of the experiment. This may be significantly less than 100% if the experiment is terminated early or the basic reproductive ratio R0 is low.
In some circumstances allele A may dominate over B, or vice-versa. This affect can but accounted for by adjusting the relevant dominance slider.
The first number shows the total number of individuals for the experiment.
Model parameters ag, af, and ar represent the relative differences in susceptibility, infectivity and recovery rate for individuals with an A compared to a B allele at the SNP under investigation. Specifically ag=0.1 represents the case in which individuals with genotype AA are approximately 20% more susceptible to disease than those with BB (see paper for a more precise definition).
Due to the stochastic nature of the data obtained from disease transmission experiments (i.e. infection and recovery times), precise estimates for ag, af, and ar are not possible. What this tool shows is the expected standard deviation in the posterior distribution of these quantities, which provide estimates for their accuracy. Small numbers represent a higher degree of precision, and so the experimental design should be chosen to minimise these standard deviations as much as possible.
Parameters Δg, Δf, and Δr represent the scaled dominance of allele A over B. Note, the standard deviations in these quantities are each divided by the corresponding effect size (because if the SNP effect size is small, it becomes harder to establish dominance).
It is important to note that the estimates provided by this tool represent lower bounds on the standard deviations of model parameters. As shown in section 3.4 of the paper, residual contributions, group and fixed effects and incomplete data will all act to increase these standard deviations, and so reduce the statistical power with which SNP-based associations can be made.