Working - please wait...

Bayesian Latent Class Models - study design

Sample size, precision and identifiability by simulation

Sample sizes
Number of individuals sampled and tested in each population.

Target precision
The precision you require of the study, given as the half-width of the credible interval: a target of 0.10 means an interval no wider than ±10 percentage points. Leave a cell empty to drop the target for that parameter.
How to cite this tool: Tool developed by Alberto Gomez-Buendia (Universidad Complutense de Madrid) and Simon Firestone (University of Melbourne).
Until our paper specific to this tool is available, please cite:
Cheung et al. (2021). Bayesian latent class analysis when there is an imperfect reference test. Rev Sci Tech Off Int Epiz.

We also make use of R, JAGS and the R packages R2jags, runjags and mcmcplots. Please check their websites for how to cite their contributions.

Note - WOAH Terrestrial Animal Health Manual (Ch. 1.1.6):
Because Bayesian latent class models are complex and require adherence to critical assumptions, statistical assistance should be sought to help guide the analysis and describe the sampling from the target population(s), the characteristics of other tests included in the analysis, the appropriate choice of model and the estimation methods should be based on peer-reviewed literature.
Read the full chapter
Unobserved parameters to be recaptured
These are the values used to simulate the data. The BLCM will then try to recover them, which is what makes the identifiability check meaningful.

Double-click any cell in the True Value column to edit it.
Choose values you believe are plausible for your tests and populations. If you are unsure, run the app more than once across the range you consider credible.
Covariance terms in the fitted model
Select the conditional dependence terms your BLCM will estimate.
Diseased
Non-Diseased

Important: If higher-order terms are selected (e.g. All Tests or Tests 1 & 2 & 3) then all nested lower-order terms (e.g. Tests 1 & 2) must also be selected. Selection should reflect biological plausibility, data availability, and convergence.
True dependence used to simulate the data
Pairwise correlations between tests, conditional on disease status. These are converted to covariances using the true Se and Sp.

Resulting covariances and their feasible bounds. The bounds follow from the accuracies alone (Se for covD, Sp for covN), not from the covariance, so they differ between the two strata:

Higher-order covariances have no simple correlation equivalent, so enter them directly below. Keep these values for higher-order terms low (<0.1) or you will have issues with initial values and chains not running.
You can simulate data with dependence and fit a model without it (or the reverse). Doing so shows you how much bias a misspecified dependence structure would introduce in your study.
Simulate datasets
Generates cross-classified test results from the true values, sample sizes and dependence structure you specified.




Prior distributions

Double-click any cell in the table immediately below to adjust either shape parameter, or you can enter them with out BetaBuster tool below the table.


Derive shape parameters from an elicited opinion
Give the most likely value and a bound the expert is confident the parameter lies beyond. A bound below the mode is read as a lower bound, one above it as an upper bound.



Prior strength
The strength of a prior is directly represented by the sum of its hyperparameters, a + b. Use this table to evaluate whether the application of a given prior may produce a posterior driven by it rather than by the data.

Note - Prior misspecification is a critical issue that influences the validity of the inference.
The default Beta(1, 1) priors are flat and may not be the most suitable choice for most analyses. Statistical assistance should be sought to help guide decisions on prior specification and prior sensitivity analyses. Appropriate choices of priors can be based on peer-reviewed literature and expert opinion.

Our online Beta-Buster tool also derives prior Beta distribution shape parameters
MCMC settings
Burn-in samples
Iterations
Thinning
Increasing iterations improves mixing but extends runtime, and here the runtime is multiplied by the number of datasets you analyse.
Initial values
Two over-dispersed chains. Deliberately different starting points make the convergence check informative.
Run the BLCMs
Each analysed dataset is fitted with the model formulation, priors and MCMC settings you specified.



Design performance
precision is the mean half-width of the credible interval, worst is the widest one across datasets, coverage is the proportion of datasets whose interval contained the true value, and bias is the mean difference between the posterior median and the true value.

n.eff should be > 200 Rhat must be < 1.01
Terms held at a constant, such as a covariance you switched off, are ignored by the convergence summary. If the requirements are not met, re-run with more iterations before drawing any conclusion about the design.
How to cite this tool: Tool developed by Alberto Gomez-Buendia (Universidad Complutense de Madrid) and Simon Firestone (University of Melbourne).
Until our paper specific to this tool is available, please cite:
Cheung et al. (2021). Bayesian latent class analysis when there is an imperfect reference test. Rev Sci Tech Off Int Epiz.

We also make use of R, JAGS and the R packages R2jags, runjags and mcmcplots. Please check their websites for how to cite their contributions.

Note - WOAH Terrestrial Animal Health Manual (Ch. 1.1.6):
Because Bayesian latent class models are complex and require adherence to critical assumptions, statistical assistance should be sought to help guide the analysis and describe the sampling from the target population(s), the characteristics of other tests included in the analysis, the appropriate choice of model and the estimation methods should be based on peer-reviewed literature.
Read the full chapter
Is the prior or the data driving the posterior?
Orange is the prior, blue is the posterior pooled over the analysed datasets, and the red line is the true value. A posterior that sits on top of its prior means the study adds little information about that parameter.

Download plot
Inputted values vs posterior estimates
One interval per analysed dataset. If the intervals do not bracket the red true value, or drift with the starting values, the parameter is not being identified by this design.

Download plot
Achieved precision vs target

Download plot
Diagnostic plots
Trace, density and autocorrelation plots for the last dataset analysed.