absolute risk increase (ARI)
The increase in risk with a new therapy compared with the risk without the new therapy.
absolute risk reduction (ARR)
The reduction in risk with a new therapy compared with the risk without the new therapy; it is the absolute value of the difference between the experimental event rate and the control event rate (|EER – CER|).
absolute value
The positive value of a number, regardless of whether the number is positive or negative. The absolute value of a is symbolized |a|.
actuarial analysis
See life table analysis.
addition rule
The rule which states the probability that two or more mutually exclusive events all occur is the sum of the probabilities of each individual event.
adjusted rate
A rate adjusted so that it is independent of the distribution of a possible confounding variable. For example, age-adjusted rates are independent of the age distribution in the population to which they apply.
age-specific mortality rate
The mortality rate in a specific age group.
alpha (α) error
See type I error.
alpha (α) value
The level of alpha (α) selected in a hypothesis test.
alternative hypothesis
The opposite of the null hypothesis. It is the conclusion when the null hypothesis is rejected.
analysis of covariance (ANCOVA)
A special type of analysis of variance or regression used to control for the effect of a possible confounding factor.
analysis of residuals
In regression, an analysis of the differences between Y and Y' to evaluate assumptions and provide guidance on how well the equation fits the data.
analysis of variance (ANOVA)
A statistical procedure that determines whether any differences exist among two or more groups of subjects on one or more factors. The Ftest is used in ANOVA.
backward elimination
A method to select variables in multiple regression that enters all variables into the regression equation and then eliminates the variable that adds the least to the prediction, followed by the other variables one at a time that decrease the multiple R by the least amount until all statistically significant variables are removed from the equation.
bar chart or bar graph
A chart or graph used with nominal characteristics to display the numbers or percentages of observations with the characteristic of interest.
Bayes' theorem
A formula for calculating the conditional probability of one event, P(A|B), from the conditional probability of the other event, P(B|a).
bell-shaped distribution
A term used to describe the shape of the normal (gaussian) distribution.
beta (β) error
See type II error.
bias
The error related to the ways the targeted and sampled populations differ; also called measurement error, it threatens the validity of a study.
binary observation
A nominal measure that has only two outcomes (examples are gender: male or female; survival: yes or no).
binomial distribution
The probability distribution that describes the number of successes X observed in n independent trials, each with the same probability of occurrence.
biometrics
The study of measurement and statistical analysis in medicine and biology.
biostatistics
The application of research study design and statistical analysis to applications in medicine and biology.
bivariate plot
A two-dimensional plot or scatterplot of the values of two characteristics measured on the same set of subjects.
blind study
An experimental study in which subjects do not know the treatment they are receiving; investigators may also be blind to the treatment subjects are receiving; see also double-blind trial.
block design
In analysis of variance, a design in which subjects within each block (or stratum) are assigned to a different treatment.
P.404
Bonferroni t
A method for comparing means in analysis of variance; also called the Dunn multiple-comparison procedure.
bootstrap
A method for estimating standard errors or confidence intervals in which a small sample of observations is randomly selected from the original sample, estimates are calculated, and the sample is returned to the original sample. This process continues many times to produce a distribution upon which to base the estimates.
box plot
A graph that displays both the frequencies and the distribution of observations. It is useful for comparing two distributions.
box-and-whisker plot
The same as box plot.
canonical correlation analysis
An advanced statistical method for examining the relationships between two sets of interval or numerical measurements made on the same set of subjects.
case–control
An observational study that begins with patient cases who have the outcome or disease being investigated and control subjects who do not have the outcome or disease. It then looks backward to identify possible precursors or risk factors.
case–series study
A simple descriptive account of interesting or intriguing characteristics observed in a group of subjects.
categorical observation
A variable whose values are categories (an example is type of anemia). See also nominal scale.
cause-specific mortality rate
The mortality rate from a specific disease.
cell
A category of counts or value in a contingency table.
censored observation
An observation whose value is unknown, generally because the subject has not been in the study long enough for the outcome of interest, such as death, to occur.
central limit theorem
A theorem that states that the distribution of means is approximately normal if the sample size is large enough (n ≥ 30), regardless of the underlying distribution of the original measurements.
chance agreement
A measure of the proportion of times two or more raters agree in their measurement or assessment of a phenomenon.
chi-square (χ2) distribution
The distribution used to analyze counts in frequency tables.
chi-square (χ2) test
The statistical test used to test the null hypothesis that proportions are equal or, equivalently, that factors or characteristics are independent or not associated.
classes or class limits
The subdivisions of a numerical characteristic (or the widths of the classes) when it is displayed in a frequency table or graph (an example is ages by decades).
classification and regression tree (CART) analysis
A multivariate method used to detect significant relationships among variables which are then used to develop predictive models for classifying future subjects.
clinical epidemiology
The application of the science of epidemiology to clinical medicine and decision making.
clinical trial
An experimental study of a drug or procedure in which the subjects are humans.
closed question
A question on an interview or questionnaire in which a specific set of response options are provided.
cluster analysis
An advanced statistical method that determines a classification or taxonomy from multiple measures of a set of objects or subjects.
cluster random sample
A two-stage sampling process in which the population is divided into clusters, a random sample of clusters is chosen, and then random samples of subjects within the clusters are selected.
coefficient of determination (r2)
The square of the correlation coefficient. It is interpreted as the amount of variance in one variable that is accounted for by knowing the second variable.
coefficient of variation (CV)
The standard deviation divided by the mean (generally multiplied by 100). It is used to obtain a measure of relative variation.
cohort
A group of subjects who remain together in the same study over time.
cohort study
An observational study that begins with a set of subjects who have a risk factor (or have been exposed to an agent) and a second set of subjects who do not have the risk factor or exposure. Both sets are followed prospectively through time to learn how many in each set develop the outcome or consequences of interest.
combination
A formula in probability that it gives the number of ways a specific number of items, say X, can be selected from the total number of items, say n, in the entire population or sample.
complementary event
An event opposite to the event being investigated.
computer package
A set of statistical computer programs for analyzing data.
concurrent controls
Control subjects assigned to a placebo or control condition during the same period that an experimental treatment or procedure is being evaluated.
conditional probability
The probability of an event, such as A, given that another event, such as B, has occurred, denoted P(A|B).
confidence bands
Lines on each side of a regression line or curve that have a given probability of containing the line or curve in the population.
confidence coefficient
The term in the formula for a confidence interval that determines probability level associated with the interval, such as 90%, 95%, and 99%.
confidence interval (CI)
The interval computed from sample data that has a given probability that the unknown parameter, such as the mean or proportion, is contained within the interval. Common confidence intervals are 90%, 95%, and 99%.
confidence limits
The limits of a confidence interval. These limits are computed from sample data and have a given probability that the unknown parameter is located between them.
confounded
A term used to describe a study or observation that has one or more nuisance variables present that may lead to incorrect interpretations.
confounding variable
A variable more likely to be present in one group of subjects than another that is related to the outcome of interest and thus potentially confuses, or “confounds,” the results.
conservative
A term used to describe a statistical test if it reduces the chances of a type I error.
construct validity
A demonstration that the measurement of a characteristic is related to similar measures of the same characteristic and not related to measures of other characteristics.
content validity
A measure of the degree to which the items on a test or measurement scale are representative of the characteristic being measured.
contingency table
A table used to display counts or frequencies for two or more nominal or quantitative variables.
continuity correction
An adaptation to a test statistic when a continuous probability distribution is used to estimate a discrete probability distribution; eg, using the chi-square distribution for analyzing contingency tables.
continuous scale
A scale used to measure a numerical characteristic with values that occur on a continuum (an example is age).
control event rate (CER)
The number of subjects in the control group who develop the outcome being studied.
control subjects
In a clinical trial, subjects assigned to the placebo or control condition; in a case–control study, subjects without the disease or outcome.
controlled for
A term used to describe a confounding variable that is taken into consideration in the design or the analysis of the study.
controlled trial
A trial in which subjects are assigned to a control condition as well as to an experimental condition.
corrected chi-square test
A chi-square test for a 2 × 2 table that uses Yates' correction, making it more conservative.
correlation coefficient (r)
A measure of the linear relationship between two numerical measurements made on the same set of subjects. It ranges from -1 to +1, with 0 indicating no relationship. Also called the Pearson product moment.
cost–benefit analysis
A quantified methods to evaluate the trade-offs between the costs (or disadvantages) and the benefits (or advantages) of a procedure or management strategy.
cost-effectiveness analysis
A quantitative method to evaluate the cost of a procedure or management strategy that takes into account the outcome as well in order to select the lowest-cost option.
covariate
A potentially confounding variable controlled for in analysis of covariance.
Cox proportional hazard model or Cox model
A regression method used when the outcome is censored. The regression coefficients are interpreted as adjusted relative risk or odds ratios.
criterion validity
An indication of how well a test or scale predicts another related characteristic, ideally a “gold standard” if one exists.
criterion variable
The outcome (or dependent variable) that is predicted in a regression problem.
critical ratio
The term for the z score used in statistical tests.
critical region
The region (or set of values) in which a test statistic must occur for the null hypothesis to be rejected.
critical value
The value that a test statistic must exceed (in an absolute value sense) for the null hypothesis to be rejected.
crossover study
A clinical trial in which each group of subjects receives two or more treatments, but in different sequences.
cross-product ratio
See relative risk.
cross-sectional study
An observational study that examines a characteristic (or set of characteristics) in a set of subjects at one point in time; a “snap-shot” of a characteristic or condition of interest; also called survey or poll.
cross-validation
A procedure for applying the results of an analysis from one sample of subjects to a new sample of subjects to evaluate how well they generalize. It is frequently used in regression.
crude rate
A rate for the entire population that is not specific or adjusted for any given subset of the population.
cumulative frequency or percentage
In a frequency table, the frequency (or percentage) of observations having a given value plus all lower values.
curvilinear relationship (between X and Y)
A relationship that indicates that X and Y vary together, but not in constant increments.
decision analysis
A formal model for describing and analyzing a decision; also called medical decision making.
decision tree
A diagram of a set of possible actions, with their probabilities and the values of the outcomes listed. It is used to analyze a decision process.
degrees of freedom (df)
A parameter in some commonly used probability distributions; eg, the t distribution and the chi-square distribution.
dependent groups or samples
Samples in which the values in one group can be predicted from the values in the other group.
dependent variable
The variable whose values are the outcomes in a study; also called response or criterion variable.
dependent-groups t test
See paired t test.
descriptive statistics
Statistics, such as the mean, the standard deviation, the proportion, and the rate, used to describe attributes of a set of data.
dichotomous observation
A nominal measure that has only two outcomes (examples are gender: male or female; survival: yes or no); also called binary.
direct method of rate standardization
A method of adjusting rates when comparing two or more populations; it requires knowledge of the specific rates for each category in the populations to be adjusted and the frequencies in at least one population.
directional test
See one-tailed test.
discrete scale
A scale used to measure a numerical characteristic that has integer values (an example is number of pregnancies).
discriminant analysis
A regression technique for predicting a nominal outcome that has more than two values; a method used to classify subjects or objects into groups; also called discriminant function analysis.
distribution
The values of a characteristic or variable along with the frequency of their occurrence. Distributions may be based on empirical observations or may be theoretical probability distributions (eg, normal, binomial, chi-square).
distribution-free
Statistical methods that make no assumptions regarding the distribution of the observations; ie, nonparametric.
dot plot
A graphic method for displaying the frequency distribution of numerical observations for one or more groups.
double-blind trial
A clinical trial in which neither the subjects nor the investigator(s) know which treatment subjects have received.
dummy coding
A procedure in which a code of 0 or 1 is assigned to a nominal predictor variable used in regression analysis.
Dunnett's procedure
A multiple-comparison method for comparing multiple treatment groups with a single control group following a significant F test in analysis of variance.
effect or effect size
The magnitude of a difference or relationship. It is used for determining sample sizes and for combining results across studies in meta-analysis.
error mean square
(MSE) The mean square in the denominator of F in ANOVA.
error term
See residual.
estimation
The process of using information from a sample to draw conclusions about the values of parameters in a population.
event
A single outcome (or set of outcomes) from an experiment.
evidence-based medicine (EBM)
The application of the evidence based on clinical research and clinical expertise to decide optimal patient management.
expected frequencies
In contingency tables, the frequencies observed if the null hypothesis is true.
expected value
Used in decision making to denote the probability of a given outcome over the long run.
experiment
(in probability) A planned process of data collection.
experimental event rate (EER)
The number of subjects in the experimental or treatment group who develop the outcome being studied.
experimental study
A comparative study involving an intervention or manipulation. It is called a trial when human subjects are involved.
explanatory variable
See independent variable.
exponential probability distribution
A probability distribution used in models of survival or decay.
F distribution
The probability distribution used to test the equality of two estimates of the variance. It is the distribution used with the F test in ANOVA.
F test
The statistical test for comparing two variances. It is used in ANOVA.
face validity
An interview or survey that has questions on it that look related to the purpose.
factor
A characteristic that is the focus of inquiry in a study; used in analysis of variance.
factor analysis
An advanced statistical method for analyzing the relationships among a set of items or indicators to determine the factors or dimensions that underlie them.
factorial design
In ANOVA, a design in which each subject (or object) receives one level of each factor.
false-negative (FN)
A test result that is negative in a person who has the disease.
false-positive (FP)
A test result that is positive in a person who does not have the disease.
first quartile
The 25th percentile.
Fisher's exact test
An exact test for 2 × 2 contingency tables. It is used when the sample size is too small to use the chi-square test.
Fisher's z transformation
A transformation of the correlation coefficient so that it is normally distributed.
focus groups
A process in which a small group of people are interviewed about a topic or issue; often used to help generate questions for a survey, but may be used independently in qualitative research.
forward selection
A model-building method in multiple regression that first enters into the regression equation the variable with the highest correlation, followed by the other variables one at a time that increase the multiple R by the greatest amount, until all statistically significant variables are included in the equation.
frequency
The number of times a given value of an observation occurs. It is also called counts.
frequency distribution
In a set of numerical observations, the list of values that occur along with the frequency of their occurrence. It may be set up as a frequency table or as a graph.
frequency polygon
A line graph connecting the midpoints of the tops of the columns of a histogram. It is useful in comparing two frequency distributions.
frequency table
A table showing the number or percentage of observations occurring at different values (or ranges of values) of a characteristic or variable.
functional status
A measure of a person's ability to perform his or her daily activities, often called activities of daily living.
game theory
A process of assigning subjective probabilities to outcomes from a decision.
gaussian distribution
See normal distribution.
Gehan's test
A statistical test of the equality of two survival curves.
Generalized estimating equations (GEE)
A complex multivariate method used to analyze situations in which subjects are nested within groups when observations between subjects are not independent.
generalized Wilcoxon test
See Gehan's test.
geometric mean (GM)
The nth root of the product of n observations, symbolized GM or G. It is used with logarithms or skewed distributions.
gold standard
In diagnostic testing, a procedure that always identifies the true condition—diseased or disease-free of a patient.
Hawthorne effect
A bias introduced into an observational study when the subjects know they are in a study, and it is this knowledge that affects their behavior.
hazard function
The probability that a person dies in a certain time interval, given that the person has lived until the beginning of the interval. Its reciprocal is mean survival time.
hazard ratio
Similar to the risk ratio, it is the ratio of risk of the outcome (such as death) occurring at any time in one group compared with another group.
hierarchical design
A study design in which one or more of the treatments is nested within levels of another factor, such as patients within hospitals.
hierarchical regression
A logical model-building method in multiple regression in which the investigators group variables according to their function and add them to the regression equation as a group or block.
histogram
A graph of a frequency distribution of numerical observations.
historical cohort study
A cohort study that uses existing records or historical data to determine the effect of a risk factor or exposure on a group of patients.
historical controls
In clinical trials, previously collected observations on patients that are used as the control values against which the treatment is compared.
homogeneity
The situation in which the standard deviation of the dependent (Y) variable is the same, regardless of the value of the independent (X) variable; an assumption in ANOVA and regression.
homoscedasticity
See homogeneity.
Hosmer and Lemeshow's Goodness of Fit Test
A multivariate test used to test the significance of the overall results from a logistic regression analysis.
hypothesis test
An approach to statistical inference resulting in a decision to reject or not to reject the null hypothesis.
incidence
A rate giving the proportion of people who develop a given disease or condition within a specified period of time.
independent events
Events whose occurrence or outcome has no effect on the probability of the other.
independent groups or samples
Samples for which the values in one group cannot be predicted from the values in the other group.
independent observations
Observations determined at different times or by different individuals without knowledge of the value of the first observation.
independent variable
The explanatory or predictor variable in a study. It is sometimes called a factor in ANOVA.
independent-groups t test
See two-sample t test.
index of suspicion
See prior probability.
inference
(statistical) The process of drawing conclusions about a population of observations from a sample of observations.
intention-to-treat
(principle) The statistical analysis of all subjects according to the group to which they were originally assigned or belonged.
interaction
A relationship between two independent variables such that they have a different effect on the dependent variable; ie, the effect of one level of a factor a depends on the level of factor B.
intercept
In a regression equation, the predicted value of Y when X is equal to zero.
internal consistency
(reliability) The degree to which the items on an instrument or test are related to each other and provide a measure of a single characteristic.
interquartile range
The difference between the 25th percentile and the 75th percentile.
interrater reliability
The reliability between measurements made by two different persons (or raters).
intervention
The maneuver used in an experimental study. It may be a drug or a procedure.
intrarater reliability
The reliability between measurements made by the same person (or rater) at two different points in time.
jackknife
A method of cross-validation in which one observation at a time is left out of the sample; regression is performed on the remaining observations, and the results are applied to the original observation.
joint probability
The probability of two events both occurring.
Kaplan–Meier product limit method
A method for analyzing survival for censored observations. It uses exact survival times in the calculations.
kappa (κ)
A statistic used to measure interrater or intrarater agreement for nominal measures.
key concepts
Concepts and topics identified in each chapter as being the key take-home messages.
least squares regression
The most common form of regression in which the values for the intercept and regression coefficient(s) are found by minimizing the squared difference between the actual and predicted values of the outcome variable.
length of time to event
A term used in outcome and cost-effectiveness studies; it measures the length of time from a treatment or assessment until the outcome of interest occurs.
level of significance
The probability of incorrectly rejecting the null hypothesis in a test of hypothesis. Also see alpha value and P value.
Levene's test
A test of the equality of two variances. It is less sensitive to departures from normality than the F test and is often recommended by statisticians.
life table analysis
A method for analyzing survival times for censored observations that have been grouped into intervals.
likelihood
The probability of an outcome or event happening, given the parameters of the distribution of the outcome, such as the mean and standard deviation.
likelihood ratio
In diagnostic testing, the ratio of true-positives to false-positives.
linear combination
A weighted average of a set of variables or measures. For example, the prediction equation in multiple regression is a linear combination of the predictor variables.
linear regression
(of Y on X) The process of determining a regression or prediction equation to predict Y from X.
linear relationship
(between X and Y) A relationship indicating that X and Y vary together according to constant increments.
logarithm (ln)
The exponent indicating the power to which e (2.718) is raised to obtain a given number; also called the natural logarithm.
logistic regression
The regression technique used when the outcome is a binary, or dichotomous, variable.
log-linear analysis
A statistical method for analyzing the relationships among three or more nominal variables. It may be used as a regression method to predict a nominal outcome from nominal independent variables.
logrank test
A statistical method for comparing two survival curves when censored observations occur.
longitudinal study
A study that takes place over an extended period of time.
Mann-Whitney–Wilcoxon test
See Wilcoxon rank sum test.
Mantel–Haenszel chi-square test
A statistical test of two or more 2 × 2 tables. It is used to compare survival distributions or to control for confounding factors.
marginal frequencies
The row and column frequencies in a contingency table; ie, the frequencies listed on the margins of the table.
marginal probability
The row and column probabilities in a contingency table; ie, the probabilities listed on the margins of the table.
matched-groups t test
See paired t test.
matching
(or matched groups) The process of making two groups homogeneous on possible confounding factors. It is sometimes done prior to randomization in clinical trials.
McNemar's test
The chi-square test for comparing proportions from two dependent or paired groups.
mean (X̅)
The most common measure of central tendency, denoted by ľ in the population and by in the sample. In a sample, the mean is the sum of the X values divided by the number n in the sample (σX/n).
mean square among groups (MSA)
An estimate of the variation in analysis of variance. It is used in the numerator of the F statistic.
mean square within groups (MSW)
An estimate of the variation in analysis of variance. It is used in the denominator of the F statistic.
measurement error
The amount by which a measurement is incorrect because of problems inherent in the measuring process; also called bias.
measures of central tendency
Index or summary numbers that describe the middle of a distribution. See mean, median, and mode.
measures of dispersion
Index or summary numbers that describe the spread of observations about the mean. See range; standard deviation.
median (M or Md)
A measure of central tendency. It is the middle observation; ie, the one that divides the distribution of values into halves. It is also equal to the 50th percentile.
medical decision making or analysis
The application of probabilities to the decision process in medicine. It is the basis for cost-benefit analysis.
MEDLINE
A system that permits search of the bibliographic database of all articles in journals included in Index Medicus. Articles that meet specific criteria or contain specific key words are extracted for the researcher's perusal
meta-analysis
A method for combining the results from several independent studies of the same outcome so that an overall P value may be determined.
minimum variance
A desirable characteristic of a statistic that estimates a population parameter, meaning that its variance or standard deviation is less than that of another statistic; ie, the sample mean (X̅) has a smaller standard deviation than the median, although both are estimators of the population mean ľ.
modal class
The interval (generally from a frequency table or histogram) that contains the highest frequency of observations.
mode
The value of a numerical variable that occurs the most frequently.
model or modeling
A statistical statement of the relationship among variables, sometimes based upon a theoretical model.
morbidity rate
The number of patients in a defined population who develop a morbid condition over a specified period of time.
mortality rate
The number of deaths in a defined population over a specified period. It is the number of people who die during a given period of time divided by the number of people at risk during the period.
multiple comparisons
Comparisons resulting from many statistical tests performed for the same observations.
multiple R
In multiple regression, the correlation between actual and predicted values of Y (ie, rYY').
multiple regression
A multivariate method for determining a regression or prediction equation to predict an outcome from a set of independent variables.
multiple-comparison procedure
A method for comparing several means.
multiplication rule
The rule that states the probability that two or more independent events all occur is the product of the probabilities of each individual event.
multivariate
A term that refers to a study or analysis involving multiple independent or dependent variables.
multivariate analysis of variance (MANOVA)
An advanced statistical method that provides a global test when there are multiple dependent variables and the independent variables are nominal. It is analogous to analysis of variance with multiple outcome measures.
mutually exclusive events
Two or more events for which the occurrence of one event precludes the occurrence of the others.
natural log (ln)
A logarithm with the base e (e ≈ 2.718) compared with the other well-known logarithm to base 10 (log). e describes population growth patterns and is very important in logistic and Cox regression where the inverse (the antilog) of the regression coefficients are adjusted odds ratios.
Newman–Keuls procedure
A multiple-comparison method for making pairwise comparisons between means following a significant F test in analysis of variance.
nominal scale
The simplest scale of measurement. It is used for characteristics that have no numerical values (examples are race and gender). It is also called a categorical or qualitative scale.
nondirectional test
See two-tailed test.
nonmutually exclusive events
Two or more events for which the occurrence of one event does not preclude the occurrence of the others.
nonparametric method
A statistical test that makes no assumptions regarding the distribution of the observations.
nonprobability sample
A sample selected in such a way that the probability that a subject is selected is unknown.
nonrandomized trial
A clinical trial in which subjects are assigned to treatments on other than a randomized basis. It is subject to several biases.
normal distribution
A symmetric, bell-shaped probability distribution with mean ľ and standard deviation σ. If observations follow a normal distribution, the interval (ľ ą 2σ) contains 95% of the observations. It is also called the gaussian distribution.
null hypothesis
The hypothesis being tested about a population. Null generally means “no difference” and thus refers to a situation in which no difference exists (eg, between the means in a treatment group and a control group).
number needed to harm (NNH)
The number of patients that need to be treated with a proposed therapy in order to cause one undesirable outcome.
number needed to treat (NNT)
The number of patients that need to be treated with a proposed therapy in order to prevent or cure one individual; it is the reciprocal of the absolute risk reduction (1/ARR).
numerical scale
The highest level of measurement. It is used for characteristics that can be given numerical values; the differences between numbers have meaning (examples are height, weight, blood pressure level). It is also called an interval or ratio scale.
objective probability
An estimate of probability from observable events or phenomena.
observational study
A study that does not involve an intervention or manipulation. It is called case–control, cross-sectional, or cohort, depending on the design of the study.
observed frequencies
The frequencies that occur in a study. They are generally arranged in a contingency table.
odds ratio (OR)
An estimate of the relative risk calculated in case–control studies. It is the odds that a patient was exposed to a given risk factor divided by the odds that a control was exposed to the risk factor.
odds
The probability that an event will occur divided by the probability that the event will not occur; ie, odds = P/(1 – P), where P is the probability.
one-tailed test
A test in which the alternative hypothesis specifies a deviation from the null hypothesis in one direction only. The critical region is located in one end of the distribution of the test statistic. It is also called a directional test.
open-ended question
A question on an interview or questionnaire that has a fill-in-the-blank answer.
ordinal scale
Used for characteristics that have an underlying order to their values; the numbers used are arbitrary (an example is Apgar scores).
orphan P
A P value given without reference to the statistical method used to determine it.
orthogonal
Independent, nonredundant, or non overlapping, such as orthogonal comparisons in ANOVA or orthogonal factors in factor analysis.
outcome
(in an experiment) The result of an experiment or trial.
outcome assessment
The process of including quality-of-life or physical-function variables in clinical outcomes. Studies that focus on outcomes often emphasize how patients view and value their health, the care they receive, and the results or outcomes of this care.
outcome variable
The dependent or criterion variable in a study.
P value
The probability of observing a result as extreme as or more extreme than the one actually observed from chance alone (ie, if the null hypothesis is true).
paired design
See repeated-measures design.
paired t
test The statistical method for comparing the difference (or change) in a numerical variable observed for two paired (or matched) groups. It also applies to before-and-after measurements made on the same group of subjects.
parameter
The population value of a characteristic of a distribution (eg, the mean ľ).
patient satisfaction
Refers to outcome measures of patient's liking and approval of health care facilities and operations, providers, and other components of the entities that provide patient care.
percentage
A proportion multiplied by 100.
percentage polygon
A line graph connecting the midpoints of the tops of the columns of a histogram based on percentages instead of counts. It is useful in comparing two or more sets of observations when the frequencies in each group are not equal.
percentile
A number that indicates the percentage of a distribution that is less than or equal to that number.
person-years
Found by adding the length of time subjects are in a study. This concept is frequently used in epidemiology but is not recommended by statisticians because of difficulties in interpretation and analysis.
piecewise linear regression
A method used to estimate sections of a regression line when a curvilinear relationship exists between the independent and dependent variables.
placebo
A sham treatment or procedure. It is used to reduce bias in clinical studies.
point estimate
A general term for any statistic (eg, mean, standard deviation, proportion).
Poisson distribution
A probability distribution used to model the number of times a rare event occurs.
poll
A questionnaire administered to a sample of people, often about a single issue.
polynomial regression
A special case of multiple regression in which each term in the equation is a power of the independent variable X. Polynomial regression provides a way to fit a regression model to curvilinear relationships and is an alternative to transforming the data to a linear scale.
pooled standard deviation
The standard deviation used in the independent-groups t test when the standard deviations in the two groups are equal.
population
The entire collection of observations or subjects that have something in common and to which conclusions are inferred.
post hoc comparison
Method for comparing means following analysis of variance.
post hoc method
See post hoc comparison.
posterior probability
The conditional probability calculated by using Bayes' theorem. It is the predictive value of a positive test (true-positives divided by all positives) or a negative test (true-negatives divided by all negatives).
posteriori method
See post hoc comparison.
posttest odds
In diagnostic testing, the odds that a patient has a given disease or condition after a diagnostic procedure is performed and interpreted. They are similar to the predictive value of a diagnostic test.
power
The ability of a test statistic to detect a specified alternative hypothesis or difference of a specified size when the alternative hypothesis is true (ie, 1 - β, where β is the probability of a type II error). More loosely, it is the ability of a study to detect an actual effect or difference.
predictive value of a negative test (PV-)
The proportion of time that a patient with a negative diagnostic test result does not have the disease being investigated.
predictive value of a positive test (PV+)
The proportion of time that a patient with a positive diagnostic test result has the disease being investigated.
pretest odds
In diagnostic testing, the odds that a patient has a given disease or condition before a diagnostic procedure is performed and interpreted. They are similar to prior probabilities.
prevalence
The proportion of people who have a given disease or condition at a specified point in time. It is not truly a rate, although it is often incorrectly called prevalence rate.
prior probability
The unconditional probability used in the numerator of Bayes' theorem. It is the prevalence of a disease prior to performing a diagnostic procedure. Clinicians often refer to it as the index of suspicion.
probability distribution
A frequency distribution of a random variable, which may be empirical or theoretical (eg, normal, binomial).
probability sample
See random sample.
probability
The number of times an outcome occurs in the total number of trials. If a is the outcome, the probability of a is denoted P(a).
product limit method
See Kaplan–Meier product limit method.
progressively censored
A situation in which patients enter a study at different points in time and remain in the study for varying lengths of time. See censored observation.
propensity score
An advanced statistical method to control for an entire group of confounding variables.
proportion
The number of observations with the characteristic of interest divided by the total number of observations. It is used to summarize counts.
proportional hazards model
See Cox proportional hazard model.
prospective study
A study designed before data are collected.
PUBMED
See MEDLINE.
qualitative observations
Characteristics measured on a nominal scale.
quality of life (QOL)
A measure of a person's subjective assessment of the value of his or her health and functional abilities.
quantitative observations
Characteristics measured on a numerical scale; the resulting numbers have inherent meaning.
quartile
The 25th percentile or the 75th percentile, called the first and third quartiles, respectively.
random assignment
The use of random methods to assign different treatments to patients or vice versa.
random error or variation
The variation in a sample that can be expected to occur by chance.
random sample
A sample of n subjects (or objects) selected from a population so that each has a known chance of being in the sample.
random variable
A variable in a study in which subjects are randomly selected or randomly assigned to treatments.
randomization
The process of assigning subjects to different treatments (or vice versa) by using random numbers.
randomized block design
A study design used in ANOVA to help control for potential confounding.
randomized controlled trial (RCT)
An experimental study in which subjects are randomly assigned to treatment groups.
range
The difference between the largest and the smallest observation.
ranking scale
A question format on an interview or survey in which respondents are asked to rate (from 1 to …) the options listed.
rank-order scale
A scale for observations arranged according to their size, from lowest to highest or vice versa.
ranks
A set of observations arranged according to their size, from lowest to highest or vice versa.
rate
A proportion associated with a multiplier, called the base (eg, 1000, 10,000, 100,000), and computed over a specific period.
ratio
A part divided by another part. It is the number of observations with the characteristic of interest divided by the number without the characteristic.
regression
(of Y on X) The process of determining a prediction equation for predicting Y from X.
regression coefficient
The b in the simple regression equation Y = a + bX. It is sometimes interpreted as the slope of the regression line. In multiple regression, the bs are weights applied to the predictor variables.
regression toward the mean
The phenomenon in which a predicted outcome for any given person tends to be closer to the mean outcome than the person's actual observation.
relative risk (RR)
The ratio of the incidence of a given disease in exposed or at-risk persons to the incidence of the disease in unexposed persons. It is calculated in cohort or prospective studies.
relative risk reduction (RRR)
The reduction in risk with a new therapy relative to the risk without the new therapy; it is the absolute value of the difference between the experimental event rate and the control event rate divided by the control event rate (|EER – CER|/CER).
reliability
A measure of the reproducibility of a measurement. It is measured by kappa for nominal measures and by correlation for numerical measures.
repeated-measures design
A study design in which subjects are measured at more than one time. It is also called a split-plot design in ANOVA.
representative population
(or sample) A population or sample that is similar in important ways to the population to which the findings of a study are generalized.
residual
The difference between the predicted value and the actual value of the outcome (dependent) variable in regression.
response variable
See dependent variable.
retrospective cohort study
See historical cohort study.
retrospective study
A study undertaken in a post hoc manner, ie, after the observations have been made.
risk factor
A term used to designate a characteristic that is more prevalent among subjects who develop a given disease or outcome than among subjects who do not. It is generally considered to be causal.
risk ratio
See relative risk.
robust
A term used to describe a statistical method if the outcome is not affected to a large extent by a violation of the assumptions of the method.
ROC (receiver operating characteristic) curve
In diagnostic testing, a plot of the true-positives on the Y-axis versus the false-positives on the X-axis; used to evaluate the properties of a diagnostic test.
r-squared (r2)
The square of the correlation coefficient. It is interpreted as the amount of variance in one variable that is accounted for by knowing the second variable.
sample
A subset of the population.
sampled population
The population from which the sample is actually selected.
sampling distribution
(of a statistic) The frequency distribution of the statistic for many samples. It is used to make inferences about the statistic from a single sample.
sampling frame
A list of all possible subjects or objects in the population from which a random sample is to be drawn; required for some types of sampling, such as systematic sampling.
scale of measurement
The degree of precision with which a characteristic is measured. It is generally categorized into nominal (or categorical), ordinal, and numerical (or interval and ratio) scales.
scatterplot
A two-dimensional graph displaying the relationship between two numerical characteristics or variables.
Scheffé's procedure
A multiple-comparison method for comparing means following a significant F test in analysis of variance. It is the most conservative multiple-comparison method.
self-controlled study
A study in which the subjects serve as their own controls, achieved by measuring the characteristic of interest before and after an intervention.
sensitivity analysis
In decision analysis, a method for determining the way the decision changes as a function of probabilities and utilities used in the analysis.
sensitivity
The proportion of time a diagnostic test is positive in patients who have the disease or condition. A sensitive test has a low false-negative rate.
sex-specific mortality rate
A mortality rate specific to either males or females.
sign test
The nonparametric test used for testing a hypothesis about the median in a single group.
simple random sample
A random sample in which each of the n subjects (or objects) in the sample has an equal chance of being selected.
skewed distribution
A distribution in which a few outlying observations occur in one direction only. If the outlying observations are small, the distribution is skewed to the left, or negatively skewed; if they are large, the distribution is skewed to the right, or positively skewed.
slope
(of the regression line) The amount Y changes for each unit that X changes. It is designated by b in the sample.
Spearman's rank correlation (rho)
A nonparametric correlation that measures the tendency for two measurements to vary together.
specific rate
A rate that pertains to a specific group or segment of the observations (examples are age-specific mortality rate and cause-specific mortality rate).
specificity
The proportion of time that a diagnostic test is negative in patients who do not have the disease or condition. A specific test has a low false-positive rate.
standard deviation (SD)
The most common measure of dispersion or spread, denoted by σ in the population and SD or s in the sample. It can be used with the mean to describe the distribution of observations. It is the square root of the average of the squared deviations of the observations from their mean.
standard error (SE)
The standard deviation of the sampling distribution of a statistic.
standard error of the estimate (Sy.x)
A measure of the variation in a regression line. It is based on the differences between the predicted and actual values of the dependent variable Y.
standard error of the mean (SEM)
The standard deviation of the mean in a large number of samples.
standard normal distribution
The normal distribution with mean 0 and standard deviation 1, also called the z distribution.
standardized mortality ratio
The number of observed deaths divided by the number of expected deaths.
standardized regression coefficient
A regression coefficient that has the effect of the measurement scale removed so that the size of the coefficient can be interpreted.
statistic
A summary number for a sample (eg, the mean), often used as an estimate of a parameter in the population.
statistical significance
Generally interpreted as a result that would occur by chance, eg, 1 time in 20, with a P value less than or equal to 0.05. It occurs when the null hypothesis is rejected.
statistical test
The procedure used to test a null hypothesis (eg, t test, chi-square test).
stem-and-leaf plot
A graphic display for numerical data. It is similar to both a frequency table and a histogram.
stepwise regression
In multiple regression, a sequential method of selecting the variables to be included in the prediction equation.
stratified random sample
A sample consisting of random samples from each subpopulation (or stratum) in a population. It is used so that the investigator can be sure that each subpopulation is appropriately represented in the sample.
structured abstract
A journal article abstract that contains short, precise descriptions of the context of the study, the objective, design, the setting and participants, the methods or interventions, main outcomes, results, and conclusions.
subjective probability
An estimate of probability that reflects a person's opinion or best guess from previous experience.
sums of squares (SS)
Quantities calculated in analysis of variance and used to obtain the mean squares for the F test.
suppression of zero
A term used to describe a misleading graph that does not have a break (a jagged line) in the Y-axis to indicate that part of the scale is missing.
survey
An observational study that generally has a cross-sectional design; a commonly used design to collect opinions.
survival analysis
The statistical method for analyzing survival data when there are censored observations.
symbols
Often Greek letters that stand for the population parameters and the Latin letters that stand for the sample statistics.
symmetric distribution
A distribution that has the same shape on both sides of the mean. The mean, median, and mode are all equal. It is the opposite of a skewed distribution.
systematic error
A measurement error that is the same (or constant) over all observations. See also bias.
systematic random sample
A random sample obtained by selecting each kth subject or object.
t distribution
A symmetric distribution with mean zero and a standard deviation larger than that for the normal distribution for small sample sizes. As nincreases, the t distribution approaches the normal distribution.
t test
The statistical test for comparing a mean with a norm or for comparing two means with small sample sizes (n ≤ 30). It is also used for testing whether a correlation coefficient or a regression coefficient is zero.
target population
The population to which the investigator wishes to generalize.
test statistic
The specific statistic used to test the null hypothesis (eg, the t statistic or chi-square statistic).
testing threshold
In diagnostic testing, the point at which the optimal decision is to perform a diagnostic test.
test–retest reliability
A measure of the degree to which an instrument or test provides a consistent measure of a characteristic on different occasions.
third quartile
The 75th percentile.
threshold model
A model for deciding when a diagnostic test should be ordered, as opposed to doing nothing or treating the patient without performing the test.
transformation
A change in the scale for the values of a variable.
treatment threshold
In diagnostic testing, the point at which the optimal decision is to treat the patient without first performing a diagnostic test.
trial
An experiment involving humans, commonly called a clinical trial. It is also a replication (repetition) of an experiment.
true-negative (TN)
A test result that is negative in a person who does not have the disease.
true-positive (TP)
A test result that is positive in a person who has the disease.
Tukey's HSD (honestly significant difference) test
A post hoc test for making multiple pairwise comparisons between means following a significant F test in analysis of variance. It is a method highly recommended by statisticians.
two-sample t test
The statistical test used to test the null hypothesis that two independent (or unrelated) groups have the same mean.
two-tailed test
A test in which the alternative hypothesis specifies a deviation from the null hypothesis in either direction. The critical region is located in both ends of the distribution of the test statistic. It is also called a directional test.
two-way analysis of variance
ANOVA with two independent variables.
type I error
The error that results if a true null hypothesis is rejected or if a difference is concluded when no difference exists.
type II error
The error that results if a false null hypothesis is not rejected or if a difference is not detected when a difference exists.
unbiasedness
(of a statistic) A term used to describe a statistic whose mean based on a large number of samples is equal to the population parameter.
uncontrolled study
An experimental study that has no control subjects.
utility
The value of different outcomes in a decision tree.
validity
The property of a measurement that indicates how well it measures the characteristic.
variable
A characteristic of interest in a study that has different values for different subjects or objects.
variance
(σ2 in the population, s2 in the sample) The square of the standard deviation.
variation (within subject)
The variability in measurements of the same object or subject. It may occur naturally or may represent an error.
vital statistics
Mortality and morbidity rates used in epidemiology and public health.
weighted average
A number formed by multiplying each number in a set of numbers by a value called a weight, adding the resulting products, and then dividing by the sum of the weight.
Wilcoxon rank sum test
A nonparametric test for comparing two independent samples with ordinal data or with numerical observations that are not normally distributed.
Wilcoxon signed ranks test
A nonparametric test for comparing two dependent samples with ordinal data or with numerical observations that are not normally distributed.
Yates' correction
The process of subtracting 0.5 from the numerator at each term in the chi-square statistic for 2 × 2 tables prior to squaring the term.
z approximation
(to the binomial) The z test used to test the equality of two independent proportions.
z distribution
The normal distribution with mean 0 and standard deviation 1. It is also called the standard normal distribution.
z ratio
The test statistic used in the z test. It is formed by subtracting the hypothesized mean from the observed mean and dividing by the standard error of the mean.
z score
The deviation of X from the mean divided by the standard deviation.
z test
The statistical test for comparing a mean with a norm or comparing two means for large samples (n ≥ 30).
z transformation
A transformation that changes a normally distributed variable with mean and standard deviation SD to the z distribution with mean 0 and standard deviation 1.
Editors: Dawson, Beth; Trapp, Robert G.
Title: Basic & Clinical Biostatistics, 4th Edition
Copyright Š2004 McGraw-Hill
> Back of Book > Symbol
Symbol
=
equals
≠
does not equal
<
less than
≤
less than or equal to
>
greater than
≥
greater than or equal to
√a
square root of a
|a|
absolute (or positive) value of a
P(A)
probability of event A
P(A|B)
probability of event A given that event B has happened
H0
null hypothesis
H1
alternative hypothesis
α
Greek letter alpha; probability of type I error
β
Greek letter beta; probability of type II error; also, population value of the slope of the regression line
δ
Greek letter delta; population mean difference between two paired measurements
κ
Greek letter kappa; used to denote an index of agreement or reproducibility
λ
Greek letter lambda; used to denote terms in the log-linear model
ľ
Greek letter mu; population mean
π
Greek letter pi; population proportion
ρ
Greek letter rho; population correlation
σ
Greek lowercase letter sigma; population standard deviation
Σ
Greek uppercase letter sigma; symbol indicating a sum
CV
coefficient of variation
df
degree of freedom
e
base of natural logarithm (equal to 2.718)
E
expected frequency
F
symbol for the F test and distribution
H
hazard function
LR
likelihood ratio
MSA
mean square among groups
MSE
error mean square
n
sample size
O
observed frequency
r
sample correlation
r2
squared correlation, called of the coefficient of determination
R2
squared multiple correlation in multiple regression
SD
sample standard deviation
SE
standard error of the mean
sY.X
standard error of the estimate in regression
t
symbol for the t ratio (the critical ratio that follows a t distribution)
X
independent (explanatory, predictor) variable in regression
χ2
symbol for the chi-square test
X̅
x bar; sample mean
Y
dependent (outcome, response, criterion) variable in regression
Y′
predicted value of Y regression
z
symbol for the z ratio (the critical ratio that follows a z or standard normal distribution