Basic & Clinical Biostatistics, 4th Edition

Glossary

absolute risk increase (ARI)

The increase in risk with a new therapy compared with the risk without the new therapy.

absolute risk reduction (ARR)

The reduction in risk with a new therapy compared with the risk without the new therapy; it is the absolute value of the difference between the experimental event rate and the control event rate (|EER – CER|).

absolute value

The positive value of a number, regardless of whether the number is positive or negative. The absolute value of a is symbolized |a|.

actuarial analysis

See life table analysis.

addition rule

The rule which states the probability that two or more mutually exclusive events all occur is the sum of the probabilities of each individual event.

adjusted rate

A rate adjusted so that it is independent of the distribution of a possible confounding variable. For example, age-adjusted rates are independent of the age distribution in the population to which they apply.

age-specific mortality rate

The mortality rate in a specific age group.

alpha (α) error

See type I error.

alpha (α) value

The level of alpha (α) selected in a hypothesis test.

alternative hypothesis

The opposite of the null hypothesis. It is the conclusion when the null hypothesis is rejected.

analysis of covariance (ANCOVA)

A special type of analysis of variance or regression used to control for the effect of a possible confounding factor.

analysis of residuals

In regression, an analysis of the differences between Y and Y' to evaluate assumptions and provide guidance on how well the equation fits the data.

analysis of variance (ANOVA)

A statistical procedure that determines whether any differences exist among two or more groups of subjects on one or more factors. The Ftest is used in ANOVA.

backward elimination

A method to select variables in multiple regression that enters all variables into the regression equation and then eliminates the variable that adds the least to the prediction, followed by the other variables one at a time that decrease the multiple R by the least amount until all statistically significant variables are removed from the equation.

bar chart or bar graph

A chart or graph used with nominal characteristics to display the numbers or percentages of observations with the characteristic of interest.

Bayes' theorem

A formula for calculating the conditional probability of one event, P(A|B), from the conditional probability of the other event, P(B|a).

bell-shaped distribution

A term used to describe the shape of the normal (gaussian) distribution.

beta (β) error

See type II error.

bias

The error related to the ways the targeted and sampled populations differ; also called measurement error, it threatens the validity of a study.

binary observation

A nominal measure that has only two outcomes (examples are gender: male or female; survival: yes or no).

binomial distribution

The probability distribution that describes the number of successes X observed in n independent trials, each with the same probability of occurrence.

biometrics

The study of measurement and statistical analysis in medicine and biology.

biostatistics

The application of research study design and statistical analysis to applications in medicine and biology.

bivariate plot

A two-dimensional plot or scatterplot of the values of two characteristics measured on the same set of subjects.

blind study

An experimental study in which subjects do not know the treatment they are receiving; investigators may also be blind to the treatment subjects are receiving; see also double-blind trial.

block design

In analysis of variance, a design in which subjects within each block (or stratum) are assigned to a different treatment.

P.404

Bonferroni t

A method for comparing means in analysis of variance; also called the Dunn multiple-comparison procedure.

bootstrap

A method for estimating standard errors or confidence intervals in which a small sample of observations is randomly selected from the original sample, estimates are calculated, and the sample is returned to the original sample. This process continues many times to produce a distribution upon which to base the estimates.

box plot

A graph that displays both the frequencies and the distribution of observations. It is useful for comparing two distributions.

box-and-whisker plot

The same as box plot.

canonical correlation analysis

An advanced statistical method for examining the relationships between two sets of interval or numerical measurements made on the same set of subjects.

case–control

An observational study that begins with patient cases who have the outcome or disease being investigated and control subjects who do not have the outcome or disease. It then looks backward to identify possible precursors or risk factors.

case–series study

A simple descriptive account of interesting or intriguing characteristics observed in a group of subjects.

categorical observation

A variable whose values are categories (an example is type of anemia). See also nominal scale.

cause-specific mortality rate

The mortality rate from a specific disease.

cell

A category of counts or value in a contingency table.

censored observation

An observation whose value is unknown, generally because the subject has not been in the study long enough for the outcome of interest, such as death, to occur.

central limit theorem

A theorem that states that the distribution of means is approximately normal if the sample size is large enough (n ≥ 30), regardless of the underlying distribution of the original measurements.

chance agreement

A measure of the proportion of times two or more raters agree in their measurement or assessment of a phenomenon.

chi-square (χ2) distribution

The distribution used to analyze counts in frequency tables.

chi-square (χ2) test

The statistical test used to test the null hypothesis that proportions are equal or, equivalently, that factors or characteristics are independent or not associated.

classes or class limits

The subdivisions of a numerical characteristic (or the widths of the classes) when it is displayed in a frequency table or graph (an example is ages by decades).

classification and regression tree (CART) analysis

A multivariate method used to detect significant relationships among variables which are then used to develop predictive models for classifying future subjects.

clinical epidemiology

The application of the science of epidemiology to clinical medicine and decision making.

clinical trial

An experimental study of a drug or procedure in which the subjects are humans.

closed question

A question on an interview or questionnaire in which a specific set of response options are provided.

cluster analysis

An advanced statistical method that determines a classification or taxonomy from multiple measures of a set of objects or subjects.

cluster random sample

A two-stage sampling process in which the population is divided into clusters, a random sample of clusters is chosen, and then random samples of subjects within the clusters are selected.

coefficient of determination (r2)

The square of the correlation coefficient. It is interpreted as the amount of variance in one variable that is accounted for by knowing the second variable.

coefficient of variation (CV)

The standard deviation divided by the mean (generally multiplied by 100). It is used to obtain a measure of relative variation.

cohort

A group of subjects who remain together in the same study over time.

cohort study

An observational study that begins with a set of subjects who have a risk factor (or have been exposed to an agent) and a second set of subjects who do not have the risk factor or exposure. Both sets are followed prospectively through time to learn how many in each set develop the outcome or consequences of interest.

combination

A formula in probability that it gives the number of ways a specific number of items, say X, can be selected from the total number of items, say n, in the entire population or sample.

complementary event

An event opposite to the event being investigated.

computer package

A set of statistical computer programs for analyzing data.

concurrent controls

Control subjects assigned to a placebo or control condition during the same period that an experimental treatment or procedure is being evaluated.

conditional probability

The probability of an event, such as A, given that another event, such as B, has occurred, denoted P(A|B).

confidence bands

Lines on each side of a regression line or curve that have a given probability of containing the line or curve in the population.

confidence coefficient

The term in the formula for a confidence interval that determines probability level associated with the interval, such as 90%, 95%, and 99%.

confidence interval (CI)

The interval computed from sample data that has a given probability that the unknown parameter, such as the mean or proportion, is contained within the interval. Common confidence intervals are 90%, 95%, and 99%.

confidence limits

The limits of a confidence interval. These limits are computed from sample data and have a given probability that the unknown parameter is located between them.

confounded

A term used to describe a study or observation that has one or more nuisance variables present that may lead to incorrect interpretations.

confounding variable

A variable more likely to be present in one group of subjects than another that is related to the outcome of interest and thus potentially confuses, or “confounds,” the results.

conservative

A term used to describe a statistical test if it reduces the chances of a type I error.

construct validity

A demonstration that the measurement of a characteristic is related to similar measures of the same characteristic and not related to measures of other characteristics.

content validity

A measure of the degree to which the items on a test or measurement scale are representative of the characteristic being measured.

contingency table

A table used to display counts or frequencies for two or more nominal or quantitative variables.

continuity correction

An adaptation to a test statistic when a continuous probability distribution is used to estimate a discrete probability distribution; eg, using the chi-square distribution for analyzing contingency tables.

continuous scale

A scale used to measure a numerical characteristic with values that occur on a continuum (an example is age).

control event rate (CER)

The number of subjects in the control group who develop the outcome being studied.

control subjects

In a clinical trial, subjects assigned to the placebo or control condition; in a case–control study, subjects without the disease or outcome.

controlled for

A term used to describe a confounding variable that is taken into consideration in the design or the analysis of the study.

controlled trial

A trial in which subjects are assigned to a control condition as well as to an experimental condition.

corrected chi-square test

A chi-square test for a 2 × 2 table that uses Yates' correction, making it more conservative.

correlation coefficient (r)

A measure of the linear relationship between two numerical measurements made on the same set of subjects. It ranges from -1 to +1, with 0 indicating no relationship. Also called the Pearson product moment.

cost–benefit analysis

A quantified methods to evaluate the trade-offs between the costs (or disadvantages) and the benefits (or advantages) of a procedure or management strategy.

cost-effectiveness analysis

A quantitative method to evaluate the cost of a procedure or management strategy that takes into account the outcome as well in order to select the lowest-cost option.

covariate

A potentially confounding variable controlled for in analysis of covariance.

Cox proportional hazard model or Cox model

A regression method used when the outcome is censored. The regression coefficients are interpreted as adjusted relative risk or odds ratios.

criterion validity

An indication of how well a test or scale predicts another related characteristic, ideally a “gold standard” if one exists.

criterion variable

The outcome (or dependent variable) that is predicted in a regression problem.

critical ratio

The term for the z score used in statistical tests.

critical region

The region (or set of values) in which a test statistic must occur for the null hypothesis to be rejected.

critical value

The value that a test statistic must exceed (in an absolute value sense) for the null hypothesis to be rejected.

crossover study

A clinical trial in which each group of subjects receives two or more treatments, but in different sequences.

cross-product ratio

See relative risk.

cross-sectional study

An observational study that examines a characteristic (or set of characteristics) in a set of subjects at one point in time; a “snap-shot” of a characteristic or condition of interest; also called survey or poll.

cross-validation

A procedure for applying the results of an analysis from one sample of subjects to a new sample of subjects to evaluate how well they generalize. It is frequently used in regression.

crude rate

A rate for the entire population that is not specific or adjusted for any given subset of the population.

cumulative frequency or percentage

In a frequency table, the frequency (or percentage) of observations having a given value plus all lower values.

curvilinear relationship (between X and Y)

A relationship that indicates that X and Y vary together, but not in constant increments.

decision analysis

A formal model for describing and analyzing a decision; also called medical decision making.

decision tree

A diagram of a set of possible actions, with their probabilities and the values of the outcomes listed. It is used to analyze a decision process.

degrees of freedom (df)

A parameter in some commonly used probability distributions; eg, the t distribution and the chi-square distribution.

dependent groups or samples

Samples in which the values in one group can be predicted from the values in the other group.

dependent variable

The variable whose values are the outcomes in a study; also called response or criterion variable.

dependent-groups t test

See paired t test.

descriptive statistics

Statistics, such as the mean, the standard deviation, the proportion, and the rate, used to describe attributes of a set of data.

dichotomous observation

A nominal measure that has only two outcomes (examples are gender: male or female; survival: yes or no); also called binary.

direct method of rate standardization

A method of adjusting rates when comparing two or more populations; it requires knowledge of the specific rates for each category in the populations to be adjusted and the frequencies in at least one population.

directional test

See one-tailed test.

discrete scale

A scale used to measure a numerical characteristic that has integer values (an example is number of pregnancies).

discriminant analysis

A regression technique for predicting a nominal outcome that has more than two values; a method used to classify subjects or objects into groups; also called discriminant function analysis.

distribution

The values of a characteristic or variable along with the frequency of their occurrence. Distributions may be based on empirical observations or may be theoretical probability distributions (eg, normal, binomial, chi-square).

distribution-free

Statistical methods that make no assumptions regarding the distribution of the observations; ie, nonparametric.

dot plot

A graphic method for displaying the frequency distribution of numerical observations for one or more groups.

double-blind trial

A clinical trial in which neither the subjects nor the investigator(s) know which treatment subjects have received.

dummy coding

A procedure in which a code of 0 or 1 is assigned to a nominal predictor variable used in regression analysis.

Dunnett's procedure

A multiple-comparison method for comparing multiple treatment groups with a single control group following a significant F test in analysis of variance.

effect or effect size

The magnitude of a difference or relationship. It is used for determining sample sizes and for combining results across studies in meta-analysis.

error mean square

(MSE) The mean square in the denominator of F in ANOVA.

error term

See residual.

estimation

The process of using information from a sample to draw conclusions about the values of parameters in a population.

event

A single outcome (or set of outcomes) from an experiment.

evidence-based medicine (EBM)

The application of the evidence based on clinical research and clinical expertise to decide optimal patient management.

expected frequencies

In contingency tables, the frequencies observed if the null hypothesis is true.

expected value

Used in decision making to denote the probability of a given outcome over the long run.

experiment

(in probability) A planned process of data collection.

experimental event rate (EER)

The number of subjects in the experimental or treatment group who develop the outcome being studied.

experimental study

A comparative study involving an intervention or manipulation. It is called a trial when human subjects are involved.

explanatory variable

See independent variable.

exponential probability distribution

A probability distribution used in models of survival or decay.

F distribution

The probability distribution used to test the equality of two estimates of the variance. It is the distribution used with the F test in ANOVA.

F test

The statistical test for comparing two variances. It is used in ANOVA.

face validity

An interview or survey that has questions on it that look related to the purpose.

factor

A characteristic that is the focus of inquiry in a study; used in analysis of variance.

factor analysis

An advanced statistical method for analyzing the relationships among a set of items or indicators to determine the factors or dimensions that underlie them.

factorial design

In ANOVA, a design in which each subject (or object) receives one level of each factor.

false-negative (FN)

A test result that is negative in a person who has the disease.

false-positive (FP)

A test result that is positive in a person who does not have the disease.

first quartile

The 25th percentile.

Fisher's exact test

An exact test for 2 × 2 contingency tables. It is used when the sample size is too small to use the chi-square test.

Fisher's z transformation

A transformation of the correlation coefficient so that it is normally distributed.

focus groups

A process in which a small group of people are interviewed about a topic or issue; often used to help generate questions for a survey, but may be used independently in qualitative research.

forward selection

A model-building method in multiple regression that first enters into the regression equation the variable with the highest correlation, followed by the other variables one at a time that increase the multiple R by the greatest amount, until all statistically significant variables are included in the equation.

frequency

The number of times a given value of an observation occurs. It is also called counts.

frequency distribution

In a set of numerical observations, the list of values that occur along with the frequency of their occurrence. It may be set up as a frequency table or as a graph.

frequency polygon

A line graph connecting the midpoints of the tops of the columns of a histogram. It is useful in comparing two frequency distributions.

frequency table

A table showing the number or percentage of observations occurring at different values (or ranges of values) of a characteristic or variable.

functional status

A measure of a person's ability to perform his or her daily activities, often called activities of daily living.

game theory

A process of assigning subjective probabilities to outcomes from a decision.

gaussian distribution

See normal distribution.

Gehan's test

A statistical test of the equality of two survival curves.

Generalized estimating equations (GEE)

A complex multivariate method used to analyze situations in which subjects are nested within groups when observations between subjects are not independent.

generalized Wilcoxon test

See Gehan's test.

geometric mean (GM)

The nth root of the product of n observations, symbolized GM or G. It is used with logarithms or skewed distributions.

gold standard

In diagnostic testing, a procedure that always identifies the true condition—diseased or disease-free of a patient.

Hawthorne effect

A bias introduced into an observational study when the subjects know they are in a study, and it is this knowledge that affects their behavior.

hazard function

The probability that a person dies in a certain time interval, given that the person has lived until the beginning of the interval. Its reciprocal is mean survival time.

hazard ratio

Similar to the risk ratio, it is the ratio of risk of the outcome (such as death) occurring at any time in one group compared with another group.

hierarchical design

A study design in which one or more of the treatments is nested within levels of another factor, such as patients within hospitals.

hierarchical regression

A logical model-building method in multiple regression in which the investigators group variables according to their function and add them to the regression equation as a group or block.

histogram

A graph of a frequency distribution of numerical observations.

historical cohort study

A cohort study that uses existing records or historical data to determine the effect of a risk factor or exposure on a group of patients.

historical controls

In clinical trials, previously collected observations on patients that are used as the control values against which the treatment is compared.

homogeneity

The situation in which the standard deviation of the dependent (Y) variable is the same, regardless of the value of the independent (X) variable; an assumption in ANOVA and regression.

homoscedasticity

See homogeneity.

Hosmer and Lemeshow's Goodness of Fit Test

A multivariate test used to test the significance of the overall results from a logistic regression analysis.

hypothesis test

An approach to statistical inference resulting in a decision to reject or not to reject the null hypothesis.

incidence

A rate giving the proportion of people who develop a given disease or condition within a specified period of time.

independent events

Events whose occurrence or outcome has no effect on the probability of the other.

independent groups or samples

Samples for which the values in one group cannot be predicted from the values in the other group.

independent observations

Observations determined at different times or by different individuals without knowledge of the value of the first observation.

independent variable

The explanatory or predictor variable in a study. It is sometimes called a factor in ANOVA.

independent-groups t test

See two-sample t test.

index of suspicion

See prior probability.

inference

(statistical) The process of drawing conclusions about a population of observations from a sample of observations.

intention-to-treat

(principle) The statistical analysis of all subjects according to the group to which they were originally assigned or belonged.

interaction

A relationship between two independent variables such that they have a different effect on the dependent variable; ie, the effect of one level of a factor a depends on the level of factor B.

intercept

In a regression equation, the predicted value of Y when X is equal to zero.

internal consistency

(reliability) The degree to which the items on an instrument or test are related to each other and provide a measure of a single characteristic.

interquartile range

The difference between the 25th percentile and the 75th percentile.

interrater reliability

The reliability between measurements made by two different persons (or raters).

intervention

The maneuver used in an experimental study. It may be a drug or a procedure.

intrarater reliability

The reliability between measurements made by the same person (or rater) at two different points in time.

jackknife

A method of cross-validation in which one observation at a time is left out of the sample; regression is performed on the remaining observations, and the results are applied to the original observation.

joint probability

The probability of two events both occurring.

Kaplan–Meier product limit method

A method for analyzing survival for censored observations. It uses exact survival times in the calculations.

kappa (κ)

A statistic used to measure interrater or intrarater agreement for nominal measures.

key concepts

Concepts and topics identified in each chapter as being the key take-home messages.

least squares regression

The most common form of regression in which the values for the intercept and regression coefficient(s) are found by minimizing the squared difference between the actual and predicted values of the outcome variable.

length of time to event

A term used in outcome and cost-effectiveness studies; it measures the length of time from a treatment or assessment until the outcome of interest occurs.

level of significance

The probability of incorrectly rejecting the null hypothesis in a test of hypothesis. Also see alpha value and P value.

Levene's test

A test of the equality of two variances. It is less sensitive to departures from normality than the F test and is often recommended by statisticians.

life table analysis

A method for analyzing survival times for censored observations that have been grouped into intervals.

likelihood

The probability of an outcome or event happening, given the parameters of the distribution of the outcome, such as the mean and standard deviation.

likelihood ratio

In diagnostic testing, the ratio of true-positives to false-positives.

linear combination

A weighted average of a set of variables or measures. For example, the prediction equation in multiple regression is a linear combination of the predictor variables.

linear regression

(of Y on X) The process of determining a regression or prediction equation to predict Y from X.

linear relationship

(between X and Y) A relationship indicating that X and Y vary together according to constant increments.

logarithm (ln)

The exponent indicating the power to which e (2.718) is raised to obtain a given number; also called the natural logarithm.

logistic regression

The regression technique used when the outcome is a binary, or dichotomous, variable.

log-linear analysis

A statistical method for analyzing the relationships among three or more nominal variables. It may be used as a regression method to predict a nominal outcome from nominal independent variables.

logrank test

A statistical method for comparing two survival curves when censored observations occur.

longitudinal study

A study that takes place over an extended period of time.

Mann-Whitney–Wilcoxon test

See Wilcoxon rank sum test.

Mantel–Haenszel chi-square test

A statistical test of two or more 2 × 2 tables. It is used to compare survival distributions or to control for confounding factors.

marginal frequencies

The row and column frequencies in a contingency table; ie, the frequencies listed on the margins of the table.

marginal probability

The row and column probabilities in a contingency table; ie, the probabilities listed on the margins of the table.

matched-groups t test

See paired t test.

matching

(or matched groups) The process of making two groups homogeneous on possible confounding factors. It is sometimes done prior to randomization in clinical trials.

McNemar's test

The chi-square test for comparing proportions from two dependent or paired groups.

mean (X̅)

The most common measure of central tendency, denoted by ľ in the population and by in the sample. In a sample, the mean is the sum of the X values divided by the number n in the sample (σX/n).

mean square among groups (MSA)

An estimate of the variation in analysis of variance. It is used in the numerator of the F statistic.

mean square within groups (MSW)

An estimate of the variation in analysis of variance. It is used in the denominator of the F statistic.

measurement error

The amount by which a measurement is incorrect because of problems inherent in the measuring process; also called bias.

measures of central tendency

Index or summary numbers that describe the middle of a distribution. See mean, median, and mode.

measures of dispersion

Index or summary numbers that describe the spread of observations about the mean. See range; standard deviation.

median (M or Md)

A measure of central tendency. It is the middle observation; ie, the one that divides the distribution of values into halves. It is also equal to the 50th percentile.

medical decision making or analysis

The application of probabilities to the decision process in medicine. It is the basis for cost-benefit analysis.

MEDLINE

A system that permits search of the bibliographic database of all articles in journals included in Index Medicus. Articles that meet specific criteria or contain specific key words are extracted for the researcher's perusal

meta-analysis

A method for combining the results from several independent studies of the same outcome so that an overall P value may be determined.

minimum variance

A desirable characteristic of a statistic that estimates a population parameter, meaning that its variance or standard deviation is less than that of another statistic; ie, the sample mean (X̅) has a smaller standard deviation than the median, although both are estimators of the population mean ľ.

modal class

The interval (generally from a frequency table or histogram) that contains the highest frequency of observations.

mode

The value of a numerical variable that occurs the most frequently.

model or modeling

A statistical statement of the relationship among variables, sometimes based upon a theoretical model.

morbidity rate

The number of patients in a defined population who develop a morbid condition over a specified period of time.

mortality rate

The number of deaths in a defined population over a specified period. It is the number of people who die during a given period of time divided by the number of people at risk during the period.

multiple comparisons

Comparisons resulting from many statistical tests performed for the same observations.

multiple R

In multiple regression, the correlation between actual and predicted values of Y (ie, rYY').

multiple regression

A multivariate method for determining a regression or prediction equation to predict an outcome from a set of independent variables.

multiple-comparison procedure

A method for comparing several means.

multiplication rule

The rule that states the probability that two or more independent events all occur is the product of the probabilities of each individual event.

multivariate

A term that refers to a study or analysis involving multiple independent or dependent variables.

multivariate analysis of variance (MANOVA)

An advanced statistical method that provides a global test when there are multiple dependent variables and the independent variables are nominal. It is analogous to analysis of variance with multiple outcome measures.

mutually exclusive events

Two or more events for which the occurrence of one event precludes the occurrence of the others.

natural log (ln)

A logarithm with the base e (e ≈ 2.718) compared with the other well-known logarithm to base 10 (log). e describes population growth patterns and is very important in logistic and Cox regression where the inverse (the antilog) of the regression coefficients are adjusted odds ratios.

Newman–Keuls procedure

A multiple-comparison method for making pairwise comparisons between means following a significant F test in analysis of variance.

nominal scale

The simplest scale of measurement. It is used for characteristics that have no numerical values (examples are race and gender). It is also called a categorical or qualitative scale.

nondirectional test

See two-tailed test.

nonmutually exclusive events

Two or more events for which the occurrence of one event does not preclude the occurrence of the others.

nonparametric method

A statistical test that makes no assumptions regarding the distribution of the observations.

nonprobability sample

A sample selected in such a way that the probability that a subject is selected is unknown.

nonrandomized trial

A clinical trial in which subjects are assigned to treatments on other than a randomized basis. It is subject to several biases.

normal distribution

A symmetric, bell-shaped probability distribution with mean ľ and standard deviation σ. If observations follow a normal distribution, the interval (ľ ą 2σ) contains 95% of the observations. It is also called the gaussian distribution.

null hypothesis

The hypothesis being tested about a population. Null generally means “no difference” and thus refers to a situation in which no difference exists (eg, between the means in a treatment group and a control group).

number needed to harm (NNH)

The number of patients that need to be treated with a proposed therapy in order to cause one undesirable outcome.

number needed to treat (NNT)

The number of patients that need to be treated with a proposed therapy in order to prevent or cure one individual; it is the reciprocal of the absolute risk reduction (1/ARR).

numerical scale

The highest level of measurement. It is used for characteristics that can be given numerical values; the differences between numbers have meaning (examples are height, weight, blood pressure level). It is also called an interval or ratio scale.

objective probability

An estimate of probability from observable events or phenomena.

observational study

A study that does not involve an intervention or manipulation. It is called case–control, cross-sectional, or cohort, depending on the design of the study.

observed frequencies

The frequencies that occur in a study. They are generally arranged in a contingency table.

odds ratio (OR)

An estimate of the relative risk calculated in case–control studies. It is the odds that a patient was exposed to a given risk factor divided by the odds that a control was exposed to the risk factor.

odds

The probability that an event will occur divided by the probability that the event will not occur; ie, odds = P/(1 – P), where P is the probability.

one-tailed test

A test in which the alternative hypothesis specifies a deviation from the null hypothesis in one direction only. The critical region is located in one end of the distribution of the test statistic. It is also called a directional test.

open-ended question

A question on an interview or questionnaire that has a fill-in-the-blank answer.

ordinal scale

Used for characteristics that have an underlying order to their values; the numbers used are arbitrary (an example is Apgar scores).

orphan P

A P value given without reference to the statistical method used to determine it.

orthogonal

Independent, nonredundant, or non overlapping, such as orthogonal comparisons in ANOVA or orthogonal factors in factor analysis.

outcome

(in an experiment) The result of an experiment or trial.

outcome assessment

The process of including quality-of-life or physical-function variables in clinical outcomes. Studies that focus on outcomes often emphasize how patients view and value their health, the care they receive, and the results or outcomes of this care.

outcome variable

The dependent or criterion variable in a study.

P value

The probability of observing a result as extreme as or more extreme than the one actually observed from chance alone (ie, if the null hypothesis is true).

paired design

See repeated-measures design.

paired t

test The statistical method for comparing the difference (or change) in a numerical variable observed for two paired (or matched) groups. It also applies to before-and-after measurements made on the same group of subjects.

parameter

The population value of a characteristic of a distribution (eg, the mean ľ).

patient satisfaction

Refers to outcome measures of patient's liking and approval of health care facilities and operations, providers, and other components of the entities that provide patient care.

percentage

A proportion multiplied by 100.

percentage polygon

A line graph connecting the midpoints of the tops of the columns of a histogram based on percentages instead of counts. It is useful in comparing two or more sets of observations when the frequencies in each group are not equal.

percentile

A number that indicates the percentage of a distribution that is less than or equal to that number.

person-years

Found by adding the length of time subjects are in a study. This concept is frequently used in epidemiology but is not recommended by statisticians because of difficulties in interpretation and analysis.

piecewise linear regression

A method used to estimate sections of a regression line when a curvilinear relationship exists between the independent and dependent variables.

placebo

A sham treatment or procedure. It is used to reduce bias in clinical studies.

point estimate

A general term for any statistic (eg, mean, standard deviation, proportion).

Poisson distribution

A probability distribution used to model the number of times a rare event occurs.

poll

A questionnaire administered to a sample of people, often about a single issue.

polynomial regression

A special case of multiple regression in which each term in the equation is a power of the independent variable X. Polynomial regression provides a way to fit a regression model to curvilinear relationships and is an alternative to transforming the data to a linear scale.

pooled standard deviation

The standard deviation used in the independent-groups t test when the standard deviations in the two groups are equal.

population

The entire collection of observations or subjects that have something in common and to which conclusions are inferred.

post hoc comparison

Method for comparing means following analysis of variance.

post hoc method

See post hoc comparison.

posterior probability

The conditional probability calculated by using Bayes' theorem. It is the predictive value of a positive test (true-positives divided by all positives) or a negative test (true-negatives divided by all negatives).

posteriori method

See post hoc comparison.

posttest odds

In diagnostic testing, the odds that a patient has a given disease or condition after a diagnostic procedure is performed and interpreted. They are similar to the predictive value of a diagnostic test.

power

The ability of a test statistic to detect a specified alternative hypothesis or difference of a specified size when the alternative hypothesis is true (ie, 1 - β, where β is the probability of a type II error). More loosely, it is the ability of a study to detect an actual effect or difference.

predictive value of a negative test (PV-)

The proportion of time that a patient with a negative diagnostic test result does not have the disease being investigated.

predictive value of a positive test (PV+)

The proportion of time that a patient with a positive diagnostic test result has the disease being investigated.

pretest odds

In diagnostic testing, the odds that a patient has a given disease or condition before a diagnostic procedure is performed and interpreted. They are similar to prior probabilities.

prevalence

The proportion of people who have a given disease or condition at a specified point in time. It is not truly a rate, although it is often incorrectly called prevalence rate.

prior probability

The unconditional probability used in the numerator of Bayes' theorem. It is the prevalence of a disease prior to performing a diagnostic procedure. Clinicians often refer to it as the index of suspicion.

probability distribution

A frequency distribution of a random variable, which may be empirical or theoretical (eg, normal, binomial).

probability sample

See random sample.

probability

The number of times an outcome occurs in the total number of trials. If a is the outcome, the probability of a is denoted P(a).

product limit method

See Kaplan–Meier product limit method.

progressively censored

A situation in which patients enter a study at different points in time and remain in the study for varying lengths of time. See censored observation.

propensity score

An advanced statistical method to control for an entire group of confounding variables.

proportion

The number of observations with the characteristic of interest divided by the total number of observations. It is used to summarize counts.

proportional hazards model

See Cox proportional hazard model.

prospective study

A study designed before data are collected.

PUBMED

See MEDLINE.

qualitative observations

Characteristics measured on a nominal scale.

quality of life (QOL)

A measure of a person's subjective assessment of the value of his or her health and functional abilities.

quantitative observations

Characteristics measured on a numerical scale; the resulting numbers have inherent meaning.

quartile

The 25th percentile or the 75th percentile, called the first and third quartiles, respectively.

random assignment

The use of random methods to assign different treatments to patients or vice versa.

random error or variation

The variation in a sample that can be expected to occur by chance.

random sample

A sample of n subjects (or objects) selected from a population so that each has a known chance of being in the sample.

random variable

A variable in a study in which subjects are randomly selected or randomly assigned to treatments.

randomization

The process of assigning subjects to different treatments (or vice versa) by using random numbers.

randomized block design

A study design used in ANOVA to help control for potential confounding.

randomized controlled trial (RCT)

An experimental study in which subjects are randomly assigned to treatment groups.

range

The difference between the largest and the smallest observation.

ranking scale

A question format on an interview or survey in which respondents are asked to rate (from 1 to …) the options listed.

rank-order scale

A scale for observations arranged according to their size, from lowest to highest or vice versa.

ranks

A set of observations arranged according to their size, from lowest to highest or vice versa.

rate

A proportion associated with a multiplier, called the base (eg, 1000, 10,000, 100,000), and computed over a specific period.

ratio

A part divided by another part. It is the number of observations with the characteristic of interest divided by the number without the characteristic.

regression

(of Y on X) The process of determining a prediction equation for predicting Y from X.

regression coefficient

The b in the simple regression equation Y = a + bX. It is sometimes interpreted as the slope of the regression line. In multiple regression, the bs are weights applied to the predictor variables.

regression toward the mean

The phenomenon in which a predicted outcome for any given person tends to be closer to the mean outcome than the person's actual observation.

relative risk (RR)

The ratio of the incidence of a given disease in exposed or at-risk persons to the incidence of the disease in unexposed persons. It is calculated in cohort or prospective studies.

relative risk reduction (RRR)

The reduction in risk with a new therapy relative to the risk without the new therapy; it is the absolute value of the difference between the experimental event rate and the control event rate divided by the control event rate (|EER – CER|/CER).

reliability

A measure of the reproducibility of a measurement. It is measured by kappa for nominal measures and by correlation for numerical measures.

repeated-measures design

A study design in which subjects are measured at more than one time. It is also called a split-plot design in ANOVA.

representative population

(or sample) A population or sample that is similar in important ways to the population to which the findings of a study are generalized.

residual

The difference between the predicted value and the actual value of the outcome (dependent) variable in regression.

response variable

See dependent variable.

retrospective cohort study

See historical cohort study.

retrospective study

A study undertaken in a post hoc manner, ie, after the observations have been made.

risk factor

A term used to designate a characteristic that is more prevalent among subjects who develop a given disease or outcome than among subjects who do not. It is generally considered to be causal.

risk ratio

See relative risk.

robust

A term used to describe a statistical method if the outcome is not affected to a large extent by a violation of the assumptions of the method.

ROC (receiver operating characteristic) curve

In diagnostic testing, a plot of the true-positives on the Y-axis versus the false-positives on the X-axis; used to evaluate the properties of a diagnostic test.

r-squared (r2)

The square of the correlation coefficient. It is interpreted as the amount of variance in one variable that is accounted for by knowing the second variable.

sample

A subset of the population.

sampled population

The population from which the sample is actually selected.

sampling distribution

(of a statistic) The frequency distribution of the statistic for many samples. It is used to make inferences about the statistic from a single sample.

sampling frame

A list of all possible subjects or objects in the population from which a random sample is to be drawn; required for some types of sampling, such as systematic sampling.

scale of measurement

The degree of precision with which a characteristic is measured. It is generally categorized into nominal (or categorical), ordinal, and numerical (or interval and ratio) scales.

scatterplot

A two-dimensional graph displaying the relationship between two numerical characteristics or variables.

Scheffé's procedure

A multiple-comparison method for comparing means following a significant F test in analysis of variance. It is the most conservative multiple-comparison method.

self-controlled study

A study in which the subjects serve as their own controls, achieved by measuring the characteristic of interest before and after an intervention.

sensitivity analysis

In decision analysis, a method for determining the way the decision changes as a function of probabilities and utilities used in the analysis.

sensitivity

The proportion of time a diagnostic test is positive in patients who have the disease or condition. A sensitive test has a low false-negative rate.

sex-specific mortality rate

A mortality rate specific to either males or females.

sign test

The nonparametric test used for testing a hypothesis about the median in a single group.

simple random sample

A random sample in which each of the n subjects (or objects) in the sample has an equal chance of being selected.

skewed distribution

A distribution in which a few outlying observations occur in one direction only. If the outlying observations are small, the distribution is skewed to the left, or negatively skewed; if they are large, the distribution is skewed to the right, or positively skewed.

slope

(of the regression line) The amount Y changes for each unit that X changes. It is designated by b in the sample.

Spearman's rank correlation (rho)

A nonparametric correlation that measures the tendency for two measurements to vary together.

specific rate

A rate that pertains to a specific group or segment of the observations (examples are age-specific mortality rate and cause-specific mortality rate).

specificity

The proportion of time that a diagnostic test is negative in patients who do not have the disease or condition. A specific test has a low false-positive rate.

standard deviation (SD)

The most common measure of dispersion or spread, denoted by σ in the population and SD or s in the sample. It can be used with the mean to describe the distribution of observations. It is the square root of the average of the squared deviations of the observations from their mean.

standard error (SE)

The standard deviation of the sampling distribution of a statistic.

standard error of the estimate (Sy.x)

A measure of the variation in a regression line. It is based on the differences between the predicted and actual values of the dependent variable Y.

standard error of the mean (SEM)

The standard deviation of the mean in a large number of samples.

standard normal distribution

The normal distribution with mean 0 and standard deviation 1, also called the z distribution.

standardized mortality ratio

The number of observed deaths divided by the number of expected deaths.

standardized regression coefficient

A regression coefficient that has the effect of the measurement scale removed so that the size of the coefficient can be interpreted.

statistic

A summary number for a sample (eg, the mean), often used as an estimate of a parameter in the population.

statistical significance

Generally interpreted as a result that would occur by chance, eg, 1 time in 20, with a P value less than or equal to 0.05. It occurs when the null hypothesis is rejected.

statistical test

The procedure used to test a null hypothesis (eg, t test, chi-square test).

stem-and-leaf plot

A graphic display for numerical data. It is similar to both a frequency table and a histogram.

stepwise regression

In multiple regression, a sequential method of selecting the variables to be included in the prediction equation.

stratified random sample

A sample consisting of random samples from each subpopulation (or stratum) in a population. It is used so that the investigator can be sure that each subpopulation is appropriately represented in the sample.

structured abstract

A journal article abstract that contains short, precise descriptions of the context of the study, the objective, design, the setting and participants, the methods or interventions, main outcomes, results, and conclusions.

subjective probability

An estimate of probability that reflects a person's opinion or best guess from previous experience.

sums of squares (SS)

Quantities calculated in analysis of variance and used to obtain the mean squares for the F test.

suppression of zero

A term used to describe a misleading graph that does not have a break (a jagged line) in the Y-axis to indicate that part of the scale is missing.

survey

An observational study that generally has a cross-sectional design; a commonly used design to collect opinions.

survival analysis

The statistical method for analyzing survival data when there are censored observations.

symbols

Often Greek letters that stand for the population parameters and the Latin letters that stand for the sample statistics.

symmetric distribution

A distribution that has the same shape on both sides of the mean. The mean, median, and mode are all equal. It is the opposite of a skewed distribution.

systematic error

A measurement error that is the same (or constant) over all observations. See also bias.

systematic random sample

A random sample obtained by selecting each kth subject or object.

t distribution

A symmetric distribution with mean zero and a standard deviation larger than that for the normal distribution for small sample sizes. As nincreases, the t distribution approaches the normal distribution.

t test

The statistical test for comparing a mean with a norm or for comparing two means with small sample sizes (n ≤ 30). It is also used for testing whether a correlation coefficient or a regression coefficient is zero.

target population

The population to which the investigator wishes to generalize.

test statistic

The specific statistic used to test the null hypothesis (eg, the t statistic or chi-square statistic).

testing threshold

In diagnostic testing, the point at which the optimal decision is to perform a diagnostic test.

test–retest reliability

A measure of the degree to which an instrument or test provides a consistent measure of a characteristic on different occasions.

third quartile

The 75th percentile.

threshold model

A model for deciding when a diagnostic test should be ordered, as opposed to doing nothing or treating the patient without performing the test.

transformation

A change in the scale for the values of a variable.

treatment threshold

In diagnostic testing, the point at which the optimal decision is to treat the patient without first performing a diagnostic test.

trial

An experiment involving humans, commonly called a clinical trial. It is also a replication (repetition) of an experiment.

true-negative (TN)

A test result that is negative in a person who does not have the disease.

true-positive (TP)

A test result that is positive in a person who has the disease.

Tukey's HSD (honestly significant difference) test

A post hoc test for making multiple pairwise comparisons between means following a significant F test in analysis of variance. It is a method highly recommended by statisticians.

two-sample t test

The statistical test used to test the null hypothesis that two independent (or unrelated) groups have the same mean.

two-tailed test

A test in which the alternative hypothesis specifies a deviation from the null hypothesis in either direction. The critical region is located in both ends of the distribution of the test statistic. It is also called a directional test.

two-way analysis of variance

ANOVA with two independent variables.

type I error

The error that results if a true null hypothesis is rejected or if a difference is concluded when no difference exists.

type II error

The error that results if a false null hypothesis is not rejected or if a difference is not detected when a difference exists.

unbiasedness

(of a statistic) A term used to describe a statistic whose mean based on a large number of samples is equal to the population parameter.

uncontrolled study

An experimental study that has no control subjects.

utility

The value of different outcomes in a decision tree.

validity

The property of a measurement that indicates how well it measures the characteristic.

variable

A characteristic of interest in a study that has different values for different subjects or objects.

variance

2 in the population, s2 in the sample) The square of the standard deviation.

variation (within subject)

The variability in measurements of the same object or subject. It may occur naturally or may represent an error.

vital statistics

Mortality and morbidity rates used in epidemiology and public health.

weighted average

A number formed by multiplying each number in a set of numbers by a value called a weight, adding the resulting products, and then dividing by the sum of the weight.

Wilcoxon rank sum test

A nonparametric test for comparing two independent samples with ordinal data or with numerical observations that are not normally distributed.

Wilcoxon signed ranks test

A nonparametric test for comparing two dependent samples with ordinal data or with numerical observations that are not normally distributed.

Yates' correction

The process of subtracting 0.5 from the numerator at each term in the chi-square statistic for 2 × 2 tables prior to squaring the term.

z approximation

(to the binomial) The z test used to test the equality of two independent proportions.

z distribution

The normal distribution with mean 0 and standard deviation 1. It is also called the standard normal distribution.

z ratio

The test statistic used in the z test. It is formed by subtracting the hypothesized mean from the observed mean and dividing by the standard error of the mean.

z score

The deviation of X from the mean divided by the standard deviation.

z test

The statistical test for comparing a mean with a norm or comparing two means for large samples (n ≥ 30).

z transformation

A transformation that changes a normally distributed variable with mean and standard deviation SD to the z distribution with mean 0 and standard deviation 1.

Editors: Dawson, Beth; Trapp, Robert G.

Title: Basic & Clinical Biostatistics, 4th Edition

Copyright Š2004 McGraw-Hill

> Back of Book > Symbol

Symbol

=

equals

does not equal

<

less than

less than or equal to

>

greater than

greater than or equal to

√a

square root of a

|a|

absolute (or positive) value of a

P(A)

probability of event A

P(A|B)

probability of event A given that event B has happened

H0

null hypothesis

H1

alternative hypothesis

α

Greek letter alpha; probability of type I error

β

Greek letter beta; probability of type II error; also, population value of the slope of the regression line

δ

Greek letter delta; population mean difference between two paired measurements

κ

Greek letter kappa; used to denote an index of agreement or reproducibility

λ

Greek letter lambda; used to denote terms in the log-linear model

ľ

Greek letter mu; population mean

π

Greek letter pi; population proportion

ρ

Greek letter rho; population correlation

σ

Greek lowercase letter sigma; population standard deviation

Σ

Greek uppercase letter sigma; symbol indicating a sum

CV

coefficient of variation

df

degree of freedom

e

base of natural logarithm (equal to 2.718)

E

expected frequency

F

symbol for the F test and distribution

H

hazard function

LR

likelihood ratio

MSA

mean square among groups

MSE

error mean square

n

sample size

O

observed frequency

r

sample correlation

r2

squared correlation, called of the coefficient of determination

R2

squared multiple correlation in multiple regression

SD

sample standard deviation

SE

standard error of the mean

sY.X

standard error of the estimate in regression

t

symbol for the t ratio (the critical ratio that follows a t distribution)

X

independent (explanatory, predictor) variable in regression

χ2

symbol for the chi-square test

x bar; sample mean

Y

dependent (outcome, response, criterion) variable in regression

Y′

predicted value of Y regression

z

symbol for the z ratio (the critical ratio that follows a z or standard normal distribution



If you find an error or have any questions, please email us at admin@doctorlib.org. Thank you!