Please read the instructions thoroughly!
Specific methods of data collection (e.g., surveys, interviews, observations) produce specific types of data that will answer particular research questions, but not others; so here too, as covered in previous weeks, the research questions inform how the data will be obtained. Furthermore, the method used to collect the data may impact the reliability and the validity of that data.
For this Discussion, you will first consider sampling strategies. Then, you will turn your attention to data collection methods, including their strengths, limitations, and ethical implications. Last, you will consider measurement reliability and validity in the context of your discipline.
Position A: Probability sampling represents the best strategy for selecting research participants.
Post a restatement of your assigned position on sampling strategies. Explain why this position is the best strategy for selecting research participants. Support your explanation with an example and support from the scholarly literature. Next, select a data collection method (e.g., surveys, interviews, observations) and briefly explain at least one strength and at least one limitation. Then, identify a potential ethical issue with this method and describe a strategy to address it. Last, explain the relationship between measurement reliability and measurement validity using an example from your discipline.
Be sure to support your Main Issue Post and Response Post with reference to the week’s Learning Resources and other scholarly evidence in APA Style.
Mixed Methods Sampling: A Typology with Examples by Teddlie, C., & Yu, F., in Journal of Mixed Methods Research, Vol. 1/Issue 1. Copyright 2007 by Sage Publications Inc. Reprinted by permission of Sage Publications Inc. via the Copyright Clearance Center.
TheQualitative Report The Qualitative Report
Volume 12 Number 2 Article 9
6-1-2007
A Typology of Mixed Methods Sampling Designs in Social Science A Typology of Mixed Methods Sampling Designs in Social Science
Research
Research
Anthony J. Onwuegbuzie
Sam Houston State University, tonyonwuegbuzie@aol.com
Kathleen M.T. Collins
Unversity of Arkansas
Follow this and additional works at: https://nsuworks.nova.edu/tqr
Part of the Quantitative, Qualitative, Comparative, and Historical Methodologies Commons, and the
Social Statistics Commons
Recommended APA Citation Recommended APA Citation
Onwuegbuzie, A. J., & Collins, K. M. (2007). A Typology of Mixed Methods Sampling Designs in Social
Science Research . The Qualitative Report, 12(2), 281-316. Retrieved from https://nsuworks.nova.edu/tqr/
vol12/iss2/9
This Article is brought to you for free and open access by the The Qualitative Report at NSUWorks. It has been
accepted for inclusion in The Qualitative Report by an authorized administrator of NSUWorks. For more
information, please contact nsuworks@nova.edu.
http://nsuworks.nova.edu/tqr/
http://nsuworks.nova.edu/tqr/
https://nsuworks.nova.edu/tqr
https://nsuworks.nova.edu/tqr/vol12
https://nsuworks.nova.edu/tqr/vol12/iss2
https://nsuworks.nova.edu/tqr/vol12/iss2/9
https://nsuworks.nova.edu/tqr?utm_source=nsuworks.nova.edu%2Ftqr%2Fvol12%2Fiss2%2F9&utm_medium=PDF&utm_campaign=PDFCoverPages
http://network.bepress.com/hgg/discipline/423?utm_source=nsuworks.nova.edu%2Ftqr%2Fvol12%2Fiss2%2F9&utm_medium=PDF&utm_campaign=PDFCoverPages
http://network.bepress.com/hgg/discipline/1275?utm_source=nsuworks.nova.edu%2Ftqr%2Fvol12%2Fiss2%2F9&utm_medium=PDF&utm_campaign=PDFCoverPages
https://nsuworks.nova.edu/tqr/vol12/iss2/9?utm_source=nsuworks.nova.edu%2Ftqr%2Fvol12%2Fiss2%2F9&utm_medium=PDF&utm_campaign=PDFCoverPages
https://nsuworks.nova.edu/tqr/vol12/iss2/9?utm_source=nsuworks.nova.edu%2Ftqr%2Fvol12%2Fiss2%2F9&utm_medium=PDF&utm_campaign=PDFCoverPages
mailto:nsuworks@nova.edu
A Typology of Mixed Methods Sampling Designs in Social Science Research
Abstract Abstract
This paper provides a framework for developing sampling designs in mixed methods research. First, we
present sampling schemes that have been associated with quantitative and qualitative research. Second,
we discuss sample size considerations and provide sample size recommendations for each of the major
research designs for quantitative and qualitative approaches. Third, we provide a sampling design
typology and we demonstrate how sampling designs can be classified according to time orientation of
the components and relationship of the qualitative and quantitative sample. Fourth, we present four major
crises to mixed methods research and indicate how each crisis may be used to guide sampling design
considerations. Finally, we emphasize how sampling design impacts the extent to which researchers can
generalize their findings.
Keywords Keywords
Sampling Schemes, Qualitative Research, Generalization, Parallel Sampling Designs, Pairwise Sampling
Designs, Subgroup Sampling Designs, Nested Sampling Designs, and Multilevel Sampling Designs
Creative Commons License Creative Commons License
This work is licensed under a Creative Commons Attribution-Noncommercial-Share Alike 4.0 License.
This article is available in The Qualitative Report: https://nsuworks.nova.edu/tqr/vol12/iss2/9
https://goo.gl/u1Hmes
https://goo.gl/u1Hmes
https://creativecommons.org/licenses/by-nc-sa/4.0/
https://creativecommons.org/licenses/by-nc-sa/4.0/
https://creativecommons.org/licenses/by-nc-sa/4.0/
https://nsuworks.nova.edu/tqr/vol12/iss2/9
The Qualitative Report Volume 12 Number 2 June 2007 281-316
http://www.nova.edu/ssss/QR/QR12-2/onwuegbuzie2
A Typology of Mixed Methods Sampling Designs in Social
Science Research
Anthony J. Onwuegbuzie
Sam Houston State University, Huntsville, Texas
Kathleen M. T. Collins
University of Arkansas, Fayetteville, Arkansas
This paper provides a framework for developing sampling designs in mixed
methods research. First, we present sampling schemes that have been
associated with quantitative and qualitative research. Second, we discuss
sample size considerations and provide sample size recommendations for
each of the major research designs for quantitative and qualitative
approaches. Third, we provide a sampling design typology and we
demonstrate how sampling designs can be classified according to time
orientation of the components and relationship of the qualitative and
quantitative sample. Fourth, we present four major crises to mixed methods
research and indicate how each crisis may be used to guide sampling design
considerations. Finally, we emphasize how sampling design impacts the
extent to which researchers can generalize their findings. Key Words:
Sampling Schemes, Qualitative Research, Generalization, Parallel Sampling
Designs, Pairwise Sampling Designs, Subgroup Sampling Designs, Nested
Sampling Designs, and Multilevel Sampling Designs
Sampling, which is the process of selecting “a portion, piece, or segment that is
representative of a whole” (The American Heritage College Dictionary, 1993, p. 1206), is an
important step in the research process because it helps to inform the quality of inferences
made by the researcher that stem from the underlying findings. In both quantitative and
qualitative studies, researchers must decide the number of participants to select (i.e., sample
size) and how to select these sample members (i.e., sampling scheme). While the decisions
can be difficult for both qualitative and quantitative researchers, sampling strategies are even
more complex for studies in which qualitative and quantitative research approaches are
combined either concurrently or sequentially. Studies that combine or mix qualitative and
quantitative research techniques fall into a class of research that are appropriately called
mixed methods research or mixed research. Sampling decisions typically are more
complicated in mixed methods research because sampling schemes must be designed for
both the qualitative and quantitative research components of these studies.
Despite the fact that mixed methods studies have now become popularized, and
despite the number of books (Brewer & Hunter, 1989; Bryman, 1989; Cook & Reichardt,
1979; Creswell, 1994; Greene & Caracelli, 1997; Newman & Benz, 1998; Reichardt &
Rallis, 1994; Tashakkori & Teddlie, 1998, 2003a), book chapters (Creswell, 1999, 2002;
Jick, 1983; Li, Marquart, & Zercher, 2000; McMillan & Schumacher, 2001; Onwuegbuzie,
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 282
Jiao, & Bostick, 2004; Onwuegbuzie & Johnson, 2004; Smith, 1986), and methodological
articles (Caracelli & Greene, 1993; Dzurec & Abraham, 1993; Greene, Caracelli, & Graham,
1989; Greene, & McClintock, 1985; Gueulette, Newgent, & Newman, 1999; Howe, 1988,
1992; Jick, 1979; Johnson & Onwuegbuzie, 2004; Laurie & Sullivan, 1991; Morgan, 1998;
Morse, 1991, 1996; Onwuegbuzie, 2002a; Onwuegbuzie & Leech, 2004b, 2005a; Rossman
& Wilson, 1985; Sandelowski, 2001; Sechrest & Sidana, 1995; Sieber, 1973; Tashakkori &
Teddlie, 2003b; Waysman & Savaya, 1997) devoted to mixed methods research, relatively
little has been written on the topic of sampling. In fact, at the time of writing1, with the
exception of Kemper, Stringfield, and Teddlie (2003) and Onwuegbuzie and Leech (2005a),
discussion of sampling schemes has taken place in ways that link research paradigm to
method. Specifically, random sampling schemes are presented as belonging to the
quantitative paradigm, whereas non-random sampling schemes are presented as belonging to
the qualitative paradigm. As noted by Onwuegbuzie and Leech (2005a), this represents a
false dichotomy. Rather, both random and non-random sampling can be used in quantitative
and qualitative studies.
Similarly, discussion of sample size considerations tends to be dichotomized, with
small samples being associated with qualitative research and large samples being associated
with quantitative studies. Although this represents the most common way of linking sample
size to research paradigm, this representation is too simplistic and thereby misleading.
Indeed, there are times when it is appropriate to use small samples in quantitative research,
while there are occasions when it is justified to use large samples in qualitative research.
With this in mind, the purpose of this paper is to provide a framework for developing
sampling designs in mixed methods research. First, we present the most common sampling
schemes that have been associated with both quantitative and qualitative research. We
contend that although sampling schemes traditionally have been linked to research paradigm
(e.g., random sampling has been associated with quantitative research) in research
methodology textbooks (Onwuegbuzie & Leech, 2005b), this is not consistent with practice.
Second, we discuss the importance of researchers making sample size considerations in both
quantitative and qualitative research. We then provide sample size recommendations for
each of the major research designs for both approaches. Third, we provide a typology of
sampling designs in mixed methods research. Here, we demonstrate how sampling designs
can be classified according to: (a) the time orientation of a study’s components (i.e., whether
the qualitative and quantitative components occur simultaneously or sequentially) and (b) the
relationship of the qualitative and quantitative samples (e.g., identical vs. nested). Fourth, we
present the four major crises or challenges to mixed methods research: representation,
legitimation, integration, and politics. These crises are then used to provide guidelines for
making sampling design considerations. Finally, we emphasize how choice of sampling
design helps to determine the extent to which researchers can generalize their findings and
make what Tashakkori and Teddlie (2003c, p. 687) refer to as “meta-inferences;” namely,
the term they give to describe the integration of generalizable inferences that are derived on
the basis of findings stemming from the qualitative and quantitative components of a mixed
methods study.
1 Since this article was accepted for publication, the following three articles in the area of mixed methods
sampling have emerged: Teddlie and Yu (2007) and Collins et al. (2006, 2007). Each of these three articles
cites the present article, and the latter two articles used the framework of the current article. However, despite
these additions to the literature, it is still accurate for us to state that relatively little has been written in this area.
283 The Qualitative Report June 2007
For the purposes of the present article, we distinguish between sampling schemes and
sampling designs. We define sampling schemes as specific strategies used to select units
(e.g., people, groups, events, settings). Conversely, sampling designs represent the
framework within which the sampling takes place, including the number and types of
sampling schemes as well as the sample size.
The next section presents the major sampling schemes. This is directly followed by a
section on sample size considerations. After discussing sampling schemes and sample sizes,
a presentation of sampling designs ensues. Indeed, a typology of sampling designs is
outlined that incorporates all of the available sampling schemes.
Sampling Schemes
According to Curtis, Gesler, Smith, and Washburn (2000) and Onwuegbuzie and
Leech (2005c, 2007a), some kind of generalizing typically occurs in both quantitative and
qualitative research. Quantitative researchers tend to make “statistical” generalizations,
which involve generalizing findings and inferences from a representative statistical sample to
the population from which the sample was drawn. In contrast, many qualitative researchers,
although not all, tend to make “analytic” generalizations (Miles & Huberman, 1994), which
are “applied to wider theory on the basis of how selected cases ‘fit’ with general constructs”
(Curtis et al., 2000, p. 1002); or they make generalizations that involve case-to-case transfer
(Firestone, 1993; Kennedy, 1979). In other words, statistical generalizability refers to
representativeness (i.e., some form of universal generalizability), whereas analytic
generalizability and case-to-case transfer relate to conceptual power (Miles & Huberman,
1994). Therefore, the process of sampling is important to both quantitative and qualitative
research. Unfortunately, a false dichotomy appears to prevail with respect to sampling
schemes available to quantitative and qualitative researchers. As noted by Onwuegbuzie and
Leech (2005b), random sampling tends to be associated with quantitative research, whereas
non-random sampling typically is linked to qualitative research. However, choice of
sampling class (i.e., random vs. non-random) should be based on the type of generalization
of interest (i.e., statistical vs. analytic). In fact, qualitative research can involve random
sampling. For example, Carrese, Mullaney, and Faden (2002) used random sampling
techniques to select 20 chronically ill housebound patients (aged 75 years or older), who
were subsequently interviewed to examine how elderly patients think about and approach
future illness and the end of life. Similarly, non-random sampling techniques can be used in
quantitative studies. Indeed, although this adversely affects the external validity (i.e.,
generalizability) of findings, the majority of quantitative research studies utilize non-random
samples (cf. Leech & Onwuegbuzie, 2002). Breaking down this false dichotomy
significantly increases the options that both qualitative and quantitative researchers have for
selecting their samples.
Building on the work of Patton (1990) and Miles and Huberman (1994),
Onwuegbuzie and Leech (2007a) identified 24 sampling schemes that they contend both
qualitative and quantitative researchers have available for use. All of these sampling schemes
fall into one of two classes: random sampling (i.e., probabilistic sampling) schemes or non-
random sampling (i.e., non-probabilistic sampling) schemes. These sampling schemes
encompass methods for selecting samples that have been traditionally associated with the
qualitative paradigm (i.e., non-random sampling schemes) and those that have been typically
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 284
associated with the quantitative paradigm (i.e., random sampling schemes). Table 1 (below)
presents a matrix that crosses type of sampling scheme (i.e., random vs. non-random) and
research approach (qualitative vs. quantitative). Because the vast majority of both qualitative
and quantitative studies use non-random samples, Type 4 (as shown in Table 1) is by far the
most common combination of sampling schemes in mixed methods used, regardless of
mixed methods research goal (i.e., to predict; add to the knowledge base; have a personal,
social, institutional, and/or organizational impact; measure change; understand complex
phenomena; test new ideas; generate new ideas; inform constituencies; or examine the past;
Newman, Ridenour, Newman, & DeMarco, 2003), research objective (i.e., exploration,
description, explanation, prediction, or influence; Johnson & Christensen, 2004), research
purpose (i.e., triangulation, or seeking convergence of findings; complementarity, or
examining different overlapping aspects of a phenomenon; initiation, or discerning
paradoxes and contradictions; development, or using the results from the first method to
inform the use of the second method; or expansion, adding breath and scope to a study;
Greene et al., 1989), and research question. Conversely, Type 1, involving random sampling
for both the qualitative and quantitative components of a mixed methods study, is the least
common. Type 3, involving random sampling for the qualitative component(s) and non-
random sampling for the quantitative component(s) also is rare. Finally, Type 2, consisting
of non-random sampling for the qualitative component(s) and random sampling for the
quantitative component(s) is the second most common combination.
Table 1
Matrix Crossing Type of Sampling Scheme by Research Approach
Qualitative Component(s)
Random Sampling
Non-Random
Sampling
Quantitative
Component(s)
Random Sampling
Rare
Combination
(Type 1)
Occasional
Combination
(Type 2)
Non-Random Sampling
Very Rare
Combination
(Type 3)
Frequent
Combination
(Type 4)
285 The Qualitative Report June 2007
Random (Probability) Sampling
Before deciding on the sampling scheme, mixed methods researchers must decide
what the objective of the study is. For example, if the objective of the study is to generalize
the quantitative and/or qualitative findings to the population from which the sample was
drawn (i.e., make inferences), then the researcher should attempt to select a sample for that
component that is random. In this situation, the mixed method researcher can select one of
five random (i.e., probability) sampling schemes at one or more stages of the research
process: simple random sampling, stratified random sampling, cluster random sampling,
systematic random sampling, and multi-stage random sampling. Each of these strategies is
summarized in Table 2.
Table 2
Major Sampling Schemes in Mixed Methods Research
Sampling Scheme
Description
Simplea
Stratifieda
Clustera
Systematica
Multi-Stage Randoma
Maximum Variation
Homogeneous
Critical Case
Every individual in the sampling frame (i.e., desired
population) has an equal and independent chance of being
chosen for the study.
Sampling frame is divided into sub-sections comprising
groups that are relatively homogeneous with respect to one or
more characteristics and a random sample from each stratum
is selected.
Selecting intact groups representing clusters of individuals
rather than choosing individuals one at a time.
Choosing individuals from a list by selecting every kth
sampling frame member, where k typifies the population
divided by the preferred sample size.
Choosing a sample from the random sampling schemes in
multiple stages.
Choosing settings, groups, and/or individuals to maximize the
range of perspectives investigated in the study.
Choosing settings, groups, and/or individuals based on similar
or specific characteristics.
Choosing settings, groups, and/or individuals based on
specific characteristic(s) because their inclusion provides the
researcher with compelling insight about a phenomenon of
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 286
Theory-Based
Confirming
Disconfirming
Snowball/Chain
Extreme Case
interest.
Choosing settings, groups, and/or individuals because their
inclusion helps the researcher to develop a theory.
After beginning data collection, the researcher conducts
subsequent analyses to verify or contradict initial results.
Participants are asked to recruit individuals to join the study.
Selecting outlying cases and conducting comparative
analyses.
Typical Case
Intensity
Politically Important
Case
Random Purposeful
Stratified Purposeful
Criterion
Opportunistic
Mixed Purposeful
Convenience
Selecting and analyzing average or normal cases.
Choosing settings, groups, and/or individuals because their
experiences relative to the phenomena of interest are viewed
as
intense but not extreme.
Choosing settings, groups, and/or individuals to be included
or excluded based on their political connections to the
phenomena of interest.
Selecting random cases from the sampling frame and
randomly choosing a desired number of individuals to
participate in the study.
Sampling frame is divided into strata to obtain relatively
homogeneous sub-groups and a purposeful sample is selected
from each stratum.
Choosing settings, groups, and/or individuals because they
represent one or more criteria.
Researcher selects a case based on specific characteristics
(i.e., typical, negative, or extreme) to capitalize on developing
events occurring during data collection.
Choosing more than one sampling strategy and comparing the
results emerging from both samples.
Choosing settings, groups, and/or individuals that are
conveniently available and willing to participate in the study.
287 The Qualitative Report June 2007
Quota
Multi-Stage Purposeful
Random
Multi-Stage Purposeful
Researcher identifies desired characteristics and quotas of
sample members to be included in the study.
Choosing settings, groups, and/or individuals representing a
sample in two or more stages. The first stage is random
selection and the following stages are purposive selection of
participants.
Choosing settings, groups, and/or individuals representing a
sample in two or more stages in which all stages reflect
purposive sampling of participants.
a Represent random (i.e., probabilistic) sampling schemes. All other schemes are non-random.
Non-Random (Non-Probability) Sampling
If the goal is not to generalize to a population but to obtain insights into a
phenomenon, individuals, or events (as will often be the case in the qualitative component of
a mixed methods study), then the researcher purposefully selects individuals, groups, and
settings for this phase that maximize understanding of the underlying phenomenon. Thus,
many mixed methods studies utilize some form of purposeful sampling. Here, individuals,
groups, and settings are considered for selection if they are “information rich” (Patton, 1990,
p. 169). There are currently 19 purposive sampling schemes. These schemes differ with
respect to whether they are implemented before data collection has started or after data
collection begins (Creswell, 2002). Also, the appropriateness of each scheme is dependent on
the research goal, objective, purpose, and question. Each of these non-random sampling
schemes is summarized in Table 2.
Thus, mixed methods researchers presently have 24 sampling schemes from which to
choose. These 24 designs comprise 5 probability sampling schemes and 19 purposive
sampling schemes. For a discussion of these sampling schemes, we refer readers to Collins,
Onwuegbuzie, and Jiao (2006, in press), Kemper et al. (2003), Miles and Huberman (1994),
Onwuegbuzie and Leech (2007a), Patton (1990), and Teddlie and Yu (2007). As Kemper et
al. concluded, “the understanding of a wide range of sampling techniques in one’s
methodological repertoire greatly increases the likelihood of one’s generating findings that
are both rich in content and inclusive in scope” (p. 292).
Sample Size
In addition to deciding how to select the samples for the qualitative and quantitative
components of a study, mixed methods researchers also should determine appropriate sample
sizes for each phase. The choice of sample size is as important as is the choice of sampling
scheme because it also determines the extent to which the researcher can make statistical
and/or analytic generalizations. Unfortunately, as has been the case with sampling schemes,
discussion of sample size considerations has tended to be dichotomized, with small samples
being associated with qualitative research and large samples being linked to quantitative
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 288
studies. Yet, small samples can be used in quantitative research that represents exploratory
research or basic research. In fact, single-subject designs, which routinely utilize quantitative
approaches, are characterized by small samples. Conversely, qualitative research can utilize
large samples, as in the case of program evaluation research. Moreover, to associate
qualitative data analyses with small samples is to ignore the growing body of literature in the
area of text mining, the process of analyzing naturally occurring text in order to discover and
capture semantic information (see, for example, Del Rio, Kostoff, Garcia, Ramirez, &
Humenik, 2002; Liddy, 2000; Powis & Cairns, 2003; Srinivasan, 2004).
The size of the sample should be informed primarily by the research objective,
research question(s), and, subsequently, the research design. Table 3 presents minimum
sample sizes for several of the most common research designs. The sample sizes
corresponding to the traditional quantitative research designs (i.e., correlational, causal-
comparative, experimental) are the result of the statistical power analysis undertaken by
Onwuegbuzie et al. (2004). According to Onwuegbuzie et al. (2004), many of the sample
size guidelines provided in virtually every introductory research methodology and statistics
textbook, such as the recommendation of sample sizes of 30 for both correlational and
causal-comparative designs (e.g., Charles & Mertler, 2002; Creswell, 2002; Gall, Borg, &
Gall, 1996; Gay & Airasian, 2003; McMillan & Schumacher, 2001), if followed, would lead
to statistical tests with inadequate power because they are not based on power analyses. For
example, for correlational research designs, a minimum sample size of 30 represents a
statistical power of only .51 for one-tailed tests for detecting a moderate relationship (i.e., r =
.30) between two variables at the 5% level of statistical significance, and a power of .38 for
two-tailed tests of moderate relationships (Erdfelder, Faul, & Buchner, 1996; Onwuegbuzie
et al., 2004). Therefore, the proposed sample sizes in Table 3 represent sizes for detecting
moderate effect sizes with .80 statistical power at the 5% level of significance.
Table 3
Minimum Sample Size Recommendations for Most Common Quantitative and Qualitative
Research Designs
Research Design/Method
Minimum Sample Size Suggestion
Research Design1
Correlational
Causal-Comparative
Experimental
Case Study
64 participants for one-tailed hypotheses; 82 participants
for two-tailed hypotheses (Onwuegbuzie et al., 2004)
51 participants per group for one-tailed hypotheses; 64
participants for two-tailed hypotheses (Onwuegbuzie et al.,
2004)
21 participants per group for one-tailed hypotheses
(Onwuegbuzie et al., 2004)
3-5 participants (Creswell, 2002)
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
289 The Qualitative Report June 2007
Phenomenological
Grounded Theory
Ethnography
Ethological
Sampling Design
Subgroup Sampling
Design
Nested Sampling Design
Data Collection Procedure
Interview
Focus Group
≤ 10 interviews (Creswell, 1998); ≥ 6 (Morse, 1994)
15-20 (Creswell, 2002); 20-30 (Creswell, 2007)
1 cultural group (Creswell, 2002); 30-50 interviews
(Morse, 1994)
100-200 units of observation (Morse, 1994)
≥ 3 participants per subgroup (Onwuegbuzie & Leech,
2007c)
≥ 3 participants per subgroup (Onwuegbuzie & Leech,
2007c)
12 participants (Guest, Bunce, & Johnson, 2006)
6-9 participants (Krueger, 2000); 6-10 participants
(Langford, Schoenfeld, & Izzo, 2002; Morgan, 1997); 6-12
participants (Johnson & Christensen, 2004); 6-12
participants (Bernard, 1995); 8–12 participants
(Baumgartner, Strong, & Hensley, 2002)
3 to 6 focus groups (Krueger, 1994; Morgan, 1997;
Onwuegbuzie, Dickinson, Leech, & Zoran, 2007)
1 For correlational, causal-comparative, and experimental research designs, the recommended sample sizes
represent those needed to detect a medium (using Cohen’s [1988] criteria), one-tailed statistically significant
relationship or difference with .80 power at the 5% level of significance.
As Sandelowski (1995) stated, “a common misconception about sampling in
qualitative research is that numbers are unimportant in ensuring the adequacy of a sampling
strategy” (p. 179). However, some methodologists have provided guidelines for selecting
samples in qualitative studies based on the research design (e.g., case study, ethnography,
phenomenology, grounded theory), sampling design (i.e., subgroup sampling design, nested
sampling design), or data collection procedure (i.e., interview, focus group). These
recommendations also are summarized in Table 3. In general, sample sizes in qualitative
research should not be so small as to make it difficult to achieve data saturation, theoretical
saturation, or informational redundancy. At the same time, the sample should not be so large
that it is difficult to undertake a deep, case-oriented analysis (Sandelowski, 1995).
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 290
Mixed Methods Sampling Designs
The sampling schemes described in previous sections could be used in isolation.
Indeed, each of these sampling schemes could be used in monomethod research that
characterizes either solely qualitative or quantitative studies. That is, both qualitative and
quantitative researchers can use any of the 24 sampling schemes, as appropriate, to address
their research questions. However, in mixed methods research, sampling schemes must be
chosen for both the qualitative and quantitative components of the study. Therefore,
sampling typically is much more complex in mixed methods studies than in monomethod
studies.
In fact, the mixed methods sampling process involves the following seven distinct
steps: (a) determine the goal of the study, (b) formulate the research objective(s), (c)
determine the research purpose, (d) determine the research question(s), (e) select the research
design, (f) select the sampling design, and (g) select the sampling scheme. These steps are
presented in Figure 1. From this figure, it can be seen that these steps are linear. That is, the
study’s goal (e.g., understand complex phenomena, test new ideas) leads to the research
objective(s) (e.g., exploration, prediction), which, in turn, leads to a determination of the
research purpose (e.g., triangulation, complementarity), which is followed by the selection of
the mixed methods research design.
Currently, there are many mixed methods research designs in existence. In the
Tashakkori and Teddlie (2003a) book alone, approximately 35 mixed methods research
designs are outlined. Thus, in order to simplify researchers’ design choices, several
typologies have been developed (e.g., Creswell, 1994, 2002; Creswell, Plano Clark,
Guttmann, & Hanson, 2003; Greene & Caracelli, 1997; Greene et al., 1989; Johnson &
Onwuegbuzie, 2004; Maxwell & Loomis, 2003; McMillan & Schumacher, 2001; Morgan,
1998; Morse, 1991, 2003; Onwuegbuzie & Johnson, 2004; Patton, 1990; Tashakkori &
Teddlie, 1998, 2003c). These typologies differ in their levels of complexity. However, most
mixed method designs utilize time orientation dimension as its base. Time orientation refers
to whether the qualitative and quantitative phases of the study occur at approximately the
same point in time such that they are independent of one another (i.e., concurrent) or whether
these two components occur one after the other such that the latter phase is dependent, to
some degree, on the former phase (i.e., sequential). An example of a concurrent mixed
methods design is a study examining attitudes toward reading and reading strategies among
fifth-grade students that involves administering a survey containing both closed-ended items
(e.g., Likert-format responses that measure attitudes toward reading) and open-ended
questions (i.e., that elicit qualitative information about the students’ reading strategies).
Conversely, an example of a sequential mixed methods design is a descriptive assessment of
reading achievement levels among 30 fifth-grade students (quantitative phase), followed by
an interview (i.e., qualitative phase) of the highest and lowest 3 fifth-grade students who
were identified in the quantitative phase in order to examine their reading strategies. Thus, in
order to select a mixed method design, the researcher should decide whether one wants to
conduct the phases concurrently (i.e., independently) or sequentially (i.e., dependently). As
noted earlier, another decision that the researcher should make relates to the purpose of
mixing the quantitative and qualitative approaches (e.g., triangulation, complementarity,
initiation, development, expansion).
291 The Qualitative Report June 2007
Figure 1. Steps in the mixed methods sampling process.
Determine the
Goal of the Study
Formulate Research
Objectives
Determine Research
Purpose
Determine Research
Question(s)
Select Research Design
Select the
Sampling Design
Select the Individual
Sampling Schemes
Crossing these two dimensions (i.e., time order and purpose of mixing) produces a 2
(concurrent vs. sequential) x 5 (triangulation vs. complementarity vs. initiation vs.
development vs. expansion) matrix that produces 10 cells. This matrix is presented in Table
4. This matrix matches the time orientation to the mixed methods purpose. For instance, if
the purpose of the mixed methods research is triangulation, then a concurrent design is
appropriate such that the quantitative and qualitative data can be triangulated. As noted by
Creswell et al. (2003),
In concurrently gathering both forms of data at the same time, the researcher
seeks to compare both forms of data to search for congruent findings (e.g.,
how the themes identified in the qualitative data collection compare with the
statistical results in the quantitative analysis, pp. 217-218).
However, sequential designs are not appropriate for triangulation because when they are
utilized either the qualitative or quantitative data are gathered first, such that findings from
the first approach might influence those from the second approach, thereby positively biasing
any comparisons. On the other hand, if the mixed methods purpose is development, then
sequential designs are appropriate because development involves using the methods
sequentially, such that the findings from the first method inform the use of the second
method. For this reason, concurrent designs do not address development purposes. Similarly,
sequential designs only are appropriate for expansion purposes. Finally, both concurrent and
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 292
sequential designs can be justified if the mixed method purpose either is complementarity or
initiation.
Table 4
Matrix Crossing Purpose of Mixed Methods Research by Time Orientation
Purpose of Mixed Methods
Research
Concurrent Design
Appropriate?
Sequential Design
Appropriate?
Triangulation
Complementarity
Development
Initiation
Expansion
Yes
Yes
No
Yes
No
No
Yes
Yes
Yes
Yes
Once a decision has been made about the mixed method purpose and design type
(i.e., time orientation), the next step is for the researcher to select a mixed methods sampling
design. Two criteria are useful here: time orientation (i.e., concurrent vs. sequential) and
relationship of the qualitative and quantitative samples. These relationships either can be
identical, parallel, nested, or multilevel. An identical relationship indicates that exactly the
same sample members participate in both the qualitative and quantitative phases of the study
(e.g., administering a survey of reading attitudes and reading strategies to a class of fourth
graders that contains both closed- and open-ended items, yielding quantitative and
qualitative phases that occur simultaneously). A parallel relationship specifies that the
samples for the qualitative and quantitative components of the research are different but are
drawn from the same population of interest (e.g., administering a quantitative measure of
reading attitudes to one class of third-grade students for the quantitative phase and
conducting in-depth interviews and observations examining reading strategies on a small
sample of third-grade students from another class within the same school, or from another
school for the qualitative phase). A nested relationship implies that the sample members
selected for one phase of the study represent a subset of those participants chosen for the
other facet of the investigation (e.g., administering a quantitative measure of reading
attitudes to one class of third-grade students for the quantitative phase and conducting in-
depth interviews and observations examining reading strategies on the lowest- and highest-
scoring third-grade students from the same class). Finally, a multilevel relationship involves
the use of two or more sets of samples that are extracted from different levels of the study
(i.e., different populations). For example, whereas one phase of the investigation (e.g.,
quantitative phase) might involve the sampling of students within a high school, the other
phase (e.g., qualitative) might involve the sampling of their teachers, principal, and/or
parents. Thus, the multilevel relationship is similar to what Kemper et al. (2003) call
multilevel sampling in mixed methods studies, where Kemper et al. define it as occurring
“when probability and purposive sampling techniques are used on different levels of the
293 The Qualitative Report June 2007
study (e.g., student, class, school district)” (p. 287), while in the present conceptualization,
multilevel sampling could involve combining probability and purposive sampling techniques
in any of the four ways described in Table 1 (i.e., Type 1 – Type 4). Thus, for example,
multilevel sampling in mixed methods studies could involve sampling on all levels being
purposive or sampling on all levels being random. Therefore, our use of the multilevel is
more general and inclusive than that of Kemper et al. The two criteria, time orientation and
sample relationship, yield eight different types of major sampling designs that a mixed
methods researcher might use. These designs, which are labeled as Design 1 to Design 8, are
outlined in our Two-Dimensional Mixed Methods Sampling Model in Figure 2.
Design 1 involves a concurrent design using identical samples for both qualitative
and quantitative components of the study. An example of a Design 1 sampling design is the
study conducted by Daley and Onwuegbuzie (2004). These researchers examined male
juvenile delinquents’ causal attributions for others’ violent behavior, and the salient pieces of
information they utilize in arriving at these attributions. A 12-item questionnaire called the
Violence Attribution Survey, which was designed by Daley and Onwuegbuzie, was used to
assess attributions made by juveniles for the behavior of others involved in violent acts. Each
item consisted of a vignette, followed by three possible attributions (i.e., person, stimulus,
circumstance), presented using a multiple-choice format (i.e., quantitative component), and
an open-ended question asking the juveniles their reasons for choosing the responses that
they did (i.e., qualitative component). Participants included 82 male juvenile offenders who
were drawn randomly from the population of juveniles incarcerated at correctional facilities
in a large southeastern state. By collecting quantitative and qualitative data within the same
time frame from the same sample members, the researchers used a concurrent, identical
sampling design. Simple random sampling was used to select the identical samples. Because
these identical samples were selected randomly, Daley and Onwuegbuzie’s combined
sampling schemes can be classified as being Type 1 (cf. Table 1).
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 294
Figure 2. Two-dimensional mixed methods sampling model providing a typology of
mixed methods sampling designs.
Time Orientation Relationship of Samples Sampling Schemes
Sequential:
QUAL → QUAN
QUAL → quan
qual → QUAN
QUAN → QUAL
QUAN → qual
quan → QUAL
Identical (5)
Parallel (6)
Nested (7)
Multilevel (8)
Identical (1)
Parallel (2)
Nested (3)
Multilevel (4)
Select sampling
scheme (cf.
Table 2) and
sample size (cf.
Table 3) for
each qualitative
and quantitative
component
Concurrent:
QUAL + QUAN
QUAL + quan
qual + QUAN
qual + quan
Notation: “qual” stands for qualitative, “quan” stands for quantitative, “+” stands for concurrent, “ ” stands for
sequential, capital letters denote high priority or weight, and lower case letters denote lower priority or weight.
Design 2 involves a concurrent design using parallel samples for the qualitative and
quantitative components of the study. An example of a Design 2 sampling design is the study
conducted by Collins (2007). The study’s purpose was to assess the relationship between
college students’ reading abilities (i.e., reading vocabulary and reading comprehension
scores obtained on a standardized reading test) and students’ responses to three
questionnaires that measured their attitudes about reading-based assignments, such as
writing papers, using library resources, and implementing effective study habits. To
triangulate students’ responses to the questionnaires, an open-ended interview protocol also
was administered. The sample consisted of two sets of undergraduate students enrolled in
two developmental reading courses. Both samples completed the standardized test. The first
sample completed the three questionnaires that measured their attitudes about reading-based
assignments. The second sample completed the open-ended interview protocol. Collins’
combined sampling schemes can be classified as being Type 4 (cf. Table 1) because these
samples were selected purposively (i.e., homogeneous sampling scheme).
295 The Qualitative Report June 2007
Design 3 involves a concurrent design using nested samples for the qualitative and
quantitative components of the study. An example of a Design 3 sampling design is the study
conducted by Hayter (1999), whose purpose was to: (a) describe the prevalence and nature of
burnout in clinical nurse specialists in HIV/AIDS care working in community settings and
(b) examine the association between burnout and HIV/AIDS care-related factors among this
group. In the first stage of the study, the quantitative phase, 32 community HIV/AIDS nurse
specialists were administered measures of burnout and the psychological impact of working
with people with HIV/AIDS, as well as a demographic survey. In the second stage, the
qualitative phase, five nurse specialists were randomly sampled for semi-structured
interview. Because the quantitative phase involved convenient sampling and the qualitative
phase involved random sampling, Hayter’s combined sampling schemes can be classified as
being Type 3 (cf. Table 1).
Design 4 involves a concurrent design using multilevel samples for the qualitative
and quantitative components of the study. An example of a Design 4 sampling design is the
study conducted by Savaya, Monnickendam, and Waysman (2000). The purpose of the study
was to evaluate a decision support system (DSS) designed to assist youth probation officers
in choosing their recommendations to the courts. In the qualitative component, analysis of
documents and interviews of senior administrators were conducted. In the quantitative
component, youth probation officers were surveyed to determine their utilization of DSS in
the context of their work. Savaya et al.’s combined sampling schemes can be classified as
being Type 4 (cf. Table 1) because these samples were selected purposively (i.e., maximum
variation).
Design 5 involves a sequential design using identical samples for both qualitative and
quantitative components of the study. An example of a Design 5 sampling design is Taylor
and Tashakkori’s (1997) investigation, in which teachers were classified into four groups
based on their quantitative responses to measures of: (a) efficacy (low vs. high) and (b) locus
of causality for student success (i.e., internal vs. external). Then these four groups of
teachers were compared with respect to obtained qualitative data, namely, their reported
desire for and actual participation in decision making. Thus, the quantitative data collection
and analysis represented the first phase, whereas the qualitative data collection and analysis
represented the second phase. Because these identical samples were selected purposively,
Taylor and Tashakkori’s combined sampling schemes can be classified as being Type 4 (cf.
Table 1).
Design 6 involves a sequential design using parallel samples for the qualitative and
quantitative components of the study. An example of a Design 6 sampling design is the study
conducted by Scherer and Lane (1997). These researchers conducted a mixed methods study
to determine the needs and preferences of consumers (i.e., individuals with disabilities)
regarding rehabilitation services and assistive technologies. In the quantitative phase of the
study, consumers were surveyed to identify the assistive products that they perceived as
needing improvement. In the qualitative component of the study, another sample of
consumers participated in focus groups to assess the quality of the assistive products, defined
in the quantitative phase, according to specific criteria (e.g., durability, reliability,
affordability). Scherer and Lane’s combined sampling schemes can be classified as being
Type 4 (cf. Table 1) because these samples were selected purposively (i.e., homogeneous
samples).
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 296
Design 7 involves a sequential design using nested samples for the qualitative and
quantitative components of the study. An example of a Design 7 sampling design is the study
conducted by Way, Stauber, Nakkula, and London (1994). These researchers administered
questionnaires that focused in the areas of depression and substance use/abuse to students in
urban and suburban high schools (quantitative phase). On finding a positive relationship
between depression and substance use only in the suburban sample, the researchers
undertook in-depth interviews of the most depressed urban and suburban students
(qualitative phase). Here, the selection of study participants who represented the most
depressed students yielded a nested sample. The quantitative phase utilized a convenience
sample, whereas the qualitative phase employed extreme case sampling. Because both of
these sampling techniques are purposive, Way et al.’s combined sampling schemes can be
classified as being Type 4 (cf. Table 1).
Finally, Design 8 involves a sequential design using multilevel samples for the
qualitative and quantitative components of the study. An example of a Design 8 sampling
design is the study conducted by Blattman, Jensen, and Roman (2003). The study’s purpose
was to evaluate the possible socio-economic development opportunities available in a rural
community located in India. These researchers conducted both field interviews of individuals
from a variety of professional backgrounds (e.g., farmers, laborers, government workers,
educators, students) and focus groups (i.e., men, women, farmers, laborers) to obtain their
perspectives regarding the sources of development opportunities available for various
community agents, specifically farmers. Data obtained from the qualitative component were
utilized to develop a household survey questionnaire (i.e., quantitative component). The
questionnaire was distributed in two stages to two samples. The first sample (i.e., purposive;
homogeneous sampling scheme) was drawn from households representing a cross-section of
selected villages that typified the region and reflected villages of varying size, caste
composition, and access to telecommunications and agricultural and non agricultural
activities. In the second stage, random sampling procedures were used to select a different
subset of households that represented approximately 10% of the population of the selected
villages. Blattman et al.’s combined sampling schemes can be classified as being Type 2 (cf.
Table 1) because the qualitative phase involved purposive sampling, utilizing a maximum
variation sampling schema, and the quantitative phase involved stratified random sampling.
As can be seen from these mixed methods sampling examples, each of these eight
designs could involve any of the four combinations of types of sampling schemes presented
in Table 1, which, in turn could involve a combination of any of the 24 sampling schemes
presented in Table 2. Whichever of the eight sampling designs is used, careful consideration
must be made of the sample sizes needed for both the quantitative and qualitative
components of the study, depending on the type and level of generalization of interest (cf.
Table 3).
The two-dimensional mixed methods sampling model is extremely flexible because it
can be extended to incorporate studies that involve more than two components or phases. For
example, the mixed methods sampling model can be extended for a study that incorporates a
sandwich design (Sandelowski, 2003), also called a bracketed design (Greene et al., 1989),
comprising two qualitative/quantitative phases and one quantitative/qualitative phase
297 The Qualitative Report June 2007
occurring sequentially that involves either: (a) a qualitative phase followed by a quantitative
phase followed by a qualitative phase (i.e., qual → quan → qual) or (b) a quantitative phase
followed by a qualitative phase followed by a quantitative phase (i.e., quan → qual → quan).
In either case, at the third stage, the mixed methods researcher also must decide on the
relationship of the sample to the other two samples, as well as the sampling scheme and
sample size.
The exciting aspect of mixed methods sampling model is that a researcher can create
more tailored and/or more complex sampling designs than the ones outlined here to fit a
specific research context, as well as the research goal, research objective(s), research
purpose, and research question(s). Also, it is possible for a sampling design to emerge during
a study in new ways, depending on how the research evolves. However, many of these
variants can be subsumed within these eight sampling designs.
Sampling Tenets Common to Qualitative and Quantitative Research
Onwuegbuzie (2007) identified the following four crises or challenges that
researchers face when undertaking mixed methods research: representation, legitimation,
integration, and politics. The crisis of representation refers to the fact that sampling problems
characterize both quantitative and qualitative research. With respect to quantitative research,
the majority of quantitative studies utilize sample sizes that are too small to detect
statistically significant differences or relationships. That is, in the majority of quantitative
inquiries, the statistical power for conducting null hypothesis significance tests is inadequate.
As noted by Cohen (1988), the power of a null hypothesis significance test (i.e., statistical
power) is “the probability [assuming the null hypothesis is false] that it will lead to the
rejection of the null hypothesis, i.e., the probability that it will result in the conclusion that
the phenomenon exists” (p. 4). In other words, statistical power refers to the conditional
probability of rejecting the null hypothesis (i.e., accepting the alternative hypothesis) when
the alternative hypothesis actually is true (Cohen, 1988, 1992). Simply stated, power
represents how likely it is that the researcher will find a relationship or difference that really
prevails (Onwuegbuzie & Leech, 2004a).
Disturbingly, Schmidt and Hunter (1997) reported that “the average [hypothesized]
power of null hypothesis significance tests in typical studies and research literature is in the
.40 to .60 range (Cohen, 1962, 1965, 1988, 1992; Schmidt, 1996; Schmidt, Hunter, & Urry,
1976; Sedlmeier & Gigerenzer, 1989)…[with] .50 as a rough average” (p. 40).
Unfortunately, an average hypothetical power of .5 indicates that more than one-half of all
null hypothesis significance tests in the social and behavioral science literature will be
statistically non-significant. As noted by Schmidt and Hunter (p. 40), “This level of accuracy
is so low that it could be achieved just by flipping a (unbiased) coin!” Moreover, as declared
by Rossi (1997), it is possible that “at least some controversies in the social and behavioral
sciences may be artifactual in nature” (p. 178). This represents a crisis of representation.
This crisis of representation still prevails in studies in which null hypothesis
significance testing does not take place, as is the case where only effect-size indices are
reported and interpreted. Indeed, as surmised by Onwuegbuzie and Levin (2003), effect-size
statistics represent random variables that are affected by sampling variability, which is a
function of sample size. Thus, “when the sample size is small, the discrepancy between the
sample effect size and population effect size is larger (i.e., large bias) than when the sample
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 298
size is large” (p. 140). Even in descriptive research, in which no inferential analyses are
undertaken and only descriptive statistics are presented, as long as generalizations are being
made from the sample to some target population, the small sample sizes that typify
quantitative research studies still create a crisis of representation. In addition, the fact that
the majority of studies in the social and behavioral sciences do not utilize random samples
(Shaver & Norton, 1980a, 1980b), even though “inferential statistics is based on the
assumption of random sampling from populations” (Glass & Hopkins, 1984, p. 177), affects
the external validity of findings; again, yielding a crisis of representation.
In qualitative research, the crisis of representation refers to the difficulty for
researchers in capturing lived experiences. As noted by Denzin and Lincoln (2005),
Such experience, it is argued, is created in the social text written by the
researcher. This is the representational crisis. It confronts the inescapable
problem of representation, but does so within a framework that makes the
direct link between experience and text problematic. (p. 19)
Further, according to Lincoln and Denzin (2000), the crisis of representation asks who the
Other is and whether qualitative researchers can use text to represent authentically the
experience of the Other. If this is not possible, how do interpretivists establish a social
science that includes the Other? As noted by Lincoln and Denzin, these questions can be
addressed by “including the Other in the larger research processes that we have developed”
(p. 1050), which, for some, involves various types of research (e.g., action research,
participatory research, evaluation research, clinical research, policy research, racialized
discourse, ethnic epistemologies) that can occur in a variety of settings (e.g., educational,
social, clinical, familial, corporate); for some, this involves training Others to conduct their
own research of their own communities; for some, this involves positioning Others as co-
authors; and for some, this involves Others writing auto-ethnographic accounts with the
qualitative researcher assuming the role of ensuring that the Others’ voices are heard
directly. In any case, there appears to be general agreement that there is a crisis of
representation in qualitative research.
The second crisis in mixed methods research pertains to legitimation or validity. The
importance of legitimation or what is more commonly referred to as “validity,” has been
long acknowledged by quantitative researchers. For example, extending the seminal works of
Campbell and Stanley (Campbell, 1957; Campbell & Stanley, 1963), Onwuegbuzie (2003)
presented 50 threats to internal validity and external validity that occur at the research
design/data collection, data analysis, and/or data interpretation stages of the quantitative
research process. These threats are presented in Figure 3, in what was later called the
Quantitative Legitimation Model. As illustrated in Figure 3, Onwuegbuzie identified 22
threats to internal validity and 12 threats to external validity at the research design/data
collection stage of the quantitative research process. At the data analysis stage, 21 threats to
internal validity and 5 threats to external validity were conceptualized. Finally, at the data
interpretation stage, 7 and 3 threats to internal validity and external validity were identified,
respectively. In Figure 4, Onwuegbuzie, Daniel, and Collins’ (in press) schematic
representation of instrument score validity also is provided for interested readers.
Onwuegbuzie et al. build on Messick’s (1989, 1995) conceptualization of validity to yield
what they refer to as a meta-validity model that subdivides content-, criterion-, and
299 The Qualitative Report June 2007
construct-related validity into several areas of evidence. Another useful conceptualization of
validity is that of Shadish, Cook, and Campbell (2001). These authors also build on
Campbell’s earlier work and classify research validity into four major types: statistical
conclusion validity, internal validity, construct validity, and external validity. Other selected
seminal works showing the historical development of validity in quantitative research can be
found in the following references: American Educational Research Association, American
Psychological Association, and National Council on Measurement in Education (1999);
Bracht and Glass (1968); Campbell (1957); Campbell and Stanley (1963); Cook and
Campbell (1979); Messick (1989, 1995); and Smith and Glass (1987).
With respect to the qualitative research paradigm, Denzin and Lincoln (2005) argue
for “a serious rethinking of such terms as validity, generalizability, and reliability, terms
already retheorized in postpositivist…, constructivist-naturalistic…, feminist…,
interpretive…, poststructural…, and critical…discourses. This problem asks, ‘How are
qualitative studies to be evaluated in the contemporary, poststructural moment?’” (pp. 19-
20). Part of their solution has been to reconceptualize traditional validity concepts by new
labels (Lincoln & Guba, 1985, 1990). For example, Lincoln and Guba (1985) presented the
following types: credibility (replacement for quantitative concept of internal validity),
transferability (replacement for quantitative concept of external validity), dependability
(replacement for quantitative concept of reliability), and confirmability (replacement for
quantitative concept of objectivity).
Another popular classification for validity in qualitative research was provided by
Maxwell (1992), who identified the following five types of validity:
• descriptive validity (i.e., factual accuracy of the account as documented by
the researcher);
• interpretive validity (i.e., the extent to which an interpretation of the account
represents an understanding of the perspective of the underlying group and
the meanings attached to the members’ words and actions);
• theoretical validity (i.e., the degree to which a theoretical explanation
developed from research findings is consistent with the data);
• evaluative validity (i.e., the extent to which an evaluation framework can be
applied to the objects of study, as opposed to a descriptive, interpretive, or
explanatory one); and
• generalizability (i.e., the extent to which a researcher can generalize the
account of a particular situation, context, or population to other individuals,
times, settings, or context).
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 300
Figure 3. Threats to internal and external validity.
_
___________________________________
With regard to the latter validity type, Maxwell differentiates internal generalizability
from external generalizability, with the former referring to the generalizability of a
conclusion within the underlying setting or group, and the latter pertaining to generalizability
beyond the group, setting, time, or context. According to Maxwell, internal generalizability
is typically more important to qualitative researchers than is external generalizability.
Research
Design/Data
Collection
Data
Analysis
Data
Interpretation
His tory
Maturation
Testing
Instrumentation
Statistica l Regression
Differential Selection o f Participants
Mortal ity
Selection Interaction Effects
Implementation Bias
Sample Augmentation Bias
Behavior Bias
Order Bias
Observationa l Bias
Researcher Bias
Matching Bias
Treatment
Replication
Error
Evaluation Anxiety
Mul tiple-Treatment Interference
Reactive Arrangements
Treatment Diffusion
Time x Treatment Interaction
History x Treatment In teraction
Population Valid ity
Ecological Valid ity
Temporal Valid ity
Mul tiple-Treatment Interference
Researcher Bias
Reactive Arrangements
Order Bias
Matching Bias
Speci ficity of Variables
Treatment Diffusion
Pretest x Treatment In teraction
Selection x Treatment In teraction
Statistica l Regression
Restricted Range
Mortal ity
Non-In teraction Seeking Bias
Type I – Type X Error
Observationa l Bias
Researcher Bias
Matching Bias
Treatment Replication Error
Vio lated As sumptions
Multicoll inearity
Mis-Specification Error
Effect Size
Confirm ation Bias
Statistica l Regression
Distorted Graphics
Illusory Correlation
Crud Factor
Posi tive Manifold
Causal Error
Threats to External
Val idity/External
Replication
Population Valid ity
Researcher Bias
Speci ficity of Variables
Matching Bias
Mis-Specification Error
Population Valid ity
Ecological Valid ity
Temporal Valid ity
Threats to Internal
Va lidi ty/Internal
Replication
301 The Qualitative Report June 2007
Figure 4. Schematic representation of instrument score validity.
Logical ly Based Em pirically based
Content-
Related Valid ity
Cri terion-
Related Valid ity
Cons truct-
Related Valid ity
Face Valid ity Item Valid ity
Sampling
Valid ity
Concurrent
Valid ity
Predictive
Valid ity
Substanti ve
Valid ity
Structural
Valid ity
Comparati ve
Valid ity
Outcome
Valid ity
General izabil ity
Convergent
Valid ity
Discriminant
Valid ity
Divergent
Valid ity
Onwuegbuzie (2000) conceptualized what he called the Qualitative Legitimation
Model, which contains 29 elements of legitimation for qualitative research at the following
three stages of the research process: research design/data collection, data analysis, and data
interpretation. As illustrated in Figure 5, the following threats to internal credibility are
pertinent to qualitative research: ironic legitimation, paralogical legitimation, rhizomatic
legitimation, voluptuous (i.e., embodied) legitimation, descriptive validity, structural
corroboration, theoretical validity, observational bias, researcher bias, reactivity,
confirmation bias, illusory correlation, causal error, and effect size. Also in this model, the
following threats to external credibility have been identified as being pertinent to qualitative
research: catalytic validity, communicative validity, action validity, investigation validity,
interpretive validity, evaluative validity, consensual validity, population generalizability,
ecological generalizability, temporal generalizability, researcher bias, reactivity, order bias,
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 302
and effect size. (For an in-depth discussion of each of these threats to internal credibility and
external credibility, we refer the reader to Onwuegbuzie & Leech, 2007b.)
Figure 5. Qualitative legitimation model.
Threats to
External Credibility
Threats to
Internal Credibility
Data
Analysis
Research
Design/
Data
Collection
Data
Interpretation
Population Generalizability
Ecological Generalizability
Temporal Generalizability
Catalytic Validity
Communicative Validity
Action Validity
Investigation Validity
Interpretative validity
Evaluative Validity
Consensual Validity
Researcher Bias
Reactivity
Order Bias
Effect size
Theoretical
Validity
Ironic Legitimation
Paralogical Legitimation
Rhizomatic Legitimation
Embodied Legitimation
Structural Corroboration
Confirmation Bias
Illusory Correlation
Causal Error
Effect Size
Observational Bias
Researcher Bias
Reactivity
Descriptive
Validity
Observational Bias
Researcher Bias
Because of the association with the quantitative conceptualization of the research
process, qualitative researchers have, by and large, replaced the term validity by terms such
as legitimation, trustworthiness, and credibility. The major works in the area of legitimation
in qualitative research include the following: Creswell (1998), Glaser and Strauss (1967),
Kvale (1995), Lather (1986, 1993), Lincoln and Guba (1985, 1990), Longino (1995),
Maxwell (1992, 1996), Miles and Huberman (1984, 1994), Onwuegbuzie and Leech
(2007b), Schwandt (2001), Strauss and Corbin (1998), and Wolcott (1990).
303 The Qualitative Report June 2007
In mixed method research, the crises of representation and legitimation often are
exacerbated because both the quantitative and qualitative components of studies bring to the
fore their own unique crises. In mixed methods studies, the crisis of representation refers to
the difficulty in capturing (i.e., representing) the lived experience using text in general and
words and numbers in particular. The problem of legitimation refers to the difficulty in
obtaining findings and/or making inferences that are credible, trustworthy, dependable,
transferable, and/or confirmable.
The third crisis in mixed methods research pertains to integration (Onwuegbuzie,
2007). The crisis of integration refers to the extent to which combining qualitative and
quantitative approaches can address adequately the research goal, research objective(s),
research purpose(s), and research question(s). This crisis compels mixed methods
researchers to ask questions such as the following: Is it appropriate to triangulate,
consolidate, or compare quantitative data stemming from a large random sample on equal
grounds with qualitative data arising from a small purposive sample? How much weight
should be placed on qualitative data compared to quantitative data? Are quantitatively
confirmed findings more important than findings that emerge during a qualitative study
component? When quantitative and qualitative findings contradict themselves, what should
the researcher conclude?
The fourth crisis in mixed methods research is the crisis of politics (Onwuegbuzie,
2007). This crisis refers to the tensions that arise as a result of combining quantitative and
qualitative approaches. These tensions include any conflicts that arise when different
investigators are used for the quantitative and qualitative components of a study, as well as
the contradictions and paradoxes that come to the fore when the quantitative and qualitative
data are compared and contrasted. The crisis of politics also pertains to the difficulty in
persuading the consumers of mixed methods research, including stakeholders and
policymakers, to value the results stemming from both the quantitative and qualitative
components of a study. Additionally, the crisis of politics refers to tensions ensuing when
ethical standards are not addressed within the research design. These four crises are
summarized in Table 5.
Table 5
Crises Faced by Mixed Methods Researchers
Crisis
Description
Representation
The crisis of representation refers to the fact that sampling problems
characterize both quantitative and qualitative research. It refers to the
difficulty in capturing (i.e., representing) the lived experience using
text in general and words and numbers in particular.
Quantitative Phase: This crisis prevails when the sample size used is
too small to yield adequate statistical power (i.e., reduce external
validity) and/or the non-random sampling scheme used adversely
affects generalizability (i.e., reduces external validity)
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 304
Legitimation
Integration
Politics
Qualitative Phase: This crisis refers to the difficulty in capturing lived
experiences; the direct link between experience and text is problematic.
The crisis of legitimation refers to the difficulty in obtaining findings
and/or making inferences that are credible, trustworthy, dependable,
transferable, and/or confirmable.
Quantitative Phase: This crisis involves the difficulty in obtaining
quantitative findings that possess adequate internal validity and
external validity.
Qualitative Phase: This crisis leads to the following question being
asked: How are qualitative studies to be evaluated in the contemporary,
post-structural moment? It involves the difficulty in obtaining
qualitative findings that possess adequate credibility, transferability,
dependability, and/or confirmability.
The crisis of integration refers to the extent to which combining
qualitative and quantitative approaches addresses adequately the
research goal, research objective(s), research purpose(s), and research
question(s).
This crisis refers to the tensions that arise as a result of combining
quantitative and qualitative approaches, including any conflicts that
arise when different investigators are used for the quantitative and
qualitative components of a study, the contradictions and paradoxes
that come to the fore when the quantitative and qualitative data are
compared and contrasted, the difficulty in persuading the consumers of
mixed methods research (e.g., stakeholders and policymakers) to value
the results stemming from both the quantitative and qualitative
components of a study, and the tensions ensuing when ethical standards
are not addressed within the research design.
Selecting an appropriate sampling design for the qualitative and quantitative
components of the study can be a difficult choice. Thus, guidelines are needed to help mixed
methods researchers in this selection. However, we believe that keeping in mind these four
crises should help mixed methods researchers to select optimal sampling designs. That is, we
believe that an optimal sampling design in a mixed methods study is one that allows the
researcher to address simultaneously the four aforementioned crises as adequately as
possible. In particular, representation can be enhanced by ensuring that sampling decisions
stem from the research goal (e.g., predict, understand complex phenomena), research
objective (e.g., exploration, prediction), research purpose (e.g., triangulation,
complementarity), and research question(s). As displayed in Figure 1, decisions about the
research goal, research objective, research purpose, and research questions(s) are sequential
in nature. Thus, research questions arise from the research purpose, which arise from the
research objective, which, in turn, arise from the research goal. (The importance of the
305 The Qualitative Report June 2007
research question in sampling decisions is supported by Curtis et al., 2000; Kemper et al.,
2003; and Miles & Huberman, 1994.) For example, with respect to the research goal, testing
new ideas compared to understanding complex phenomena likely will lead to a different
research objective (i.e., prediction or influence vs. exploration, description, or explanation),
research purpose (e.g., triangulation vs. expansion), and research questions; and, hence,
result in different sampling designs, sampling schemes, and sample sizes being optimal.
Representation also can be enhanced by ensuring that the sample selected for each
component of the mixed methods study is compatible with the research design (cf. Table 3).
In addition, the selected samples should generate sufficient data pertaining to the
phenomenon of interest to allow thick, rich description (Curtis et al., 2000; Kemper et al.,
2003; Miles & Huberman, 1994), thereby increasing descriptive validity and interpretive
validity (Maxwell, 1992). Such samples also should help to improve representation.
Borrowing the language from qualitative researchers, both the qualitative and quantitative
components of a study should yield data that have a realistic chance of reaching data
saturation (Flick, 1998; Morse, 1995), theoretical saturation (Strauss & Corbin, 1990), or
informational redundancy (Lincoln & Guba, 1985). Representation can be further improved
by selecting samples that allow the researcher to make statistical and/or analytical
generalizations. That is, the sampling design should allow mixed methods researchers to
make generalizations to other participants, populations, settings, contexts, locations, times,
events, incidents, activities, experiences, and/or processes; that is, the sampling design
should facilitate internal and/or external generalizations (Maxwell, 1992).
Legitimation can be enhanced by ensuring that inferences stem directly from the
extracted sample of units (Curtis et al., 2000; Kemper et al., 2003; Miles & Huberman,
1994). The selected sampling design also should increase theoretical validity, where
appropriate (Maxwell, 1992). The sampling design can enhance legitimation by
incorporating audit trails (Halpern, 1983; Lincoln & Guba, 1985).
Further, the crisis of integration can be reduced by utilizing sampling designs that
help researchers to make meta-inferences that adequately represent the quantitative and
qualitative findings and which allow the appropriate weight to be assigned. Even more
importantly, the sampling design should seek to enhance what Onwuegbuzie and Johnson
(2006) refer to as “sample integration legitimation.” This legitimation type refers to
situations in which the mixed methods researcher wants to make statistical generalizations
from the sample members to the underlying population. As noted by Onwuegbuzie and
Johnson (2006), unless the relationship between the qualitative and quantitative samples is
identical (cf. Figure 2), conducting meta-inferences by pulling together the inferences from
the qualitative and quantitative phases can pose a threat to legitimation.
Finally, the crisis of politics can be decreased by employing sampling designs that
are realistic, efficient, practical, and ethical. Realism means that the data extracted from the
samples are collected, analyzed, and interpreted by either: (a) a single researcher who
possesses the necessary competencies and experiences in both qualitative and quantitative
techniques; (b) a team of investigators consisting of researchers with competency and
experience in one of the two approaches such that there is at least one qualitative and one
quantitative researcher who are able to compare and contrast effectively their respective
findings; or (c) a team of investigators consisting of researchers with minimum competency
in both qualitative and quantitative approaches and a highly specialized skill set in one of
these two procedures. According to Teddlie and Tashakkori (2003), these combinations
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 306
represent the “three current models for professional competency and collaboration” in mixed
methods research (p. 44). Moreover, a realistic sampling design is one that “provides a really
convincing account and explanation of what is observed” (Curtis et al., 2000, p. 1003).
Efficient sampling designs support studies that can be undertaken using the available
resources (e.g., money, time, effort). As such, efficiency refers more to the scope of the
researchers (i.e., manageability). In particular, the sampling design should be compatible
with the researcher’s competencies, experiences, interests, and work style (Curtis et al.,
2000; Miles & Huberman, 1994). However, even if resources are available for a chosen
sampling design, these must also be within the scope of the potential sample members. That
is, the sampling design employed must be one from which all of the data can be collected
from the sample members. For example, the sample members should not be unduly
inconvenienced. This is what is meant by utilizing a practical sampling design. Indeed, a
practical and efficient sampling design should be one that “sets an upper bound on the
internal validity/trustworthiness and external validity/transferability of the research project”
(Kemper et al., 2003, p. 277).
Finally, an ethical sampling design is one that adheres to the ethical guidelines
stipulated by organizations such as Institutional Review Boards in order for the integrity of
the research to be maintained throughout and that all sample members are protected (cf.
American Educational Research Association [AERA], 2000; Sales & Folkman, 2002).
Further, mixed methods researchers should continually evaluate their sampling designs and
procedures for ethical and scientific appropriateness throughout the course of their studies.
In particular, as specified by the Standard I.B.6 of AERA (2000), mixed methods researchers
should provide information about their sampling designs and strategies “accurately and
sufficiently in detail to allow knowledgeable, trained researchers to understand and interpret
them.” In addition, based on their sampling designs, mixed methods researchers should write
their reports in such a way that they “Communicate the practical significance for policy,
including limits in effectiveness and in generalizability to situations, problems, and contexts”
(AERA, 2000, Standard I.B.7). Even more importantly, mixed methods researchers should
undertake the following:
1. fully inform all sample members about “the likely risks involved in the research and
of potential consequences for participants” (AERA, 2000, Standard II. B.1);
2. guarantee confidentiality (Standard II. B.2) and anonymity (Standard II. B.11);
3. avoid deception (Standard II. B.3);
4. ensure that “participants have the right to withdraw from the study at any time”
(Standard II. B.5);
5. “have a responsibility to be mindful of cultural, religious, gender, and other
significant differences within the research population in the planning, conduct, and
reporting of their research” (Standard II. B.7); and
6. “carefully consider and minimize the use of research techniques that might have
negative social consequences” (Standard II. B.7).
Furthermore, mixed method researchers should consider carefully the “implications of
excluding cases because they are less articulate or less well documented, of uncertain
reliability or difficult to access” (Curtis et al., 2000, p. 1012).
307 The Qualitative Report June 2007
Summary and Conclusions
Sampling is an important step in both the qualitative and quantitative research
process. However, sampling is even more important in the mixed methods research process
because of its increased complexity arising from the fact that the quantitative and qualitative
components bring into the setting their own problems of representation, legitimation,
integration, and politics. These combined problems are likely to yield an additive effect or a
multiplicative effect that adversely impacts the quality of data collected. Thus, it is
somewhat surprising that the issue of sampling was not included as one of Teddlie and
Tashakkori’s (2003) six issues of concern in mixed methods research. Moreover, with a few
exceptions, discussion of sampling schemes has not taken place within a mixed methods
framework. Thus, the purpose of this article has been to contribute to the discussion about
sampling issues in mixed methods research. In fact, the present essay appears to represent
the most in-depth and comprehensive discussion of sampling in mixed methods research to
date. First, we presented 24 sampling schemes that have been associated with quantitative
and/or qualitative research. We contended that the present trend of methodologists and
textbook authors of linking research paradigms to sampling schemes represents a false
dichotomy that is not consistent with practice. Second, we discussed the importance of
researchers making sample size considerations for both the quantitative and qualitative
components of mixed methods studies. We then provided sample size guidelines from the
extant literature for each of the major qualitative and quantitative research designs. Third, we
provided a typology of sampling designs in mixed methods research. Specifically, we
introduced our two-dimensional mixed methods sampling model, which demonstrated how
sampling designs can be classified according to: (a) the time orientation of the components
(i.e., concurrently vs. sequentially) and (b) the relationship of the qualitative and quantitative
samples (e.g., identical vs. nested). Fourth, we presented the four major crises or challenges
to mixed methods research: representation, legitimation, integration, and politics. These
crises were then used to provide guidelines for making sampling design considerations.
The two-dimensional mixed methods sampling model presented in this paper helps to
fulfill two goals. First and foremost, this model can help mixed methods researchers to
identify an optimal sampling design. Second, the model can be used to classify mixed
methods studies in the extant literature with respect to their sampling strategies. Indeed,
future research should build on the work of Collins et al. (2006, 2007) who investigated the
prevalence of each of the eight sampling designs presented in Figure 2. Such studies also
could identify any potential misuse of sampling designs with respect to the four crises in
mixed methods research.
Virtually all researchers (whether qualitative, quantitative, or mixed methods
researchers) make some form of generalization when interpreting their data. Typically, they
make statistical generalizations, analytic generalizations, and/or generalizations that involve
case-to-case transfer (Curtis et al., 2000; Firestone, 1993; Kennedy, 1979; Miles &
Huberman, 1994). However, the generalizing process is in no way mechanical (Miles &
Huberman, 1994). Indeed, generalization represents an active process of reflection
(Greenwood & Levin, 2000). Specifically, because all findings are context-bound; (a) any
interpretations stemming from these findings should be made only after being appropriately
aware of the context under which these results were constructed, (b) generalizations of any
interpretations to another context should be made only after being adequately cognizant of
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 308
the new context and how this new context differs from the context from which the
interpretations were generated; and (c) generalizations should occur only after the researcher
has reflected carefully on the consequences that such a generalization may have. Therefore,
choosing an optimal sampling design is an essential part of the reflection process.
Selecting a sampling design involves making a series of decisions not only about how
many individuals to include in a study and how to select these individuals, but also about
conditions under which this selection will take place. These decisions are extremely
important and, as stated by Curtis et al. (2000), “It seems essential to be explicit about these
[decisions], rather than leaving them hidden, and to consider the implications of the choice
for the way that the…study can be interpreted” (p. 1012). Unfortunately, the vast majority of
qualitative and quantitative researchers do not make clear their sampling decisions. Indeed,
the exact nature of the sampling scheme rarely is specified (Onwuegbuzie, 2002b). As such,
sampling in qualitative and quantitative research appears to be undertaken as a private
enterprise that is unavailable for public inspection. However, as noted by Curtis et al. (2000,
“careful consideration of… [sampling designs] can enhance the interpretive power of a study
by ensuring that the scope and the limitations of the analysis is clearly specified” (p. 1013).
Thus, we hope that the framework that we have provided can help mixed methods
researchers in their quest to select an optimal sampling design. Further, we hope that our
framework will motivate other research methodologists to construct alternative typologies
for helping researchers in making their sampling decisions.
References
American Educational Research Association. (2000). Ethical standards of the American
Educational Research Association. Retrieved August 25, 2007, from
http://www.aera.net/AboutAERA/Default.aspx?menu_id=90&id=222
American Educational Research Association, American Psychological Association, &
National Council on Measurement in Education (1999). Standards for educational
and psychological testing (Rev. ed.). Washington, DC: American Educational
Research Association.
Baumgartner, T. A., Strong, C. H., & Hensley, L. D. (2002). Conducting and reading
research in health and human performance (3rd ed.). New York: McGraw-Hill.
Bernard, H. R. (1995). Research methods in anthropology: Qualitative and quantitative
approaches. Walnut Creek, CA: AltaMira.
Blattman, C., Jensen, R., & Roman, R. (2003). Assessing the need and potential community
networking for development in rural India. The Information Society, 19, 349-364.
Bracht, G. H., & Glass, G. V. (1968). The external validity of experiments. American
Educational Research Journal, 5, 437-474.
Brewer, J., & Hunter, A. (1989). Multimethod research: A synthesis of style. Newbury Park,
CA: Sage.
Bryman, A. (1989). Quantity and quality in social science research. London: Routledge.
Campbell, D. T. (1957). Factors relevant to the validity of experiments in social
settings. Psychological Bulletin, 54, 297-312.
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for
research. Chicago: Rand McNally.
309 The Qualitative Report June 2007
Caracelli, V. W., & Greene, J. C. (1993). Data analysis strategies for mixed-method
evaluation designs. Educational Evaluation and Policy Analysis, 15, 195-207.
Carrese, J. A., Mullaney, J. L., & Faden, R. R. (2002). Planning for death but not serious
future illness: Qualitative study of household elderly patients. British Medical
Journal, 325(7356), 125-130.
Charles, C. M., & Mertler, C. A. (2002). Introduction to educational research (4th ed.).
Boston, MA: Allyn & Bacon.
Cohen, J. (1962). The statistical power of abnormal social psychological research: A review.
Journal of Abnormal and Social Psychology, 65, 145-153.
Cohen, J. (1965). Some statistical issues in psychological research. In B. B. Wolman (Ed.),
Handbook of clinical psychology (pp. 95-121). New York: McGraw-Hill.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale,
NJ: Lawrence Erlbaum.
Cohen, J. (1992). A power primer. Psychological Bulletin, 112, 155-159.
Collins, K. M. T. (2007). Assessing the relationship between college students’ reading
abilities and their attitudes toward reading-based assignments. Manuscript
submitted for publication.
Collins, K. M. T., Onwuegbuzie, A. J., & Jiao, Q. G. (2006). Prevalence of mixed methods
sampling designs in social science research. Evaluation and Research in Education,
19, 83-101.
Collins, K. M. T., Onwuegbuzie, A. J., & Jiao, Q. G. (2007). A mixed methods investigation
of mixed methods sampling designs in social and health science research. Journal of
Mixed Methods Research, 1, 267-294.
Cook, T. D., & Campbell, D. T. (1979). Quasi-experimentation: Design and analysis issues
for field settings. Chicago: Rand McNally.
Cook, T. D., & Reichardt, C. S. (Eds.). (1979). Qualitative and quantitative methods in
evaluation research. Beverly Hills, CA: Sage.
Creswell, J. W. (1994). Research design: Qualitative and quantitative approaches.
Thousand Oaks, CA: Sage.
Creswell, J. W. (1998). Qualitative inquiry and research design: Choosing among five
traditions. Thousand Oaks, CA: Sage.
Creswell, J. W. (1999). Mixed-method research: Introduction and application. In C. Ciznek
(Ed.), Handbook of educational policy (pp. 455-472). San Diego, CA: Academic
Press.
Creswell, J. W. (2002). Educational research: Planning, conducting, and evaluating
quantitative and qualitative research. Upper Saddle River, NJ: Pearson Education.
Creswell, J. W. (2007). Qualitative inquiry and research method: Choosing among five
approaches (2nd. ed.). Thousand Oaks, CA: Sage.
Creswell, J. W., Plano Clark, V. L., Guttmann, M. L., & Hanson, E. E. (2003). Advanced
mixed methods research design. In A. Tashakkori & C. Teddlie (Eds.), Handbook of
mixed methods in social and behavioral research (pp. 209-240). Thousand Oaks,
CA: Sage.
Curtis, S., Gesler, W., Smith, G., & Washburn, S. (2000). Approaches to sampling and case
selection in qualitative research: Examples in the geography of health. Social Science
and Medicine, 50, 1001-1014.
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 310
Daley, C. E., & Onwuegbuzie, A. J. (2004). Attributions toward violence of male juvenile
delinquents: A concurrent mixed methods analysis. Journal of Social Psychology,
144, 549-570.
Del Rio, J. A., Kostoff, R. N., Garcia, E. O., Ramirez, A. M., & Humenik, J. A. (2002).
Phenomenological approach to profile impact of scientific research: Citation mining.
Advances in Complex Systems, 5(1), 19-42.
Denzin. N. K., & Lincoln, Y. S. (2005). The discipline and practice of qualitative research.
In N. K. Denzin & Y. S. Lincoln (Eds.), Handbook of qualitative research (3rd ed.,
pp. 1-32). Thousand Oaks, CA: Sage.
Dzurec, L. C., & Abraham, J. L. (1993). The nature of inquiry: Linking quantitative and
qualitative research. Advances in Nursing Science, 16, 73-79.
Erdfelder, E., Faul, F., & Buchner, A. (1996). GPOWER: A general power analysis
program. Behavior Research Methods, Instruments, & Computers, 28, 1-11.
Firestone, W. A. (1993). Alternative arguments for generalizing from data, as applied to
qualitative research. Educational Researcher, 22(4), 16-23.
Flick, U. (1998). An introduction to qualitative research: Theory, method, and applications.
London: Sage.
Gall, M. D., Borg, W. R., & Gall, J. P. (1996). Educational research: An introduction (6th
ed.). White Plains, NY: Longman.
Gay, L. R., & Airasian, P. (2003). Educational research: Competencies foranalysis and
application (7th ed.). Upper Saddle River, NJ: Pearson Education.
Glaser, B. G., & Strauss, A. L. (1967). The discovery of grounded theory: Strategies for
qualitative research. Chicago: Aldine.
Glass, G. V., & Hopkins, K. D. (1984). Statistical methods in education and psychology.
(2nd ed.) Englewood Cliffs, NJ: Prentice-Hall.
Greene, J. C., & Caracelli, V. J. (Eds.) (1997). Advances in mixed-method evaluation: The
challenges and benefits of integrating diverse paradigms: New directions for
evaluation (No. 74). San Francisco: Jossey-Bass.
Greene, J. C., Caracelli, V. J., & Graham, W. F. (1989). Toward a conceptual framework for
mixed-method evaluation designs. Educational Evaluation and Policy Analysis, 11,
255-274.
Greene, J. C., & McClintock, C. (1985). Triangulation in evaluation: Design and analysis
issues. Evaluation Review, 9, 523-545.
Greenwood, D. J., & Levin, M. (2000). Reconstructing the relationships between
universities and society through action research. In N. K. Denzin & Y. S. Lincoln
(Eds.), Handbook of qualitative research (2nd ed., pp. 85-106). Thousand Oaks, CA:
Sage.
Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An
experiment with data saturation and variability. Field Methods, 18, 59-82.
Gueulette, C., Newgent, R., & Newman, I. (1999, December). How much of qualitative
research is really qualitative? Paper presented at the annual meeting of the
Association for the Advancement of Educational Research (AAER), Ponte Vedra,
Florida.
Halpern, E. S. (1983). Auditing naturalistic inquiries: The development and application of a
model. Unpublished doctoral dissertation, Indiana University, Bloomington.
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/searchpost.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A000206
http://0-web6.epnet.com.library.uark.edu/authHjafDetail.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A00
311 The Qualitative Report June 2007
Hayter, M. (1999). Burnout and AIDS care-related factors in HIV community clinical nurse
specialists in the north of England. Journal of Advanced Nursing, 29, 984-993.
Howe, K. R. (1988). Against the quantitative-qualitative incompatability thesis or dogmas
die hard. Educational Researcher, 17, 10-16.
Howe, K. R. (1992). Getting over the quantitative-qualitative debate. American Journal of
Education, 100, 236-256.
Jick, T. D. (1979). Mixing qualitative and quantitative methods: Triangulation in action.
Administrative Science Quarterly, 24, 602-611.
Jick, T. D. (1983). Mixing qualitative and quantitative methods: Triangulation in action. In
J. Van Mannen (Ed.), Qualitative methodology (pp. 135-148). Beverly Hills, CA:
Sage.
Johnson, R. B., & Christensen, L. B. (2004). Educational research: Quantitative,
qualitative, and mixed approaches. Boston, MA: Allyn and Bacon.
Johnson, R. B., & Onwuegbuzie, A. J. (2004). Mixed methods research: A research
paradigm whose time has come. Educational Researcher, 33(7), 14-26.
Kemper, E. A., Stringfield, S., & Teddlie, C. (2003). Mixed methods sampling strategies in
social science research. In A. Tashakkori & C. Teddlie (Eds.), Handbook of mixed
methods in social and behavioral research (pp. 273-296). Thousand Oaks, CA: Sage.
Kennedy, M. (1979). Generalizing from single case studies. Evaluation Quarterly, 3, 661-
678.
Krueger, R. A. (1994). Focus groups: A practical guide for applied research (2nd ed.).
Thousand Oaks, CA: Sage.
Krueger, R. A. (2000). Focus groups: A practical guide for applied research (3rd ed.).
Thousand Oaks, CA: Sage.
Kvale, S. (1995). The social construction of validity. Qualitative Inquiry, 1, 19-40.
Langford, B. E., Schoenfeld, G., & Izzo, G. (2002). Nominal grouping sessions vs.
focus groups. Qualitative Market Research, 5, 58-70.
Lather, P. (1986). Issues of validity in openly ideological research: Between a rock and a
soft place. Interchange, 17, 63-84.
Lather, P. (1993). Fertile obsession: Validity after poststructuralism. Sociological Quarterly,
34, 673-693.
Laurie, H., & Sullivan, O. (1991). Combining qualitative and quantitative methods in the
longitudinal study of household allocations. Sociological Review, 39, 113-139.
Leech, N. L., & Onwuegbuzie, A. J. (2002, November). A call for greater use of
nonparametric statistics. Paper presented at the annual meeting of the Mid-South
Educational Research Association, Chattanooga, TN.
Li, S., Marquart, J. M., & Zercher, C. (2000). Conceptual issues and analytical strategies in
mixed-method studies of preschool inclusion. Journal of Early Intervention, 23, 116-
132.
Liddy, E. D. (2000). Text mining. Bulletin of the American Society for Information Science
& Technology, 27(1), 14-16.
Lincoln, Y. S., & Denzin. N. K. (2000). The seventh moment. In N. K. Denzin & Y. S.
Lincoln (Eds.), Handbook of qualitative research (2nd ed., pp. 1047-1065).
Thousand Oaks, CA: Sage.
Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Beverly Hills, CA: Sage.
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 312
Lincoln, Y. S., & Guba, E. G. (1990). Judging the quality of case study reports.
International Journal of Qualitative Studies in Education, 3, 53-59.
Longino, H. (1995). Gender, politics, and the theoretical virtues. Synthese, 104, 383-397.
Maxwell, J. A. (1992). Understanding and validity in qualitative research. Harvard
Educational Review, 62, 279-299.
Maxwell, J. A. (1996). Qualitative research design. Newbury Park, CA: Sage.
Maxwell, J. A., & Loomis, D. M. (2003). Mixed methods design: An alternative approach.
In A. Tashakkori and C. Teddlie (Eds.), Handbook of mixed methods in social and
behavioral research (pp. 241-272). Thousand Oaks, CA: Sage.
McMillan, J. H., & Schumacher, S. (2001). Research in education: A conceptual
introduction (5th ed.). New York: Longman.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13-
103). Old Tappan, NJ: Macmillan.
Messick, S. (1995). Validity of psychological assessment: Validation of inferences from
persons’ responses and performances as scientific inquiry into score meaning.
American Psychologist, 50, 741-749.
Miles, M. B., & Huberman, A. M. (1984). Qualitative data analysis: A sourcebook of new
methods. Beverly Hills, CA: Sage.
Miles, M., & Huberman, A. M. (1994). Qualitative data analysis: An expanded sourcebook
(2nd ed.). Thousand Oaks, CA: Sage.
Morgan, D. L. (1997). Focus groups as qualitative research (2nd ed.). Qualitative Research
Methods Series 16. Thousand Oaks, CA: Sage.
Morgan, D. L. (1998). Practical strategies for combining qualitative and quantitative
methods: Applications to health research. Qualitative Health Research, 3, 362-376.
Morse, J. M. (1991). Approaches to qualitative-quantitative methodological triangulation.
Nursing Research, 40, 120-123.
Morse, J. M. (1994). Designing funded qualitative research. In N. K. Denzin & Y. S.
Lincoln (Eds.), Handbook of qualitative research (pp. 220-235). Thousand Oaks,
CA: Sage.
Morse, J. M. (1995). The significance of saturation. Qualitative Health Research, 5, 147-
149.
Morse, J. M. (1996). Is qualitative research complete? Qualitative Health Research, 6, 3-5.
Morse, J. M. (2003). Principles of mixed methods and multimethod research design. In A.
Tashakkori & C. Teddlie (Eds.), Handbook of mixed methods in social and
behavioral research (pp. 189-208). Thousand Oaks, CA: Sage.
Newman, I., & Benz, C. R. (1998). Qualitative-quantitative research methodology:
Exploring the interactive continuum. Carbondale, Illinois: Southern Illinois
University Press.
Newman, I., Ridenour, C. S., Newman, C., & DeMarco, G. M. P. (2003). A typology of
research purposes and its relationship to mixed methods. In A. Tashakkori & C.
Teddlie (Eds.), Handbook of mixed methods in social and behavioral research (pp.
167-188). Thousand Oaks, CA: Sage.
Onwuegbuzie, A. J. (2000, November). Validity and qualitative research: An oxymoron?
Paper presented at the annual meeting of the Association for the Advancement of
Educational Research (AAER), Ponte Vedra, Florida.
313 The Qualitative Report June 2007
Onwuegbuzie, A. J. (2002a). Positivists, post-positivists, post-structuralists, and post-
modernists: Why can’t we all get along? Towards a framework for unifying research
paradigms. Education, 122, 518-530.
Onwuegbuzie, A. J. (2002b). Common analytical and interpretational errors in educational
research: An analysis of the 1998 volume of the British Journal of Educational
Psychology. Educational Research Quarterly, 26(1), 11-22.
Onwuegbuzie, A. J. (2003). Expanding the framework of internal and external validity in
quantitative research. Research in the Schools, 10(1), 71-90.
Onwuegbuzie, A. J. (2007). Mixed methods research in sociology and beyond. In G. Ritzer
(Ed.), Encyclopedia of sociology (Vol. 6, pp. 2978-2981). Oxford, England:
Blackwell Publishing.
Onwuegbuzie, A. J., Daniel, L. G., & Collins, K. M. T. (in press). A meta-validation model
for assessing the score-validity of student teacher evaluations. Quality & Quantity:
International Journal of Methodology.
Onwuegbuzie, A. J., Dickinson, W. B., Leech, N. L., & Zoran, A. G. (2007, February).
Toward more rigor in focus group research: A new framework for collecting and
analyzing focus group data. Paper presented at the annual meeting of the Southwest
Educational Research Association, San Antonio, TX.
Onwuegbuzie, A. J., Jiao, Q. G., & Bostick, S. L. (2004). Library anxiety: Theory, research,
and applications. Lanham, MD: Scarecrow Press.
Onwuegbuzie, A. J., & Johnson, R. B. (2004). Mixed method and mixed model research. In
R. B. Johnson & L. B. Christensen, Educational research: Quantitative, qualitative,
and mixed approaches (pp. 408-431). Needham Heights, MA: Allyn & Bacon.
Onwuegbuzie, A. J., & Johnson, R. B. (2006). The validity issue in mixed research.
Research in
the Schools, 13(1), 48-63.
Onwuegbuzie, A. J., & Leech, N. L. (2004a). Post-hoc power: A concept whose time has
come. Understanding Statistics, 3, 151-180.
Onwuegbuzie, A. J., & Leech, N. L. (2004b). Enhancing the interpretation of “significant”
findings: The role of mixed methods research. The Qualitative Report, 9(4), 770-792.
Retrieved August 27, 2007, from http://www.nova.edu/ssss/QR/QR9-4/
onwuegbuzie
Onwuegbuzie, A. J., & Leech, N. L. (2005a). Taking the “Q” out of research: Teaching
research methodology courses without the divide between quantitative and
qualitative paradigms. Quality & Quantity: International Journal of Methodology,
39, 267-296.
Onwuegbuzie, A. J., & Leech, N. L. (2005b, March 10). A typology of errors and myths
perpetuated in educational research textbooks. Current Issues in Education, 8(7).
Retrieved May 29, 2007, from http://cie.asu.edu/volume8/number7/index.html
Onwuegbuzie, A. J., & Leech, N. L. (2005c). The role of sampling in qualitative research.
Academic Exchange Quarterly, 9, 280-284.
Onwuegbuzie, A. J., & Leech, N. L. (2007a). A call for qualitative power analyses. Quality
& Quantity: International Journal of Methodology, 41, 105-121.
Onwuegbuzie, A. J., & Leech, N. L. (2007b). Validity and qualitative research: An
oxymoron? Quality & Quantity: International Journal of Methodology, 41, 233-249.
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 314
Onwuegbuzie, A. J., & Leech, N. L. (2007c). Sampling designs in qualitative research:
Making the sampling process more public. The Qualitative Report, 12(2), 238-254
Retrieved August 31, 2007 from http://www.nova.edu/ssss/QR/QR12-
2/onwuegbuzie1
Onwuegbuzie, A. J., & Levin, J. R. (2003). Without supporting statistical evidence, where
would reported measures of substantive importance lead? To no good effect. Journal
of Modern Applied Statistical Methods, 2, 133-151.
Patton, M. Q. (1990). Qualitative research and evaluation methods (2nd ed.). Newbury
Park, CA: Sage.
Powis, T., & Cairns, D. (2003). Mining for meaning: Text mining the relationship between
social representations of reconciliation and beliefs about Aboriginals. Australian
Journal of Psychology, 55, 59-62.
Reichardt, C. S., & Rallis, S. F. (1994). The qualitative-quantitative debate: New directions
for program evaluation (No. 61). San Francisco: Jossey-Bass.
Rossi, J. S. (1997). A case study in the failure of psychology as a cumulative science: The
spontaneous recovery of verbal learning. In L. L. Harlow, S. A. Mulaik, & J. H.
Steiger (Eds.), What if there were no significance tests? (pp. 175-197). Mahwah, NJ:
Erlbaum.
Rossman, G. B., & Wilson, B. L. (1985). Numbers and words: Combining quantitative and
qualitative methods in a single large-scale evaluation study. Evaluation Review, 9,
627-643.
Sales, B. D., & Folkman, S. (2002). Ethics in research with human participants.
Washington, DC: American Psychological Association.
Sandelowski, M. (1995). Focus on qualitative methods: Sample sizes in qualitative research.
Research in Nursing & Health, 18, 179-183.
Sandelowski, M. (2001). Real qualitative researchers don’t count: The use of numbers in
qualitative research. Research in Nursing & Health, 24, 230-240.
Sandelowski, M. (2003). Tables or tableaux? The challenges of writing and reading mixed
methods studies. In A. Tashakkori & C. Teddlie (Eds.), Handbook of mixed methods
in social and behavioral research (pp. 321-350). Thousand Oaks, CA: Sage.
Savaya, R., Monnickendam, M., & Waysman, M. (2000). An assessment of the utilization of
a computerized decision support system for youth probation officers. Journal of
Technology in Human Services, 17(4), 1-14.
Scherer, M. J., & Lane, J. P. (1997, December). Assessing consumer profiles of “ideal”
assistive technologies in ten categories: An integration of quantitative and qualitative
methods. Disability & Rehabilitation: An International Multidisciplinary Journal,
19(12), 528-535.
Schmidt, F. L. (1996). Statistical significance testing and cumulative knowledge in
psychology: Implications for the training of researchers. Psychological Methods, 1,
115-129.
Schmidt, F. L., & Hunter, J. E. (1997). Eight common but false objections to the
discontinuation of significance testing in the analysis of research data. In L. L.
Harlow, S. A. Mulaik, & J. H. Steiger (Eds.), What if there were no significance
tests? (pp. 37-64). Mahwah, NJ: Erlbaum.
Schmidt, F. L., Hunter, J. E., & Urry, V. E. (1976). Statistical power in criterion-related
validation studies. Journal of Applied Psychology, 61, 473-485.
315 The Qualitative Report June 2007
Schwandt, T. A. (2001). Dictionary of qualitative inquiry (2nd ed.). Thousand Oaks, CA:
Sage.
Sechrest, L., & Sidana, S. (1995). Quantitative and qualitative methods: Is there an
alternative? Evaluation and Program Planning, 18, 77-87.
Sedlmeier, P., & Gigerenzer, G. (1989). Do studies of statistical power have an effect on the
power of studies? Psychological Bulletin, 105, 309-316.
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2001). Experimental and quasi-
experimental designs for generalized causal inference. Boston: Houghton Mifflin.
Shaver, J. P., & Norton, R. S. (1980a). Populations, samples, randomness, and replication in
two social studies journals. Theory and Research in Social Education, 8(2), 1-20.
Shaver, J. P., & Norton, R. S. (1980b). Randomness and replication in ten years of the
American Educational Research Journal. Educational Researcher, 9(1), 9-15.
Sieber, S. D. (1973). The integration of fieldwork and survey methods. American Journal of
Sociology, 73, 1335-1359.
Smith, M. L. (1986). The whole is greater: Combining qualitative and quantitative
approaches in evaluation studies. In D. D. Williams (Ed.), Naturalistic evaluation
(pp. 37-54). San Francisco: Jossey-Bass.
Smith, M. L., & Glass, G. V. (1987). Research and evaluation in education and the social
sciences. Englewood Cliffs, NJ: Prentice Hall.
Srinivasan, P. (2004). Generation hypotheses from MEDLINE. Journal of the American
Society for Information Science & Technology, 55, 396-413.
Strauss, A., & Corbin, J. (1990). Basics of qualitative research: Grounded theory
procedures and techniques. Newbury Park, CA: Sage.
Strauss, A., & Corbin, J. (1998). Basics of qualitative research: Techniques and procedures
for developing grounded theory. Thousand Oaks, CA: Sage.
Tashakkori, A., & Teddlie, C. (1998). Mixed methodology: Combining qualitative and
quantitative approaches. Thousand Oaks, CA: Sage.
Tashakkori, A., & Teddlie, C. (Eds.). (2003a). Handbook of mixed methods in social and
behavioral research. Thousand Oaks, CA: Sage.
Tashakkori, A., & Teddlie, C. (2003b). Issues and dilemmas in teaching research methods
courses in social and behavioral sciences: A US perspective. International Journal
of Social Research Methodology, 6(1), 61 – 77.
Tashakkori, A., & Teddlie, C. (2003c). The past and future of mixed methods research:
From data triangulation to mixed model designs. In A. Tashakkori & C. Teddlie
(Eds.), Handbook of mixed methods in social and behavioral research (pp. 671-701).
Thousand Oaks, CA: Sage.
Taylor, D. L., & Tashakkori, A. (1997). Toward an understanding of teachers’ desire for
participation in decision making. Journal of School Leadership, 7, 609-628.
Teddlie, C., & Tashakkori, A. (2003). Major issues and controversies in the use of mixed
methods in the social and behavioral sciences. In A. Tashakkori & C. Teddlie (Eds.),
Handbook of mixed methods in social and behavioral research (pp. 3-50). Thousand
Oaks, CA: Sage.
Teddlie, C., & Yu, F. (2007). Mixed methods sampling: A typology with examples. Journal
of Mixed Methods Research, 1, 77-100.
The American Heritage College Dictionary (3rd ed.). (1993). Boston: Houghton Mifflin.
http://0-web6.epnet.com.library.uark.edu/authHjafDetail.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A00
http://0-web6.epnet.com.library.uark.edu/authHjafDetail.asp?tb=1&_ug=sid+4F2E2DB4%2D5E0C%2D49D6%2DAD91%2D6FE6AFB14A2C%40sessionmgr4+dbs+aph+cp+1+CE53&_us=hd+False+hs+True+cst+0%3B1%3B2%3B3+or+Date+fh+False+ss+SO+sm+ES+sl+0+dstb+ES+ri+KAAACB1A00
Anthony J. Onwuegbuzie and Kathleen M. T. Collins 316
Way, N., Stauber, H. Y., Nakkula, M. J., & London, P. (1994). Depression and substance
use in two divergent high school cultures: A quantitative and qualitative analysis.
Journal of Youth and Adolescence, 23, 331-357.
Waysman, M., & Savaya, R. (1997). Mixed method evaluation: A case study. Evaluation
Practice, 18, 227-237.
Wolcott, H. F. (1990). On seeking-and rejecting-validity in qualitative research. In E. W.
Eisner & A. Peshkin (Eds.), Qualitative inquiry in education: The continuing debate
(pp. 121-152). New York: Teachers College Press.
Author Note
Anthony Onwuegbuzie, Ph.D., is professor in the Department of Educational
Leadership and Counseling at Sam Houston State University. He teaches courses in
doctoral-level qualitative research, quantitative research, and mixed
methods. His research topics primarily involve disadvantaged and under-served
populations such as minorities, children living in war zones, students with special
needs, and juvenile delinquents. Also, he writes extensively on qualitative,
quantitative, and mixed methodological topics.
Kathleen M. T. Collins, Ph.D., is an associate professor in the Department of
Curriculum & Instruction, University of Arkansas. She teaches courses in doctoral-level
assessment and mixed methods. Her specializations are special populations, mixed methods
research, and education of postsecondary students.
Correspondence should be addressed to Anthony J. Onwuegbuzie, Department of
Educational Leadership and Counseling, Box 2119, Sam Houston State University,
Huntsville, Texas 77341-2119; Email: tonyonwuegbuzie@aol.com
Copyright 2007: Anthony J. Onwuegbuzie, Kathleen M. T. Collins, and Nova
Southeastern University
Article Citation
Onwuegbuzie, A. J., & Collins, K. M. T. (2007). A typology of mixed methods sampling
designs in social science research. The Qualitative Report, 12(2), 281-316. Retrieved
[Insert date], from http://www.nova.edu/ssss/QR/QR12-2/onwuegbuzie2
Shaver, J. P., & Norton, R. S. (1980b). Randomness and replication in ten years of the American Educational Research Journal. Educational Researcher, 9(1), 9-15.
Education Research and Perspectives, Vol.38, No.1
105
Validity and Reliability in Social Science Research
Ellen A. Drost
California State University, Los Angeles
Concepts of reliability and validity in social science research are
introduced and major methods to assess reliability and validity reviewed
with examples from the literature. The thrust of the paper is to provide
novice researchers with an understanding of the general problem of
validity in social science research and to acquaint them with approaches
to developing strong support for the validity of their research.
Introduction
An important part of social science research is the quantification
of human behaviour — that is, using measurement instruments to
observe human behaviour. The measurement of human behaviour
belongs to the widely accepted positivist view, or empirical-
analytic approach, to discern reality (Smallbone & Quinton, 2004).
Because most behavioural research takes place within this
paradigm, measurement instruments must be valid and reliable.
The objective of this paper is to provide insight into these two
important concepts, and to introduce the major methods to assess
validity and reliability as they relate to behavioural research. The
paper has been written for the novice researcher in the social
sciences. It presents a broad overview taken from traditional
literature, not a critical account of the general problem of validity
of research information.
The paper is organised as follows. The first section presents what
reliability of measurement means and the techniques most
frequently used to estimate reliability. Three important questions
researchers frequently ask about reliability are discussed: (1) what
Address for correspondence: Dr. Ellen Drost, Department of
Management, College of Business and Economics, California State
University, Los Angeles, 5151 State University Drive, Los Angeles, CA
90032. Email: edrost@calstatela.edu.
Ellen Drost
106
affects the reliability of a test?, (2) how can a test be made more
reliable?, and (3) what is a satisfactory level of reliability? The
second section presents what validity means and the methods to
develop strong support for validity in behavioural research. Four
types of validity are introduced: (1) statistical conclusion validity,
(2) internal validity, (3) construct validity and (4) external validity.
Approaches to substantiate them are also discussed. The paper
concludes with a summary and suggestions.
Reliability
Reliability is a major concern when a psychological test is used to
measure some attribute or behaviour (Rosenthal and Rosnow,
1991). For instance, to understand the functioning of a test, it is
important that the test which is used consistently discriminates
individuals at one time or over a course of time. In other words,
reliability is the extent to which measurements are repeatable –
when different persons perform the measurements, on different
occasions, under different conditions, with supposedly alternative
instruments which measure the same thing. In sum, reliability is
consistency of measurement (Bollen, 1989), or stability of
measurement over a variety of conditions in which basically the
same results should be obtained
(Nunnally, 1978).
Data obtained from behavioural research studies are influenced by
random errors of measurement. Measurement errors come either in
the form of systematic error or random error. A good example is a
bathroom scale (Rosenthal and Rosnow, 1991). Systematic error
would be at play if you repeatedly weighed yourself on a
bathroom scale which provided you with a consistent measure of
your weight, but was always 10lb. heavier than it should be.
Random error would be at work if the scale was accurate, but you
misread it while weighing yourself. Consequently, on some
occasions, you would read your weight as being slightly higher
and on other occasions as slightly lower than it actually was.
These random errors would, however, cancel out, on the average,
over repeated measurements on a single person. On the other
hand, systematic errors do not cancel out; these contribute to the
Validity and reliability in social science research
107
mean score of all subjects being studied, causing the mean value
to be either too big or too small. Thus, if a person repeatedly
weighed him/herself on the same bathroom scale, he/she would
not get the exact same weight each time, but assuming the small
variations are random and cancel out, he/she would estimate
his/her weight by averaging the values. However, should the scale
always give a weight that is 10lb. too high, taking the average will
not cancel this systematic error, but can be compensated for by
subtracting 10lb. from the person‘s average weight. Systematic
errors are a main concern of validity.
There are many ways that random errors can influence
measurements in tests. For example, if a test only contains a small
number of items, how well students perform on the test will
depend to some extent on their luck in knowing the right answers.
Also, when a test is given on a day that the student does not feel
well, he/she might not perform as strongly as he/she would
normally. Lastly, when the student guesses answers on a test, such
guessing adds an element of randomness or unreliability to the
overall test results (Nunnally, 1978).
In sum, numerous sources of error may be introduced by the
variations in other forms of the test, by the situational factors that
influence the behaviour of the subjects under study, by the
approaches used by the different examiners, and by other factors
of influence. Hence, the researcher (or science, in general) is
limited by the reliability of the measurement instruments and/or by
the reliability with which he/she uses them.
Somewhat confusing to the novice researcher is the notion that a
reliable measure is not necessarily a valid measure. Bollen (1990)
explains that reliability is that part of a measure that is free of
purely random error and that nothing in the description of
reliability requires that the measure be valid. It is possible to have
a very reliable measure that is not valid. The bathroom scale
example described earlier clearly illustrates this point. Thus,
reliability is a necessary but not a sufficient condition for validity
(Nunnally, 1978).
Ellen Drost
108
Estimates of reliability
Because reliability is consistency of measurement over time or
stability of measurement over a variety of conditions, the most
commonly used technique to estimate reliability is with a measure
of association, the correlation coefficient, often termed reliability
coefficient (Rosnow and Rosenthal, 1991). The reliability
coefficient is the correlation between two or more variables (here
tests, items, or raters) which measure the same thing.
Typical methods to estimate test reliability in behavioural research
are: test-retest reliability, alternative forms, split-halves, inter-rater
reliability, and internal consistency. There are three main concerns
in reliability testing: equivalence, stability over time, and internal
consistency. These concerns and approaches to reliability testing
are depicted in Figure 1. Each will be discussed next.
Test-retest reliability. Test-retest reliability refers to the temporal
stability of a test from one measurement session to another. The
procedure is to administer the test to a group of respondents and
then administer the same test to the same respondents at a later
date. The correlation between scores on the identical tests given at
different times operationally defines its test-retest reliability.
Despite its appeal, the test-retest reliability technique has several
limitations (Rosenthal & Rosnow, 1991). For instance, when the
interval between the first and second test is too short, respondents
might remember what was on the first test and their answers on
the second test could be affected by memory. Alternatively, when
the interval between the two tests is too long, maturation happens.
Maturation refers to changes in the subject factors or respondents
(other than those associated with the independent variable) that
occur over time and cause a change from the initial measurements
to the later measurements (t and t + 1). During the time between
the two tests, the respondents could have been exposed to things
which changed their opinions, feelings or attitudes about the
behaviour under study.
Validity and reliability in social science research
109
Figure 1. Reliability of Measurement Tests
RELIABILITY
Alternative
Forms
Stability over
time
VALIDITY
Equivalence
VALIDITY
Internal
Consistency
VALIDITY
Test-Retest
Split-Half Inter-rater Cronbach
Alpha
Ellen Drost
110
Alternative forms. The alternative forms technique to estimate
reliability is similar to the test retest method, except that different
measures of a behaviour (rather than the same measure) are
collected at different times (Bollen, 1989). If the correlation
between the alternative forms is low, it could indicate that
considerable measurement error is present, because two different
scales were used. For example, when testing for general spelling,
one of the two independently composed tests might not test
general spelling but a more subject-specific type of spelling such
as business vocabulary. This type of measurement error is then
attributed to the sampling of items on the test. Several of the limits
of the test-retest method are also true of the alternative forms
technique.
Split-half approach. The split-half approach is another method
to test reliability which assumes that a number of items are
available to measure a behaviour. Half of the items are combined
to form one new measure and the other half is combined to form
the second new measure. The result is two tests and two new
measures testing the same behaviour. In contrast to the test-retest
and alternative form methods, the split-half approach is usually
measured in the same time period. The correlation between the
two halves tests must be corrected to obtain the reliability
coefficient for the whole test (Nunnally, 1978; Bollen,
1989).
There are several aspects that make the split-halves approach more
desirable than the test-retest and alternative forms methods. First,
the effect of memory discussed previously does not operate with
this approach. Also, a practical advantage is that the split-halves
are usually cheaper and more easily obtained than over time data
(Bollen, 1989).
A disadvantage of the split-half method is that the tests must be
parallel measures – that is, the correlation between the two halves
will vary slightly depending on how the items are divided.
Nunnally (1978) suggests using the split-half method when
measuring variability of behaviours over short periods of time
when alternative forms are not available. For example, the even
Validity and reliability in social science research
111
items can first be given as a test and, subsequently, on the second
occasion, the odd items as the alternative form. The corrected
correlation coefficient between the even and odd item test scores
will indicate the relative stability of the behaviour over that period
of time.
Interrater reliability. When raters or judges are used to measure
behaviour, the reliability of their judgments or combined internal
consistency of judgments is assessed (Rosenthal & Rosnow,
1991). Below in table format is an example of two judges rating
10 persons on a particular test (i.e., judges rating people‘s
competency in their writing skills).
Judge 1 Rating Judge 2 Rating
Subject 1 —- Subject 1 —-
—- —- —- —-
Subject 10 —- Subject 10 —-
The correlation between the ratings made by the two judges will
tell us the reliability of either judge in the specific situation. The
composite reliability of both judges, referred to as effective
reliability, is calculated using the Spearman-Brown formula (see
Rosenthal & Rosnow, 1991, pp. 51-55).
Internal consistency. Internal consistency concerns the
reliability of the test components. Internal consistency measures
consistency within the instrument and questions how well a set of
items measures a particular behaviour or characteristic within the
test. For a test to be internally consistent, estimates of reliability
are based on the average intercorrelations among all the single
items within a test.
The most popular method of testing for internal consistency in the
behavioural sciences is coefficient alpha. Coefficient alpha was
popularised by Cronbach (1951), who recognised its general
usefulness. As a result, it is often referred to as Cronbach’s alpha.
Coefficients of internal consistency increase as the number of
items goes up, to a certain point. For instance, a 5-item test might
Ellen Drost
112
correlate .40 with true scores, and a 12-item test might correlate
.80 with true scores.
Consequently, the individual item would be expected to have only
a small correlation with true scores. Thus, if coefficient alpha
proves to be very low, either the test is too short or the items have
very little in common. Coefficient alpha is useful for estimating
reliability for item-specific variance in a unidimentional test
(Cortina, 1993). That is, it is useful once the existence of a single
factor or construct has been determined (Cortina, 1993). Next in
conclusion of this section, three important questions researchers
frequently ask about reliability are considered.
What factors affect the reliability of a test?
There are many factors that prevent measurements from being
exactly repeatable or replicable. These factors depend on the
nature of the test and how the test is used (Nunnally, 1978). It is
important to make a distinction between errors of measurement
that cause variation in performance within a test, and errors of
instrumentation that are apparent only in variation in performance
on different forms of a test.
Sources of error within a test. A major source of error within a
test is attributable to the sampling of items. Because each person
has the same probability of answering an item correctly, the higher
the number of items on the test, the lower the amount of error in
the test as a whole. However, error due to item sampling is
entirely predictable from the average correlation, thus coefficient
alpha would be the correct measure of reliability. Other examples
of sources of errors on tests are: guessing on a test, marking
answers incorrectly (clerical errors), skipping a question
inadvertently, and misinterpreting test instructions.
On subjective tests, such as essay tests, measurement errors are
often caused by fluctuations in standards by the individual grader
and by the differences in standards of different graders. For
example, on an essay examination the instructor might grade all
Validity and reliability in social science research
113
answers to question 1, then grade all answers to question 2, and so
forth. If these scores are independent, then the average correlation
among the questions can be used to obtain an accurate estimate of
reliability. On the other hand, if half the questions are scored by
one person and the other half are independently scored by another
person, then the correlation between the two half-tests will provide
an estimate of the reliability. Thus, for any test, the sampling of
items from a domain includes the sampling of situational factors.
Variation between tests. There are two major sources of error
which intervene between administrations of different tests: (1)
systematic differences in content of the two tests, and (2)
respondents‘ change with regard to the attribute being measured.
Systematic differences in the content of two tests and in variations
in people from one occasion to another cannot be adequately
handled by random sampling of items. In this case, the tests should
be thought of as random samples of particular occasions, and
correlations among tests are allowed to be slightly lower than
would be predicted from the correlations among items within tests
(Nunnally, 1978). The average correlation among a number of
alternative tests completed on different occasions would then be a
better estimate of reliability than that given by coefficient alpha
for one test administered on one occasion only.
How can I make a test more reliable?
Reliability can be improved by writing items clearly, making test
instructions easily understood, and training the raters effectively
by making the rules for scoring as explicit as possible (Nunnally,
1978), for instance.
The principal method to make tests more reliable is to make them
longer, thus adding more items. For reliability and other reasons in
psychometrics, the maxim holds that, other things being equal, a
long test is a good test (from Nunnally, p. 243). However, the
longer the test, the more likely that boredom and fatigue, among
Ellen Drost
114
other factors, can produce attenuation (reduction) in the
consistency of accurate responding (Rosenthal & Rosnow, 1991).
What is a satisfactory level of reliability?
A satisfactory level of reliability depends on how a measure is
being used. The standard is taken from Nunnally (1978), who
suggests that in the early stages of research on predictor tests or
hypothesised measures of a construct, reliabilities of .70 or higher
will be sufficient. During this stage, Nunnally (1978) maintains
that increasing reliabilities much beyond .80 are often wasteful of
time and funds, because correlations at that level are attenuated
very little by measurement error. To obtain a higher reliability of
.90, for instance, requires strenuous efforts at standardisation and
probably an addition of items.
On the other hand, in applied settings where important decisions
are made with respect to specific test scores, Nunnally (1978)
recommends that a reliability of at least .90 is desirable, because a
great deal depends on the exact score made by a person on a test.
A good example is given for children with low IQs below 70 who
are placed in special classes. In this case, it makes a big difference
whether the child has an IQ of 65 or 75 on a particular test. Next,
the discussion will focus on validity in research.
Validity
Validity is concerned with the meaningfulness of research
components. When researchers measure behaviours, they are
concerned with whether they are measuring what they intended to
measure. Does the IQ test measure intelligence? Does the GRE
actually predict successful completion of a graduate study
program? These are questions of validity and even though they
can never be answered with complete certainty, researchers can
develop strong support for the validity of their measures (Bollen,
1989).
Validity and reliability in social science research
115
There are four types of validity that researchers should consider:
statistical conclusion validity, internal validity, construct validity,
and external validity. Each type answers an important question
and is discussed next.
Statistical conclusion validity
Does a relationship exist between the two variables? Statistical
conclusion validity pertains to the relationship being tested.
Statistical conclusion validity refers to inferences about whether it
is reasonable to presume covariation given a specified alpha level
and the obtained variances (Cook & Campbell, 1979). There are
some major threats to statistical conclusion validity such as low
statistical power, violation of assumptions, reliability of measures,
reliability of treatment, random irrelevancies in the experimental
setting, and random heterogeneity of respondents.
Internal validity
Given that there is a relationship, is the relationship a causal one?
Are there no confounding factors in my study? Internal validity
speaks to the validity of the research itself. For example, a
manager of a company tests employees on leadership satisfaction.
Only 50% of the employees responded to the survey and all of
them liked their boss. Does the manager have a representative
sample of employees or a bias sample? Another example would be
to collect a job satisfaction survey before Christmas just after
everybody received a nice bonus. The results showed that all
employees were happy. Again, do the results really indicate job
satisfaction in the company or do the results show a bias?
There are many threats to internal validity of a research design.
Some of these threats are: history, maturation, testing,
instrumentation, selection, mortality, diffusion of treatment and
compensatory equalisation, rivalry and demoralisation. A
discussion of each threat is beyond the scope of this paper.
Ellen Drost
116
Construct validity
If a relationship is causal, what are the particular cause and effect
behaviours or constructs involved in the relationship? Construct
validity refers to how well you translated or transformed a
concept, idea, or behaviour – that is a construct – into a
functioning and operating reality, the operationalisation (Trochim,
2006). To substantiate construct validity involves accumulating
evidence in six validity types: face validity, content validity,
concurrent and predictive validity, and convergent and
discriminant validity. Trochim (2006) divided these six types into
two categories: translation validity and criterion-related validity.
These two categories and their respective validity types are
depicted in Figure 2 and discussed in turn, next.
Translation Validity. Translation validity centres on whether the
operationalisation reflects the true meaning of the construct.
Translation validity attempts to assess the degree to which
constructs are accurately ―translated‖ into the operationalisation,
using subjective judgment – face validity – and examining content
domain – content validity.
Face Validity. Face validity is a subjective judgment on the
operationalisation of a construct. For instance, one might look at a
measure of reading ability, read through the paragraphs, and
decide that it seems like a good measure of reading ability. Even
though subjective judgment is needed throughout the research
process, the aforementioned method of validation is not very
convincing to others as a valid judgment. As a result, face validity
is often seen as a weak form of construct validity.
Validity and reliability in social science research
117
Figure 2. Construct Validity Types
CONSTRUCT
VALIDITY
Convergent
Validity
Content
Validity
Predictive
Validity
Face Validity
Concurrent
Validity
Discriminant
Validity
Translation
Validity
VALIDITY
Criterion-
Related
Validity
VALIDITY
Ellen Drost
118
Content validity. Bollen (1989) defined content validity as ―a
qualitative type of validity where the domain of the concept is
made clear and the analyst judges whether the measures fully
represent the domain (p.185). According to Bollen, for most
concepts in the social sciences, no consensus exists on theoretical
definitions, because the domain of content is ambiguous.
Consequently, the burden falls on the researcher not only to
provide a theoretical definition (of the concept) accepted by
his/her peers but also to select indicators that thoroughly cover its
domain and dimensions. Thus, content validity is a qualitative
means of ensuring that indicators tap the meaning of a concept as
defined by the researcher. For example, if a researcher wants to
test a person‘s knowledge on elementary geography with a paper-
and-pencil test, the researcher needs to be assured that the test is
representative of the domain of elementary geography. Does the
survey really test a person‘s knowledge in elementary geography
(i.e. the location of major continents in the world) or does the test
require a more advanced knowledge in geography (i.e. continents‘
topography and their effect on climates, etc.)? There are basically
two ways of assessing content validity: (1) ask a number of
questions about the instrument or test; and/or (2) ask the opinion
of expert judges in the field.
Criterion-related validity. Criterion-related validity is the degree
of correspondence between a test measure and one or more
external referents (criteria), usually measured by their correlation.
For example, suppose we survey employees in a company and ask
them to report their salaries. If we had access to their actual salary
records, we could assess the validity of the survey (salaries
reported by the employees) by correlating the two measures. In
this case, the employee records represent an (almost) ideal
standard for comparison.
Concurrent Validity and Predictive Validity. When the criterion
exists at the same time as the measure, we talk about concurrent
validity. Concurrent ability refers to the ability of a test to predict
Validity and reliability in social science research
119
an event in the present. The previous example of employees‘
salary is an example of concurrent validity.
When the criterion occurs in the future, we talk about predictive
validity. For example, predictive validity refers to the ability of a
test to measure some event or outcome in the future. A good
example of predictive validity is the use of students‘ GMAT
scores to predict their successful completion of an MBA program.
Another example is to use students‘ GMAT scores to predict their
GPA in a graduate program. We would use correlations to assess
the strength of the association between the GMAT score with the
criterion (i.e., GPA).
Convergent and Discriminant Validity. Campbell and Fiske‘s
(1951) proposed to assess construct validity by examining their
convergent and discriminant validity. The authors posited that
construct validity can be best understood through two construct-
validation processes: first, testing for convergence across different
measures or manipulations of the same ―thing‖, and second,
testing for divergence between measures and manipulations of
related but conceptually distinct ―things‖ (Cook & Campbell,
1979, p. 61). In order to accumulate such evidence, Campbell and
Fiske proposed the use of a multitrait-multimethod (MTMM)
correlation matrix. This MTMM matrix allows one to zero in on
the convergent and discriminant validity of a construct by
investigating the intercorrelations of the matrix. The principle
behind the MTMM matrix of measuring the same and differing
behaviour is that it avoids the difficulty that high or low
correlations may be due to their common method of measurement
rather than convergent or discriminant validity. The table below
shows how the MTMM matrix works.
Method 1 Method 2 Method 3 Method 4 Method 5
Behaviours 12345 12345 12345 12345 12345
The matrix represents 5 different methods and 5 different
behaviours. Convergent validity for Trait 1 is established if Trait 1
Ellen Drost
120
(T1) measured by Method 1 (M1) correlates highly with Trait 1
(T1) and Method 2 (M2), resulting in (T1M2) and so on for
T1M3, T1M4, and T1M5. Discriminant validity for Trait 1 is
established when there are no correlations among TI Ml and the
other four traits (T2,T3,T4,T5) measured by all five methods
(M2,M3,M4,M5).
A prevalent threat to construct validity is common method
variance. Common method variance is defined as the overlap in
variance between two variables ascribed to the type of
measurement instrument used rather than due to a relationship
between the underlying constructs (Avolio, Yammarino & Bass,
1991). Cook and Campbell (1979) used the terms mono-operation
bias and mono-method bias, while Fiske (1982) adopted the term
methods variance when discussing convergent and discriminant
validation in research. Mono-operation bias represents the single
operationalisation of a construct (behaviour) rather than gathering
additional data from alternative measures of a construct.
External validity
If there is a causal relationship from construct X to construct Y,
how generalisable is this relationship across persons, settings, and
times? External validity of a study or relationship implies
generalising to other persons, settings, and times. Generalising to
well-explained target populations should be clearly differentiated
from generalising across populations. Each is truly relevant to
external validity: the former is critical in determining whether any
research objectives which specified populations have been met,
and the latter is crucial in determining which different populations
have been affected by a treatment to assess how far one can
generalise (Cook & Campbell, 1979).
For instance, if there is an interaction between an educational
treatment and the social class of children, then we cannot infer that
the same result holds across social classes. Thus, Cook and
Campbell (1979) prefer generalising across achieved (my
emphasis) populations, in which case threats to external validity
Validity and reliability in social science research
121
relate to statistical interaction effects. This implies that
interactions of selection and treatment refer to the categories of
persons to which a cause-effect relationship can be generalised.
Interactions of setting and treatment refer to whether a causal
relationship obtained in one setting can be generalised to another.
For example, can the causal relationship observed in a
manufacturing plant be replicated in a public institution, in a
bureaucracy, or on a military base? This question could be
addressed by varying settings and then analysing for a causal
relationship within each setting.
Conclusion
This paper was written to provide the novice researcher with
insight into two important concepts in research methodology:
reliability and validity. Based upon recognised and classical works
from the literature, the paper has clarified the meaning of
reliability of measurement and the general problem of validity in
behavioural research. The most frequently used techniques to
assess reliability and validity were presented to highlight their
conceptual relationships. Three important questions researchers
frequently ask about what affects reliability of their measures and
how to improve reliability were also discussed with examples
from the literature. Four types of validity were introduced:
statistical conclusion validity, internal validity, construct validity
and external validity or generalisability. The approaches to
substantiate the validity of measurements have also been presented
with examples from the literature. A final discussion on common
method variance has been provided to highlight this prevalent
threat to validity in behavioural research. The paper was intended
to provide an insight into these important concepts and to
encourage students in the social sciences to continue studying to
advance their understanding of research methodology.
Ellen Drost
122
References
Avolio, B. J. , Yammanno, F. J. and Bass, B. M. (1991).
Identifying Common Methods Variance With Data Collected
From A Single Source: An Unresolved Sticky Issue. Journal of
Management, 17 (3), 571-587.
Bollen, K. A. (1989). Structural Equations with Latent Variables
(pp. 179-225). John Wiley & Sons,
Brinberg, D. and McGrath, J. E. (1982). A Network of Validity
Concepts Within the Research Process. In Brinberg, D. and
Kidder, L. H., (Eds), Forms of Validity in Research, pp. 5-23.
Campbell, D.T. and Fiske, D.W. (1959). Convergent and
discriminant validation by the multitrait-multimethod matrix.
Psychological Bulletin, 56, 81-105.
Chapman, L.J. and Chapman, J.P. (1969). Illusory correlations as
an obstacle to use of valid psychodiagnostic signs. Journal of
Abnormal Psychology, 74, 271-280.
Cook, T. D. and Campbell, D. T. (1979). Quasi-Experimentation:
Design & Analysis Issues for Field Settings. Boston: Houghton
Muffin Company, pp. 37- 94.
Cortina, J. M. (1993). What is Coefficient Alpha? An Examination
of Theory and Applications. Journal of Applied Psychology, 78
(1), 98-104.
Cronbach, L. J. (1951). Coefficient alpha and the internal structure
of tests. Psychometrika, 16(3), 297-334.
Dansereau, F., Alutto, J.A. and Yammarino, F.J. (1984). Theory
testing in organizational behavior: The variant approach.
Englewood cliffs, NJ: Prentice Hall.
Fiske, Donald W. (1982). Convergent–Discriminant Validation in
Measurements and Research Strategies. In Brinberg, D. and
Kidder, L. H., (Eds), Forms of Validity in Research, pp. 77-93.
Nunnally, J. C. (1978). Psychometric Theory. McGraw-Hill Book
Company, pp. 86-113, 190-255.
Nunnaly, J. D. and Bernstein, I. H. (1994). Psychometric Theory.
New York, NY: McGraw Hill.
Podsakoff, P M and Organ, D W (1986). Self-reports in
organizational ‗ research: Problems and prospects. Journal of
Management, 12: 531-544.
Validity and reliability in social science research
123
Miller, M. B. (1995). Coefficient Alpha: A Basic Introduction
from the Perspective of Classical Test Theory. Structural
Equation Modeling, 2 (3), 255-273.
Rosenthal, R. and Rosnow, R. L. (1991). Essentials of Behavioral
Research: Methods and Data Analysis. Second Edition.
McGraw-Hill Publishing Company, pp. 46-65.
Shadish, W. R., Cook, T.D., and Campbell, D. T. (2001).
Experimental and Quasi-Experimental Designs for
Generalized Causal Inference. Boston: Houghton Mifflin.
Smallbone, T. and Quinton, S. (2004). Increasing Business
Students‘ Confidence in Questioning the Validity and
Reliability of their Research. Electronic Journal of Business
Research Methods, 2 (2): 153-162. www.ejbrm.com
Trochim, W. M. K. (2006). Introduction to Validity. Social
Research Methods, retrieved from
www.socialresearchmethods.net/kb/introval.php, September 9,
2010.
Williams, L.J., Cote, J.A. and Buckley, M.R. (1989). Lack of
method variance in self-reported affect and perceptions at
work: Reality or artifact? Journal of Applied Psychology, 74:
462-468.
http://www.ejbrm.com/
http://www.socialresearchmethods.net/kb/introval.php
Copyright of Education Research & Perspectives is the property of University of Western Australia,
Department of Education and its content may not be copied or emailed to multiple sites or posted to a listserv
without the copyright holder’s express written permission. However, users may print, download, or email
articles for individual use.
Essay Writing Service Features
Our Experience
No matter how complex your assignment is, we can find the right professional for your specific task. Achiever Papers is an essay writing company that hires only the smartest minds to help you with your projects. Our expertise allows us to provide students with high-quality academic writing, editing & proofreading services.Free Features
Free revision policy
$10Free bibliography & reference
$8Free title page
$8Free formatting
$8How Our Dissertation Writing Service Works
First, you will need to complete an order form. It's not difficult but, if anything is unclear, you may always chat with us so that we can guide you through it. On the order form, you will need to include some basic information concerning your order: subject, topic, number of pages, etc. We also encourage our clients to upload any relevant information or sources that will help.
Complete the order form
Once we have all the information and instructions that we need, we select the most suitable writer for your assignment. While everything seems to be clear, the writer, who has complete knowledge of the subject, may need clarification from you. It is at that point that you would receive a call or email from us.
Writer’s assignment
As soon as the writer has finished, it will be delivered both to the website and to your email address so that you will not miss it. If your deadline is close at hand, we will place a call to you to make sure that you receive the paper on time.
Completing the order and download