Video summary
Reliability, validity, generalizability and credibility. Pt .1 of 3: Research Quality
Main summary
Key takeaways
Main ideas and lessons conveyed
- The video introduces four major questions/criteria for judging research quality (attributed to Shipman, cited as “1988 Shipman”).
- It emphasizes that, for assessments, the goal is less about whether a paper’s theory sounds good and more about how data were collected, analyzed, and interpreted.
- The four criteria are presented as a framework to distinguish good research from “rubbish” research.
Research quality: the 4 criteria (Shipman)
1) Reliability
- Main question: If the investigation were repeated by different researchers using the same methods, would it produce the same results?
- Meaning: Research is reliable if it gives consistent answers across:
- time/occasions
- researchers
- conditions (as much as “other things” are equal)
Key threat sources (ways reliability fails)
-
Subject error
- People (participants) naturally vary day to day, so results may differ even if the world hasn’t “fundamentally changed.”
- Common in qualitative methods like:
- observations
- depth interviews (where people won’t answer identically every day)
-
Subject bias
- Participants’ responses change because they react differently to which researcher is asking (e.g., liking one researcher and not another).
- Researchers must balance:
- building rapport without “training” participants to answer as expected
-
Observer bias
- The researcher observes/records in a biased way due to expectations about what they think will happen.
- Different researchers on different occasions can have different biases, producing inconsistent results.
2) Validity (Internal validity)
- Main question: Do the findings reflect the actual reality being investigated?
- Meaning: The researcher truly measures what they claim to be measuring—results aren’t caused by something else.
Notes on worldview/assumptions
- The speaker describes a stance that a “real world exists independently of us.”
- They also note that some philosophers/researchers (e.g., constructivists) disagree, arguing the world is socially constructed/experienced through us.
Common threats to validity (with examples)
-
History threat
- Something changes in the environment during the study, affecting outcomes unrelated to the treatment/intervention.
- Example: an air disaster happens during a study testing desensitization for fear of flying; participant anxiety rises for reasons unrelated to the program.
-
Testing effect
- Pre-testing or earlier exposure to the survey/intervention context sensitizes participants, biasing later answers.
- Example: asking about attitudes (e.g., factory farming) before the main questionnaire can make people consider issues they otherwise wouldn’t, increasing negative attitudes.
- Timing matters:
- a longer time gap can reduce the effect
- Speed reading analogy: using the same passage hours apart can make performance look better due to recall, not learning from the course.
-
Instrumentation problems
- Measurement procedures effectively change (e.g., criteria used to observe/record behavior differ between pre and post).
- Example: observers use wider vs narrower definitions of behavior due to training or consistency problems.
-
Regression to the mean
- Selecting unusual cases (very high or very low scores) causes scores to move toward average on re-test naturally, even without real change.
- Explanation: scores reflect true ability/trait + random factors (like mood, sleep, day-to-day circumstances).
- Examples:
- selecting the top 10 and bottom 10 performers may shift toward average rather than due to the intervention
- a National Student Survey ranking drop next year may reflect statistical regression rather than real decline
- Olympic-style prediction analogy: an athlete’s exceptional performance one year doesn’t guarantee the same later, partly due to chance and day conditions
-
Additional validity threats from Cook & Campbell (explicitly named):
-
Compensatory equalization
- Professionals “improve” the control group because they know another group received special treatment.
- Example: hospital staff give control patients extra attention because they know another group received a special treatment, shrinking the difference.
-
Compensatory rivalry (John Henry effect)
- Participants in the control group work harder or change behavior to “catch up.”
- Example: when one group receives new computers, those without them may work harder, minimizing differences.
- Historical anecdote: John Henry overexerts to prove superiority over a machine—illustrating extreme compensatory effort.
-
Reliability vs validity metaphor
- Reliability: arrows land consistently (same pattern).
- Validity: arrows land in the correct target center (“ball’s eye”).
- A study can be:
- reliable but not valid (consistent but consistently wrong)
3) Generalizability
- Main question: Do the results matter beyond the specific situation/sample studied?
- Meaning: Generalize from a sample/group to a broader population.
Primary threat: selection problem
- The studied group is not representative of the population you want to generalize to.
Common non-representative sampling patterns
-
Volunteers
- Volunteers differ systematically (often more time, interest, motivation).
- Example: volunteering bias in studies about secretiveness (secretive people may be less likely to volunteer).
- Sensitive topics can produce extreme groups rather than a balanced middle (e.g., patterns in volunteering related to racism).
-
Snowball sampling
- Participants recruit others like themselves, which can worsen representativeness concerns.
Other generalizability threats described
-
Unusual setting
- Even with random sampling, the chosen location might be atypical due to local differences.
-
Setting + history affecting conclusions (anthropology example)
- Anthropologists studied “peasants resistance to change” in Mexico, assuming innate cultural conservatism.
- When revisiting nearby regions, they found the earlier village’s peasants had been repressed by landlords and police; other villages with different history were more innovative.
- Lesson: “typical” conclusions can fail because the sample/context has unique historical constraints.
-
Construct effect
- Some groups may have distinct ways of thinking that are subtle and hard to detect beforehand, meaning they may not represent the population.
4) Credibility
- Main question: Is there sufficient detail and transparency for the work to be judged as trustworthy?
- Meaning: Whether research seems trustworthy considering:
- researchers’ background/expertise
- credibility of their methods and reporting
- publication and scrutiny by others
- funding sources and possible conflicts of interest
Key emphasis
- Credibility goes beyond reliability/validity/generalizability.
- Readers should ask: Who did it, why, and who paid for it?
Speaker/source identification
- Primary speaker (video presenter/lecturer): Unnamed
-
Named referenced sources/authors:
- Shipman (1988) — presented as the origin of the “four major questions/criteria”
- Cook and Campbell — referenced for threat lists to validity (especially in experimental design contexts)
-
Other named concepts/examples:
- Constructivist researchers/philosophers (general reference)
- John Henry (used to illustrate compensatory rivalry)