Video summary

Reliability, validity, generalizability and credibility. Pt .1 of 3: Research Quality

Main summary

Key takeaways

Educational

Main ideas and lessons conveyed

  • The video introduces four major questions/criteria for judging research quality (attributed to Shipman, cited as “1988 Shipman”).
  • It emphasizes that, for assessments, the goal is less about whether a paper’s theory sounds good and more about how data were collected, analyzed, and interpreted.
  • The four criteria are presented as a framework to distinguish good research from “rubbish” research.

Research quality: the 4 criteria (Shipman)

1) Reliability

  • Main question: If the investigation were repeated by different researchers using the same methods, would it produce the same results?
  • Meaning: Research is reliable if it gives consistent answers across:
    • time/occasions
    • researchers
    • conditions (as much as “other things” are equal)

Key threat sources (ways reliability fails)

  • Subject error

    • People (participants) naturally vary day to day, so results may differ even if the world hasn’t “fundamentally changed.”
    • Common in qualitative methods like:
      • observations
      • depth interviews (where people won’t answer identically every day)
  • Subject bias

    • Participants’ responses change because they react differently to which researcher is asking (e.g., liking one researcher and not another).
    • Researchers must balance:
      • building rapport without “training” participants to answer as expected
  • Observer bias

    • The researcher observes/records in a biased way due to expectations about what they think will happen.
    • Different researchers on different occasions can have different biases, producing inconsistent results.

2) Validity (Internal validity)

  • Main question: Do the findings reflect the actual reality being investigated?
  • Meaning: The researcher truly measures what they claim to be measuring—results aren’t caused by something else.

Notes on worldview/assumptions

  • The speaker describes a stance that a “real world exists independently of us.”
  • They also note that some philosophers/researchers (e.g., constructivists) disagree, arguing the world is socially constructed/experienced through us.

Common threats to validity (with examples)

  • History threat

    • Something changes in the environment during the study, affecting outcomes unrelated to the treatment/intervention.
    • Example: an air disaster happens during a study testing desensitization for fear of flying; participant anxiety rises for reasons unrelated to the program.
  • Testing effect

    • Pre-testing or earlier exposure to the survey/intervention context sensitizes participants, biasing later answers.
    • Example: asking about attitudes (e.g., factory farming) before the main questionnaire can make people consider issues they otherwise wouldn’t, increasing negative attitudes.
    • Timing matters:
      • a longer time gap can reduce the effect
    • Speed reading analogy: using the same passage hours apart can make performance look better due to recall, not learning from the course.
  • Instrumentation problems

    • Measurement procedures effectively change (e.g., criteria used to observe/record behavior differ between pre and post).
    • Example: observers use wider vs narrower definitions of behavior due to training or consistency problems.
  • Regression to the mean

    • Selecting unusual cases (very high or very low scores) causes scores to move toward average on re-test naturally, even without real change.
    • Explanation: scores reflect true ability/trait + random factors (like mood, sleep, day-to-day circumstances).
    • Examples:
      • selecting the top 10 and bottom 10 performers may shift toward average rather than due to the intervention
      • a National Student Survey ranking drop next year may reflect statistical regression rather than real decline
      • Olympic-style prediction analogy: an athlete’s exceptional performance one year doesn’t guarantee the same later, partly due to chance and day conditions
  • Additional validity threats from Cook & Campbell (explicitly named):

    • Compensatory equalization

      • Professionals “improve” the control group because they know another group received special treatment.
      • Example: hospital staff give control patients extra attention because they know another group received a special treatment, shrinking the difference.
    • Compensatory rivalry (John Henry effect)

      • Participants in the control group work harder or change behavior to “catch up.”
      • Example: when one group receives new computers, those without them may work harder, minimizing differences.
      • Historical anecdote: John Henry overexerts to prove superiority over a machine—illustrating extreme compensatory effort.

Reliability vs validity metaphor

  • Reliability: arrows land consistently (same pattern).
  • Validity: arrows land in the correct target center (“ball’s eye”).
  • A study can be:
    • reliable but not valid (consistent but consistently wrong)

3) Generalizability

  • Main question: Do the results matter beyond the specific situation/sample studied?
  • Meaning: Generalize from a sample/group to a broader population.

Primary threat: selection problem

  • The studied group is not representative of the population you want to generalize to.

Common non-representative sampling patterns

  • Volunteers

    • Volunteers differ systematically (often more time, interest, motivation).
    • Example: volunteering bias in studies about secretiveness (secretive people may be less likely to volunteer).
    • Sensitive topics can produce extreme groups rather than a balanced middle (e.g., patterns in volunteering related to racism).
  • Snowball sampling

    • Participants recruit others like themselves, which can worsen representativeness concerns.

Other generalizability threats described

  • Unusual setting

    • Even with random sampling, the chosen location might be atypical due to local differences.
  • Setting + history affecting conclusions (anthropology example)

    • Anthropologists studied “peasants resistance to change” in Mexico, assuming innate cultural conservatism.
    • When revisiting nearby regions, they found the earlier village’s peasants had been repressed by landlords and police; other villages with different history were more innovative.
    • Lesson: “typical” conclusions can fail because the sample/context has unique historical constraints.
  • Construct effect

    • Some groups may have distinct ways of thinking that are subtle and hard to detect beforehand, meaning they may not represent the population.

4) Credibility

  • Main question: Is there sufficient detail and transparency for the work to be judged as trustworthy?
  • Meaning: Whether research seems trustworthy considering:
    • researchers’ background/expertise
    • credibility of their methods and reporting
    • publication and scrutiny by others
    • funding sources and possible conflicts of interest

Key emphasis

  • Credibility goes beyond reliability/validity/generalizability.
  • Readers should ask: Who did it, why, and who paid for it?

Speaker/source identification

  • Primary speaker (video presenter/lecturer): Unnamed
  • Named referenced sources/authors:

    • Shipman (1988) — presented as the origin of the “four major questions/criteria”
    • Cook and Campbell — referenced for threat lists to validity (especially in experimental design contexts)
  • Other named concepts/examples:

    • Constructivist researchers/philosophers (general reference)
    • John Henry (used to illustrate compensatory rivalry)

Original video