The Wrong Verdict Slides Lesson 1: The Wrong Verdict
The Cost of
Being Wrong
Understanding the tension between False Positives and False Negatives in statistical decision making.
Advanced Statistics: Comparing Populations
The Statistics Courtroom
The Null Hypothesis (\(H_0\))
The defendant is innocent .
The Alternative (\(H_a\))
The defendant is guilty .
"We assume innocence until proven guilty beyond a reasonable doubt." (In stats, our "doubt" threshold is \(\alpha\)).
How do we decide when the evidence is "enough"?
The Four Outcomes
Decision \ Truth \(H_0\) is True (Innocent) \(H_a\) is True (Guilty) Fail to Reject \(H_0\) Correct Decision Innocent person goes free. Type II Error (\(\beta\)) Guilty person goes free. Reject \(H_0\) Type I Error (\(\alpha\)) Innocent person jailed. Correct Decision Guilty person is convicted.
Discussion: Which error is "worse" in a trial? In a medical test for a deadly disease?
Naming the Villains
Type I Error
The "False Positive"
Finding a difference that doesn't exist . Rejecting a true null hypothesis.
Probability = \(\alpha\) (Alpha)
Type II Error
The "False Negative"
Missing a difference that actually exists . Failing to reject a false null hypothesis.
Probability = \(\beta\) (Beta)
"We can't minimize both at the same time for a fixed sample size."
The Verdict Worksheet THE VERDICT LOG
Statistical Error Analysis
Student Name:
Trial Date:
Part 1: Defining the Errors
Type I Error (\(\alpha\))
In your own words, what is happening here? When do we make this mistake?
Type II Error (\(\beta\))
In your own words, what is happening here? When do we make this mistake?
Part 2: The Stakes of Mistakes
Scenario A: The Rare Allergy Test
A pharmaceutical lab is testing a new screening for a life-threatening peanut allergy. \(H_0\): The patient is NOT allergic. \(H_a\): The patient IS allergic.
Consequences of Type I Error:
Consequences of Type II Error:
Which error is more dangerous here?
Type I
Type II
Scenario B: The Fertilizer Contest
A gardener tests a very expensive new fertilizer "BloomMax" against their cheap standard brand. \(H_0\): BloomMax is no better than cheap brand. \(H_a\): BloomMax is significantly better.
Consequences of Type I Error:
Consequences of Type II Error:
Which error is more dangerous here?
Type I
Type II
SYNTHESIS: The Alpha Setting
If you are a researcher and you want to be extremely sure you don't send an innocent person to jail (Type I Error), would you choose \(\alpha = 0.01\) or \(\alpha = 0.10\)? Explain how this choice affects the risk of a Type II error.
Power Mechanics Slides Lesson 2: Power Play Simulation
The Power to
Detect Truth
Why some studies find what they are looking for, and others are doomed from the start.
Statistical Power (\(1 - \beta\))
What is Power?
The Official Definition
"The probability that a test will correctly reject a false null hypothesis."
In plain English: If there is a real difference between two populations, what are the odds we actually notice it?
Power = \(1 - \beta\)
Power is a "Hit"
Miss
Type II Error
HIT!
Power
Power measures our "vision" in the fog of data.
The Levers of Power
How can a researcher increase their power?
Sample Size (\(n\))
More data reduces noise. Bigger samples = Clearer signals.
Increase \(n\) \(\rightarrow\) Power UP
Alpha (\(\alpha\))
Being less strict about evidence (raising \(\alpha\)) makes it easier to reject \(H_0\).
Increase \(\alpha\) \(\rightarrow\) Power UP
Effect Size
If the real difference is HUGE, it's very easy to see.
Big Effect \(\rightarrow\) Power UP
The Researcher's Dilemma
"If I want more Power, why don't I just set \(\alpha = 0.50\) and only sample 10,000 people?"
The Cost of high Alpha...
Way too many Type I Errors!
The Cost of huge samples...
Time and Millions of Dollars.
Power Play Activity LAB LOG: POWER PLAY
Statistical Simulation Record Sheet
Researcher Name
Lab Station #
Simulation Instructions
Open the provided Rossman/Chance "Power" applet. We will simulate testing a new medicine. The Medicine really works (Alternative Hypothesis is True). Our goal is to see how often we successfully Reject the Null . Each trial involves 100 simulations.
Trial 1: The Weight of Evidence (Sample Size)
Keep \(\alpha = 0.05\) and Effect Size = 0.5 (Standard) constant. Adjust Sample Size (\(n\)).
Sample Size (\(n\)) # of Rejections (out of 100) Estimated Power Observations 10 Describe the trend here... 30 100
Trial 2: The Strictness Filter (Alpha)
Keep \(n = 30\) and Effect Size = 0.5 constant. Adjust Alpha (\(\alpha\)).
Alpha (\(\alpha\)) # of Rejections (out of 100) Estimated Power Observations 0.01 How does "being picky" (\(\alpha=0.01\)) affect your ability to find an effect? 0.05 0.10
Post-Lab Synthesis
1. Based on your simulations, if a drug company is testing a treatment for a very common, non-lethal condition (like a mild headache), why might they prefer a large sample size rather than a high alpha level?
2. Draw a quick sketch of two overlapping normal curves. Color the area that represents Power. Use your simulation visual as a guide.
Power Variables Reference POWER CHEAT SHEET
The Rules of the Game
The Vocabulary
Power (\(1 - \beta\))
Probability of finding a difference when one actually exists .
Alpha (\(\alpha\))
Threshold for "significant." Probability of a False Positive .
Beta (\(\beta\))
Probability of a False Negative (missing the effect).
How to Boost Power
Increase Sample Size (\(n\))
"Bigger net catches more fish."
Increase Significance Level (\(\alpha\))
"Easier to clear the hurdle."
Decrease Standard Deviation (\(\sigma\))
"Less noise, more signal."
Increase True Effect Size
"Easier to find a giant than an ant."
The Power Summary Table
Variable
Direction
Type II Error (\(\beta\))
Power (\(1-\beta\))
Sample Size (\(n\))
UP
DOWN
UP
Alpha (\(\alpha\))
UP
DOWN
UP
Effect Size
UP
DOWN
UP
Size Matters Slides Lesson 3: Size Matters Most
Beyond the
P-Value
Distinguishing between statistical significance (is there an effect?) and practical significance (how big is it?).
Practical Importance & Cohen's d
The P-Value Paradox
Imagine a study with 1,000,000 people.
"A new vitamin increases IQ scores by an average of 0.05 points."
Because the sample is so huge, the p-value is 0.000001. It is highly statistically significant .
The Problem:
0.05 IQ points is completely meaningless in real life.
Statistical Signifiance
"Is it real?"
Practical Significance
"Does it matter?"
Measuring "Oomph": Cohen's \(d\)
Effect size tells us how far apart the two population means are in terms of standard deviations .
\(d = \frac{\bar{x}_1 - \bar{x}_2}{s_p}\)
Difference / Pooled Standard Deviation
Interpretation
0.2
Small Effect
0.5
Medium Effect
0.8+
Large Effect
Visualizing Effect Size
Small Effect (\(d = 0.2\))
Massive overlap. Most people in group A are similar to group B.
Large Effect (\(d = 0.8\))
Clear separation. A person in Group B is likely very different from Group A.
Effect Size Worksheet Practical Significance Lab
Calculation and Interpretation of Cohen's d
Analyst:
Report ID: EFF-03
The Golden Rule
A p-value tells you if a difference is likely due to chance . Effect size tells you how many standard deviations apart the averages are.
\(d = \frac{\bar{x}_1 - \bar{x}_2}{s_p}\)
Case Study 1: The Energy Drink Trial
N = 50,000
A company tests "HyperFocus" energy drink. The treatment group (\(n=25,000\)) averaged 102 on a focus test (\(s=15\)). The control group (\(n=25,000\)) averaged 100 (\(s=15\)). The result is statistically significant with \(p < 0.01\).
Calculate Cohen's d:
Interpretation (Small, Medium, Large?):
Practicality Check:
If the drink costs $5.00 per can, would you recommend it based on this effect size? Why or why not?
Case Study 2: Small Scale, Big Win
N = 20
A boutique therapist tests a "Memory Palace" technique on 10 students. Their scores improved by 12 points (\(s=4\)) compared to a control group of 10 students. The result is barely significant with \(p = 0.045\).
Calculate Cohen's d:
Interpretation (Small, Medium, Large?):
Scientific Value:
Compare this to Case Study 1. Which study shows a more impressive intervention? Why might researchers still be skeptical of this one?
Reference Table
\(d \approx 0.2\)
Small
\(d \approx 0.5\)
Medium
\(d \approx 0.8\)
Large
\(d \approx 2.0\)
Huge
Broken Science Slides Lesson 4: Broken Science Investigation
The Crisis of
Replication
Why half of published scientific findings might be false—and the dirty secrets behind "significant" results.
P-Hacking, Publication Bias, and Ethics
The P-Hacking Recipe
P-Hacking Defined
Manipulating data or running many tests until you find a p < 0.05, then only reporting that one result.
Example:
"We tested 20 different colors of jelly beans to see if they cause acne. 19 colors showed nothing. Green jelly beans showed a significant link (\(p=0.05\)). Headline: GREEN JELLY BEANS CAUSE ACNE!"
The Math of Luck
If you test 20 random variables, the probability of getting at least one significant result purely by chance is:
\(1 - (0.95)^{20} \approx 64\%\)
You're more likely to find a "fake" result than not!
The File Drawer Problem
Experiment 1
"No significant difference found."
The File Drawer
Experiment 2
"No significant difference found."
The File Drawer
Experiment 3
"p = 0.049 Significant!"
PUBLISHED!
The Problem: When we only see the "winners," we have no idea how many times researchers "lost" (failed to reject the null).
Statistical Integrity
"Science is a process of honest discovery, not a quest for a specific p-value."
Solutions:
Pre-registering studies
Reporting ALL findings
Replication by other labs
Peer Review Slides Lesson 5: Peer Review Board
The Final Word
Acting as scientific peer reviewers to evaluate comparative claims using everything we've learned.
Your Review Toolkit
1. Check for Errors
Is there a risk of Type I Error (Alpha too high)? Or Type II (Sample too small)?
2. Demand Power
Was the sample size large enough to actually detect a realistic effect?
3. Calculate "Oomph"
It's significant (\(p < 0.05\)), but is the Effect Size actually meaningful?
Be the gatekeeper of truth.
Abstract Review
Case A
"Daily Coffee Intake Linked to 0.002% Increase in Lifespan"
Method: Studied 1,200,000 adults over 40 years. Compared high-intake (\(>5\) cups) vs. zero-intake.
Results: Difference in mean age of death was 2 days. The p-value was \(p = 0.0003\).
Reviewer Verdict?
Statistically Significant? YES
Practically Significant? NO
The headline is misleading!
Peer Review Board Task OFFICIAL PEER REVIEW
Board of Statistical Integrity
Status
Review Pending
Assigned Reviewer
Filing Date
Assignment: Abstract Audit
You have been handed a draft for a high-stakes comparative study. Your job is to analyze the data provided and decide if the study should be Accepted , Revised , or Rejected based on its statistical merit.
Exhibit #A-902: The "Sleep-Smart" Method Topic: Cognitive Psychology
Study Abstract
"We hypothesized that listening to classical music while sleeping improves mathematics test performance. We recruited 20 high school volunteers. 10 students (Treatment) slept with music; 10 (Control) slept in silence. After one night, the Treatment group scored an average of 88% (\(s=10\)) on a calculus quiz, while the Control group scored 78% (\(s=10\)). We calculated a p-value of \(p = 0.041\)."
Sample Size
\(n = 20\) (total)
P-Value
\(0.041\)
Alpha Threshold
\(0.05\)
1. Effect Size Calculation
Calculate Cohen's d for this study. Show your work.
2. Power Assessment
With a sample size of only 10 per group, discuss the Power of this test. Is it high or low? What is the risk of a Type II error here?
3. Final Verdict & Justification
Would you publish this study as "definitive proof" that music makes you smarter? Use the terms "Practical Significance," "P-hacking potential," and "Replication" in your response.
REJECTED
REVISE
ACCEPTED
Critique Rubric Critique Rubric
Evaluation Standards for Peer Review Tasks
Criteria Advanced (4) Proficient (3) Developing (1-2) Calculation Accuracy Cohen's d and power estimation. All calculations are correct. Cohen's d is calculated exactly with pooled standard deviation. Cohen's d is calculated correctly, but may have minor rounding errors or use a simplified formula. Calculations contain significant errors or are missing steps. Conceptual Linkage Connecting n, alpha, and power. Clearly explains how small sample size leads to low power and high Type II error risk. Correctly identifies that a small sample size is a weakness but explanation lacks depth. Mentions sample size but cannot explain its relationship to power or error types. Practical vs. Stats Sig Interpreting the findings. Expertly distinguishes between the p-value result and the effect size's real-world meaning. Recognizes that a low p-value doesn't always mean a large or important effect. Relies solely on the p-value to make a judgment. Ethical Awareness Identifying bias/p-hacking. Critiques the study's design for potential publication bias or p-hacking vulnerabilities. Mentions that the study needs replication to be truly valid. Does not address the need for replication or the ethics of the findings.
Grading Note
A passing "Reviewer" must achieve at least a Proficient (3) in all categories to be certified for the Board of Statistical Integrity.