
Fitting Models Is like Tetris: Crash Course Statistics #35
video description
Date: 2022-04-04
Related videos
Comments and reviews: 6
Life
We got two examples of how to use ancova here, but no real explanation of what ancova actually -does-. :(
An ancova is a regression where you try to fit multiple lines at once. The covariate (e. g. the baby's age) is your X variable, and the Y variable is your response as usual (baby weight gain. In the simplest case, each group gets its own line [y=bx+a] with the same -slope- (b, but different -intercepts- (a1, a2, etc); in this case, one intercept for each baby formula type. Lines that have the same slope but differ in intercept will appear parallel, but at different -heights- on the plot: this would mean, in this case, that age has the same overall effect on weight gain, but one formula gives a higher -baseline- weight gain at any given time point than the other.
So the p value for your covariate is an answer to the question -is there a significant effect, i. e. a nonzero slope, of X (age) on Y (weight gain? -, while the p value for your grouping variable is the answer to the question -is there a significant difference in intercept between the regression line for one type of formula and the regression line of the other type of formula? - In this case, there was a significant difference - and reorganizing the data according to an X variable (age) made us able to detect this difference, whereas when age was not considered, all data around each regression line was smushed together around a single mean, obscuring the effect of formula with a whole lot of noise. This is why the p value changed when age was added.
If an interaction term is included, it tests for a difference in -slope- between the groups. So an age-by-formula interaction might mean, for example, that weight gain is similar shortly after birth, but it makes a difference which formula you use later on in the baby's life. This would appear in the regression plot as two lines (representing the formula types) that are close together near X=0, but move apart farther to the right on the X-axis (as age increases) because one line is steeper than the other.
reply
We got two examples of how to use ancova here, but no real explanation of what ancova actually -does-. :(
An ancova is a regression where you try to fit multiple lines at once. The covariate (e. g. the baby's age) is your X variable, and the Y variable is your response as usual (baby weight gain. In the simplest case, each group gets its own line [y=bx+a] with the same -slope- (b, but different -intercepts- (a1, a2, etc); in this case, one intercept for each baby formula type. Lines that have the same slope but differ in intercept will appear parallel, but at different -heights- on the plot: this would mean, in this case, that age has the same overall effect on weight gain, but one formula gives a higher -baseline- weight gain at any given time point than the other.
So the p value for your covariate is an answer to the question -is there a significant effect, i. e. a nonzero slope, of X (age) on Y (weight gain? -, while the p value for your grouping variable is the answer to the question -is there a significant difference in intercept between the regression line for one type of formula and the regression line of the other type of formula? - In this case, there was a significant difference - and reorganizing the data according to an X variable (age) made us able to detect this difference, whereas when age was not considered, all data around each regression line was smushed together around a single mean, obscuring the effect of formula with a whole lot of noise. This is why the p value changed when age was added.
If an interaction term is included, it tests for a difference in -slope- between the groups. So an age-by-formula interaction might mean, for example, that weight gain is similar shortly after birth, but it makes a difference which formula you use later on in the baby's life. This would appear in the regression plot as two lines (representing the formula types) that are close together near X=0, but move apart farther to the right on the X-axis (as age increases) because one line is steeper than the other.
reply
Shawn
1) Why bother with an ANCOVA when you can run a regression with both variables? Redhead / not redhead is binary, so it will work just fine in a regression model. If you wanted to expand it to include redheads compared to other hair colors, we could just dummy code each category and drop one from the model as the reference variable.
2) What was up with the - in the error for the RMA after accounting for the base run time? Does - represent an approximate value, even though your decimal was wicked long? At first glance I thought it was a - (negative) instead of -, but that wouldn't make sense. So?
reply
1) Why bother with an ANCOVA when you can run a regression with both variables? Redhead / not redhead is binary, so it will work just fine in a regression model. If you wanted to expand it to include redheads compared to other hair colors, we could just dummy code each category and drop one from the model as the reference variable.
2) What was up with the - in the error for the RMA after accounting for the base run time? Does - represent an approximate value, even though your decimal was wicked long? At first glance I thought it was a - (negative) instead of -, but that wouldn't make sense. So?
reply
Francois
I'm a bit late to this crash course, but as someone who works with Excel (statistics or otherwise, the tables in this video look weird to me. I recommend right aligning all numbers, add thousands separators and also align so that the decimal symbols of numbers of the same column are aligned horizontally. It will greatly improve the readability.
reply
I'm a bit late to this crash course, but as someone who works with Excel (statistics or otherwise, the tables in this video look weird to me. I recommend right aligning all numbers, add thousands separators and also align so that the decimal symbols of numbers of the same column are aligned horizontally. It will greatly improve the readability.
reply
Aditya
Does anyone understand repeated measures ANOVA and the example of the effects of music on running? Does she mean that we should normalize/standardize the values of running time with respect to each individual?
reply
Does anyone understand repeated measures ANOVA and the example of the effects of music on running? Does she mean that we should normalize/standardize the values of running time with respect to each individual?
reply
Ilya
Will you cover the various means comparisons methods. Bonferroni, Tukey, Sidak, Scheffe, Fisher, Holm-Bonferroni, and Holm-Sidak.
reply
Will you cover the various means comparisons methods. Bonferroni, Tukey, Sidak, Scheffe, Fisher, Holm-Bonferroni, and Holm-Sidak.
reply
mr.
Are you going to cover multivariate analysis? I-d love to watch it -, as much as I love the way you explain every topic: )
reply
Are you going to cover multivariate analysis? I-d love to watch it -, as much as I love the way you explain every topic: )
reply
Add a review, comment
Other channel videos















