
When Predictions Fail: Crash Course Statistics #43
video description
Date: 2022-04-04
Related videos
Comments and reviews: 10
Esben
I am not sure you did a fair job of summarising the errors made in 2016. The polls were mostly right (In the sense that the average errors of 2016 were less than the average errors of all polling since 1998, and showed a tight race in the crucial states that Trump won in the end. The polls didn't fully catch the shifts happening late in the campaign (But that is also hard, as you need several days of polling and compiling before you can get a poll out, at which point a predictor has to spend days to read and assess it, but they were showing the races were increasingly within the margins of error in certain Rust Belt states.
Two very common errors among predictors (Not pollsters) were to insert their own biases into the mix and to treat margins of error as static, uncorrelated and unimportant. Biases were often in analysis, where people who should know better because they had the numbers treated states as part of a blue wall, certain to go for Democrats. They looked at Trumps dismal favoured/unfavoured numbers without also comparing those numbers to Clintons. Making bad predictions in politics is just as often a question about interpreting which numbers are important and which aren't, and even good predictors cannot have a perfect track record; sometimes it really is just a cointoss and a gut instinct.
Margins of error in election polling are a bit complicated to get right. While the basic interpretation is as you put it (54% +/- 2 prcpts means anywhere between 52 and 56 is expected, that's not the only information such a number gives you. If a later poll of the same election gives 54% but the margin of error is increasing (Say, +/- 5 prcpts, that is a sign of increasing volatility and less predictive power in the polling.
At the same time, margins of error are often assumed to be random due to sample selection (I. e. margins of error are uncorrelated. If I did 10 polls, 3 of which gave 53%, 4 of which gave 54 and 3 of which gave 55, all with a +/- 5 prcpts, under almost all circumstances I'd be more certain that the 54% is right than with just 1 poll - I would expect the margins of error to be random, and thus cancel out the random noise. In such a case, I'd allow myself to average the polls and ignore the margin. However, if my sampling is just a bit off, and I have made essentially the same mistake in my 10 (nonrandom) samples, such a move would actually be wrong. I can't be certain that my average is more accurate than any individual poll, and thus a result of 48-50 cannot be ruled out (Likewise a result of 58-60, but that's not an election changer.
Furthermore, if there are several simultaneous elections, each with a discrete outcome, treating polls as uncorrelated doesn't just deny me my sense of certainty, but small errors might also throw off the entire result. If I have 10 elections, each with polls showing 55%-45% and a margin of error at 7 prcpts, I may be relatively convinced that, on the whole, the party ahead will win more than half of these elections, possibly all, while the party behind will certainly lose most. However, if that margin correlates 1: 1 between these elections, such that all ten elections in a sense -share- the -same- 7 prcpts margin of error, it would be an error to give 10 discrete election predictions. The most accurate answer would be a range of percentages for one outcome, namely either party winning the whole lot. That was a particular error in 2016, where most predictors did assume that margins of error were uncorrelated, which meant that 52%-48% and 3 prcpts margin in several states meant safe sailings for Clinton because she didn't need all of the states for the Electoral College.
Last, a big part of the apparent error in polling was actually journalists misunderstanding the statistics behind, and thus communicating bad analysis to the public. A 52-48 poll may be the headline, but if the journalist doesn't understand the vital role of the 3 prcpts margin of error, they will misinterpret the overall results. Seeing 52-48 as a continuum of 49-55 for one party and 45-51 for the other may not be easy to wrap your head around for anyone who doesn't know how difficult sampling is, and it may be harder to communicate to a reader, but it gives a more honest assessment of the relative chances of winning.
You cite Nate Silver several times, who has gone over these various factors in more depth than I have here, and his presidential prediction was pretty much spot on: Clinton won the popular vote as we knew she would, and it was incredibly close in the states his model predicted - those states where Trump won by a few thousand votes that decided those particular, discrete outcomes and thus the election.
Of course part of what makes Nate Silver a better statistician than most is that he embraces Bayesian logic and methodology, which are far superior to Fishers'.
reply
I am not sure you did a fair job of summarising the errors made in 2016. The polls were mostly right (In the sense that the average errors of 2016 were less than the average errors of all polling since 1998, and showed a tight race in the crucial states that Trump won in the end. The polls didn't fully catch the shifts happening late in the campaign (But that is also hard, as you need several days of polling and compiling before you can get a poll out, at which point a predictor has to spend days to read and assess it, but they were showing the races were increasingly within the margins of error in certain Rust Belt states.
Two very common errors among predictors (Not pollsters) were to insert their own biases into the mix and to treat margins of error as static, uncorrelated and unimportant. Biases were often in analysis, where people who should know better because they had the numbers treated states as part of a blue wall, certain to go for Democrats. They looked at Trumps dismal favoured/unfavoured numbers without also comparing those numbers to Clintons. Making bad predictions in politics is just as often a question about interpreting which numbers are important and which aren't, and even good predictors cannot have a perfect track record; sometimes it really is just a cointoss and a gut instinct.
Margins of error in election polling are a bit complicated to get right. While the basic interpretation is as you put it (54% +/- 2 prcpts means anywhere between 52 and 56 is expected, that's not the only information such a number gives you. If a later poll of the same election gives 54% but the margin of error is increasing (Say, +/- 5 prcpts, that is a sign of increasing volatility and less predictive power in the polling.
At the same time, margins of error are often assumed to be random due to sample selection (I. e. margins of error are uncorrelated. If I did 10 polls, 3 of which gave 53%, 4 of which gave 54 and 3 of which gave 55, all with a +/- 5 prcpts, under almost all circumstances I'd be more certain that the 54% is right than with just 1 poll - I would expect the margins of error to be random, and thus cancel out the random noise. In such a case, I'd allow myself to average the polls and ignore the margin. However, if my sampling is just a bit off, and I have made essentially the same mistake in my 10 (nonrandom) samples, such a move would actually be wrong. I can't be certain that my average is more accurate than any individual poll, and thus a result of 48-50 cannot be ruled out (Likewise a result of 58-60, but that's not an election changer.
Furthermore, if there are several simultaneous elections, each with a discrete outcome, treating polls as uncorrelated doesn't just deny me my sense of certainty, but small errors might also throw off the entire result. If I have 10 elections, each with polls showing 55%-45% and a margin of error at 7 prcpts, I may be relatively convinced that, on the whole, the party ahead will win more than half of these elections, possibly all, while the party behind will certainly lose most. However, if that margin correlates 1: 1 between these elections, such that all ten elections in a sense -share- the -same- 7 prcpts margin of error, it would be an error to give 10 discrete election predictions. The most accurate answer would be a range of percentages for one outcome, namely either party winning the whole lot. That was a particular error in 2016, where most predictors did assume that margins of error were uncorrelated, which meant that 52%-48% and 3 prcpts margin in several states meant safe sailings for Clinton because she didn't need all of the states for the Electoral College.
Last, a big part of the apparent error in polling was actually journalists misunderstanding the statistics behind, and thus communicating bad analysis to the public. A 52-48 poll may be the headline, but if the journalist doesn't understand the vital role of the 3 prcpts margin of error, they will misinterpret the overall results. Seeing 52-48 as a continuum of 49-55 for one party and 45-51 for the other may not be easy to wrap your head around for anyone who doesn't know how difficult sampling is, and it may be harder to communicate to a reader, but it gives a more honest assessment of the relative chances of winning.
You cite Nate Silver several times, who has gone over these various factors in more depth than I have here, and his presidential prediction was pretty much spot on: Clinton won the popular vote as we knew she would, and it was incredibly close in the states his model predicted - those states where Trump won by a few thousand votes that decided those particular, discrete outcomes and thus the election.
Of course part of what makes Nate Silver a better statistician than most is that he embraces Bayesian logic and methodology, which are far superior to Fishers'.
reply
MoeZeppelin
the financial crisis is a bad example. everyone in banking and finance knew a financial crisis was coming, but there was no way to predict specifically when. the banks were coerced to make bad loans by the federal government in the name of -equality opportunity. - the banks knew this was a pending disaster so they did their best to mitigate it (and/or hide it) by securitizing mortages, bundling them, and trading them back and forth as if they had real value. the whole disaster was a result of social engineering to the detriment of economic common sense. and all the players were in for a penny in for a pound, and banking on the notion they were -too big to fail- lol. how about looking at why every doom and gloom climate model since the 1980's has failed, yet all that's been done is doubling down on the next failure of a prediction. there's your irrational, unpredictable human behavior in a nutshell, lol.
reply
the financial crisis is a bad example. everyone in banking and finance knew a financial crisis was coming, but there was no way to predict specifically when. the banks were coerced to make bad loans by the federal government in the name of -equality opportunity. - the banks knew this was a pending disaster so they did their best to mitigate it (and/or hide it) by securitizing mortages, bundling them, and trading them back and forth as if they had real value. the whole disaster was a result of social engineering to the detriment of economic common sense. and all the players were in for a penny in for a pound, and banking on the notion they were -too big to fail- lol. how about looking at why every doom and gloom climate model since the 1980's has failed, yet all that's been done is doubling down on the next failure of a prediction. there's your irrational, unpredictable human behavior in a nutshell, lol.
reply
Petrico94
-Well educated voters are more likely to take surveys- idk, I don't think educated people answer spam phone calls all that much, but trying to get an idea of what EVERYONE in the country is thinking, sticking to one area where there's a clear bias, maybe you're just listening to vocal people who actually want to talk about it, and elections can be driven by factors besides raw numbers, just like this election where Trump actually had fewer votes but the electoral college said he had enough real votes to win. But it just goes to show 1 out of 100 is still a chance but getting that chance probably means better data should be collected.
reply
-Well educated voters are more likely to take surveys- idk, I don't think educated people answer spam phone calls all that much, but trying to get an idea of what EVERYONE in the country is thinking, sticking to one area where there's a clear bias, maybe you're just listening to vocal people who actually want to talk about it, and elections can be driven by factors besides raw numbers, just like this election where Trump actually had fewer votes but the electoral college said he had enough real votes to win. But it just goes to show 1 out of 100 is still a chance but getting that chance probably means better data should be collected.
reply
Raymond
-_. Statistics as-it-is is a failure as a method for representation of 'what-we-agree-we-know': Statistics as-is simply does-not formulate, agreement, knowledge, data, logic. like, what's the probability bank-shareholders will agree on including certain data, is simply -That does not compute- [famous Lost-In-Space robot line]. The scientific method itself-is lacking. _-
-_. on the other-hand, the to-be-famous mathematician Hari Seldon must have solved this. _-
-_. 'hmm', interesting question-did Asimov establish a Hari Seldon Prize in psychohistory. _-
reply
-_. Statistics as-it-is is a failure as a method for representation of 'what-we-agree-we-know': Statistics as-is simply does-not formulate, agreement, knowledge, data, logic. like, what's the probability bank-shareholders will agree on including certain data, is simply -That does not compute- [famous Lost-In-Space robot line]. The scientific method itself-is lacking. _-
-_. on the other-hand, the to-be-famous mathematician Hari Seldon must have solved this. _-
-_. 'hmm', interesting question-did Asimov establish a Hari Seldon Prize in psychohistory. _-
reply
ResortDog
Dutchsinse describes how forces migrate around the world dependent on plate tectonics & recent activities. No, it's not the exact times, but it's usually the right places & magnitudes. SO, living where this stuff, or by volcanoes which are in the system; you can expect it eventually. At least taping down the antique glass over your heads & strap the hot water heater.
reply
Dutchsinse describes how forces migrate around the world dependent on plate tectonics & recent activities. No, it's not the exact times, but it's usually the right places & magnitudes. SO, living where this stuff, or by volcanoes which are in the system; you can expect it eventually. At least taping down the antique glass over your heads & strap the hot water heater.
reply
Donitee
The polls leading up to the 2016 election were fairly accurate though. The polls in the crucial states that Trump narrowly won had results well within the margin of error. How you can look at the swing state polls on the day before the election and think -Hillary wins this 99/100 times- is beyond me. The models were wrong, more wrong than the polls at least.
reply
The polls leading up to the 2016 election were fairly accurate though. The polls in the crucial states that Trump narrowly won had results well within the margin of error. How you can look at the swing state polls on the day before the election and think -Hillary wins this 99/100 times- is beyond me. The models were wrong, more wrong than the polls at least.
reply
joemackey1950
Re: Housing problem. Part of that was Congress telling banks they had to grant home loans, regardless of the borrowers ability to repay. The sub-prime deal. Thus banks gave loans to people who couldn't afford them and defaulted on their loan. Then Congress shifted the blame to the banks for the problem that Congress created.
reply
Re: Housing problem. Part of that was Congress telling banks they had to grant home loans, regardless of the borrowers ability to repay. The sub-prime deal. Thus banks gave loans to people who couldn't afford them and defaulted on their loan. Then Congress shifted the blame to the banks for the problem that Congress created.
reply
John
About 5 years before the housing market crash, I was asking -Where is all this money coming from? - But what do I know; I'm not college educated. Turns out I was right. Before the election, I predicted that for every vocal Trump supporter, there'd be about 5 more who would silently vote for him. I hate being right.
reply
About 5 years before the housing market crash, I was asking -Where is all this money coming from? - But what do I know; I'm not college educated. Turns out I was right. Before the election, I predicted that for every vocal Trump supporter, there'd be about 5 more who would silently vote for him. I hate being right.
reply
Sean
You run polls several times during a campaign, but you hold the election once. Clinton did receive the most votes, but the electoral college isn-t won by the total votes cast. The break down of the various districts in each state matters.
reply
You run polls several times during a campaign, but you hold the election once. Clinton did receive the most votes, but the electoral college isn-t won by the total votes cast. The break down of the various districts in each state matters.
reply
SJ
Good work. Sometimes statistics go wrong because researchers source data from the wrong sources or from people who have the internet and telephones and then generalize such findings to the whole population.
reply
Good work. Sometimes statistics go wrong because researchers source data from the wrong sources or from people who have the internet and telephones and then generalize such findings to the whole population.
reply
Add a review, comment
Other channel videos















