Skip to main content

Posts

Showing posts with the label Box-and-Whisker plot

Basic Statistics Lecture #8: Anscombe's Quartet

"There are three kinds of lies: lies, damned lies, and statistics." - Unknown As promised last time , I will be covering Anscombe's Quartet.  It is an idea where people may use statistics to lie about a data set.  It is a series of data sets developed by statistician Francis Anscombe and published in the journal American Statistician in 1973. I'm going to provide you with four sets of data.  Do me a favor and apply what you know of statistical analysis to them.  If the statistical data look weird to you, don't be scared; you may have done the analysis perfectly.  Here's the four sets: I II III IV x y x y x y x y 10 8.04 10 9.14 10 7.46 8 6.58 8 6.95 8 8.14 8 6.77 8 5.76 13 7.58 13 8.74 13 12.74 8 7.71 9 8.81 9 8.77 9 7.11 8 8.84 11 8.33 11 9.26 11 7.81 ...

Basic Statistics Lecture #7: Quantitative, Continuous, and Numerical Data

As promised last time, today I will cover basic calculations of data accumulated from real data.  Please take note that all of the following is for the simple case of one group of data.  For two or more distinct groups of data, the calculations will be similar, but slightly more specific due to the nature of 2+ distinct groups of data.  I will cover that in a later post, which will be labeled as ANOVA.  As a side note, I'm a baseball fan, so I'm going to provide examples from the MLB. This information has the labels for the data of a sample, not the population.  The population is the set of all possible people or objects which falls under the category under study.  If we were studying the 2017 ERA's of pitchers, the population would be all MLB pitchers who have pitched in 2017.  The sample is the subset of the population which we are getting the data points from.  If we want to look at the 8 teams who have made it to the Division Series, then ...