Descriptive statistics: describe the data.
   Summarize with statistics (derived, calculated from the data), tabulate (frequency distribution), graph (histogram, etc.)
   Shape. Center. Spread. Outliers.

Data
Qualitative (nominal, categorical)
words
Quantitative
numbers
Discrete Continuous
integers real nos. (decimals)
counts measurements
(amount, size, time, rate)
how many how much

Levels of measurement
Level (scale) Examples What can do with
Nominal names, labels, categories: (classifying) [binary/dichotomous]: Yes/No, Agree/Disagree, True/False, Have/Havenot, Success/Failure, M/F, ...
MaritalStatus, State, County, Zipcode, Major, Brand,make,model,color, Place
race,religion,party,ideology,species,nationality,language, style
TaxFilingStatus, Blood type, Housing type, Pet
Count/tally each category. Relative frequency. Mode. Bar chart.
Chi-square Tests (independence, goodness-of-fit)
Confidence interval 1-PropZInt
Ordinal orderable/rankable categories
but differences (obtained by subtraction) between data values either cannot be determined or are meaningless.
class(frosh/soph/jun/sen), trim levels, film ratings, SML, gold/silver/bronze, letter grades, Education level, clothing sizes, pain scales, military rank, star ratings, priority/risk levels, Mohs
Percentiles.
Likert scale:
Strongly disagree / Disagree / Neutral (or Unsure) / Agree / Strongly agree
Very dissatisfied / Dissatisfied / Neutral / Satisfied / Very satisfied
Poor / Fair / Good / Very good / Excellent
0/1-10 scales (may be Interval level ↓): pain
Above + median/quartiles, Spearman.
Interval Numbers: orderable, and differences between data values can be found and are meaningful. But no natural zero (meaning none of the quantity). Temperature C or F, Years/Dates, shoe size, IQ/SAT (or ordinal ↑), FICO, pH
0 is fakish
histogram mean median SD...
Estimation, CI: t-test,
Hypothesis testing,
ANOVA,
correlation & regression
Ratio Numbers: orderable, and differences between data values can be found and are meaningful, and natural zero (meaning none of the quantity), and ratios (eg. "twice as much") are meanginful. Weight Height Age
Length Area Volume
Time Money Energy TemperatureK Angle Information
BP LDL FBS BMI    MPG MPH BPM
DJI S&P500
counts
Above + CV, GM HM

Measurements have some measuring unit, e.g. inches, pounds, meters, minutes, acres, grams, MPH, BPM, ng/L, ... but they are basically irrelevant for the statistical work.

Data "set" (but can have duplicates) consisting of observations/data values/measurements/datums/individuals/scores of a person, thing or event, all of the same meaning, e.g. weights of adults, greasiness of bags of chips, longevity of bulbs, widget regional sales, effect of pill, number of meal served daily, car make, ...
   Population: complete collection of all the measurements being considered. eg. weights of ALL adults, greasiness of ALL bags of chips
   Sample: a subset of the population.
Whole numbers vs. real (decimal) numbers: no difference calculating stats, histogram, etc.
Negative numbers: ditto. If all negative, "middles" are negative; if some positive, middles might be either or 0. Range and SD always positive.

interval: set of continuous numbers. e.g. [1.45,3.7]    on number line:
OR, where all our data is: [min,max]
range (statistic) is the length or distance of our data interval. Always positive. ↑ 2.25

Some Triola data
Some data distros


Example: Population: weights of adults in country/county.
Not possible to census this. So need a non-biased, representative sample (a teaspoon of the pot of soup).
Ideal: Simple random sample (SRS): every adult equally-likely to be in the sample and every sample of that size is equally-likely.
   The selection procedure/method to take the sample is the "random". Randomly-taken sample.
  Bad: voluntary response, convenience sample.
Collect data. Measured vs self-reported (unreliable).
Calculate/derive statistic from the data: a point estimate of the parameter. But samples have uncertainty/variability so determine [confidence] interval estimate.
Inferential statistics: use probability to understand/quantify/describe uncertainty.
If have census, i.e. population is all known, no need to sample, just describe the population. Sample(s) only useful/taken/needed to estimate population parameter(s).


Data in Excel file (.xlsx). Open it in Excel. Select column, copy, then paste into other SW.

Data in Text file (.txt, .dat) or webpage, in column (of many columns, each a different data set):
  Open it or Import it in Excel. Select column, copy, then paste into other SW.
    OR
  Open it in Notepad and then select all (Ctrl-A) then copy (Ctrl-C) and paste into Excel. Select column, copy, then paste into other SW.
BodyTemperatures.txt


Dot plot.

Stem-and-leaf displays.
   Data: 44 46 47 49 63 64 66 68 68 72 72 75 76 81 84 88 106
Primes<100
train schedule

Time series chart. temporal measurements. Run chart. Run-sequence plot. Process control chart (statistical quality control).
    
Playfair ~1800
Nightingale 1858
Minard 1869

Boxplot (box-and-whisker plot)

one per sample/experiment. Vertical. Side-by-side comparison. Outliers:


Research papers & technical reports published each day:
Source / TypeEstimated annual outputPer day (approx.)
Peer-reviewed journal articles3.5 – 5 million10,000 – 14,000
Conference papers1 – 1.5 million3,000 – 4,000
Preprints (arXiv, bioRxiv, SSRN, etc.)0.6 – 0.8 million1,600 – 2,200
Technical / government / industry reports0.5 – 1 million1,400 – 2,700
Total (all scholarly + technical)≈ 6 – 8 million≈ 16,000 – 22,000
Field% that use statisticsNotes
Biomedical / clinical90–95%Almost universal
Psychology / social sciences85–95%Very high
Economics / finance80–90%High
Biology / life sciences70–85%High
Engineering50–70%Mixed
Computer science40–60%Many theoretical or systems papers have little
Physics / chemistry40–60%Theoretical papers often have none
Mathematics / pure theory5–15%Mostly proofs
Humanities5–20%Low except digital humanities
Overall weighted average≈ 65–75%
So on a typical day, something like 12,000–15,000 new papers and reports appear that contain statistical analysis.


random stochastic aleatory chance luck mis/fortune contingent accidental fate fortuitous haphazard

random number between 0 and 1 (e.g. a probability)
#0 #1 #2 #3 #4 #5 #6 #7 #8 #9
computer PRNG (pseudo-random number generator) : 9 quadrillion of them    (@1/s → 285M years)
There will be about the same number of them in each quartile, decile, percentile, millionile.

Each pixel randomly black or white:
     Sound track of randomness: