MATH1231 5,584 words·28 min read

Introduction to Probability and Statistics

9.1 Some preliminary set theory#

Probability is built on sets, so before any probability appears we need to be fluent with the language of sets. Most of this will be revision, but the notation matters enormously later — almost every probability rule in this chapter is really a set identity in disguise.

Note

Definition
A set is a collection of objects, called its elements. We write a∈Aa \in A to mean aa is an element of AA, and a∉Aa \notin A otherwise. The set with no elements is the empty set, written ∅\varnothing.

Example. The set A={1,2,3}A = \{1,2,3\} has elements 11, 22 and 33; so 1∈A1 \in A but 4∉A4 \notin A.

Note

Definition
AA is a subset of BB, written A⊆BA \subseteq B, if every element of AA is also an element of BB. The power set P(A)\mathcal{P}(A) is the set of all subsets of AA.

Example. The set A={1,2,3}A = \{1,2,3\} has eight subsets; for instance {2,3}⊆A\{2,3\} \subseteq A. In general a set with nn elements has 2n2^n subsets, since each element is independently either in or out — this doubling is worth remembering, as it reappears when we count outcomes.

In any given problem there is a fixed set containing everything under discussion, called the universal set SS. For three-dimensional vector geometry it is usually S=R3S = \mathbb{R}^3; in probability it will be the set of all possible outcomes.

Note

Definition
For subsets A,B⊆SA, B \subseteq S, define

A∪B={x∈S:x∈A or x∈B}(union),A∩B={x∈S:x∈A and x∈B}(intersection),Ac={x∈S:x∉A}(complement),A−B=A∩Bc(difference).\begin{align*} A \cup B &= \{x \in S : x \in A \text{ or } x \in B\} && \text{(union)}, \\ A \cap B &= \{x \in S : x \in A \text{ and } x \in B\} && \text{(intersection)}, \\ A^c &= \{x \in S : x \notin A\} && \text{(complement)}, \\ A - B &= A \cap B^c && \text{(difference)}. \end{align*}

The sets AA and BB are disjoint (or mutually exclusive) if A∩B=∅A \cap B = \varnothing.

Example. Let SS be all students enrolled in MATH1231, let AA be those aged 20 or over, and let BB be those who own a bicycle. Then A∩BA \cap B is the students aged 20 or over who own a bicycle; A∪BA \cup B is those who are aged 20 or over, own a bicycle, or both; AcA^c is those under 20; and A−BA - B is those aged 20 or over who do not own a bicycle. Note that "or" in mathematics is always inclusive — A∪BA \cup B includes the students who satisfy both conditions. English "or" is frequently exclusive, and that mismatch causes real errors when translating word problems.

Note

Definition
Sets A1,A2,…,AnA_1, A_2, \dots, A_n partition a set BB if they are pairwise disjoint and their union is BB; that is, Ai∩Aj=∅A_i \cap A_j = \varnothing whenever i≠ji \neq j, and A1∪A2∪⋯∪An=BA_1 \cup A_2 \cup \cdots \cup A_n = B.

Example. The sets A1={1,3}A_1 = \{1,3\}, A2={2}A_2 = \{2\} and A3={4,5}A_3 = \{4,5\} partition B={1,2,3,4,5}B = \{1,2,3,4,5\}.

Basically, a partition chops a set into non-overlapping pieces that between them use up everything. This is the single most important idea in the chapter: every major probability rule — the addition rule, the law of total probability, Bayes' theorem — works by partitioning a set and adding up the pieces, which is only legitimate because the pieces do not overlap.

Two identities get used constantly and are worth proving once.

Note

Theorem (De Morgan's laws)
For all A,B⊆SA, B \subseteq S,

(A∪B)c=Ac∩Bcand(A∩B)c=Ac∪Bc.(A \cup B)^c = A^c \cap B^c \qquad \text{and} \qquad (A \cap B)^c = A^c \cup B^c.

Proof. We prove the first; the second is identical with the roles of ∪\cup and ∩\cap swapped. For any x∈Sx \in S,

x∈(A∪B)c  ⟺  x∉A∪B  ⟺  it is not the case that (x∈A or x∈B)  ⟺  x∉A and x∉B  ⟺  x∈Ac and x∈Bc  ⟺  x∈Ac∩Bc.\begin{align*} x \in (A\cup B)^c &\iff x \notin A \cup B \\ &\iff \text{it is not the case that } (x \in A \text{ or } x \in B) \\ &\iff x \notin A \text{ and } x \notin B \\ &\iff x \in A^c \text{ and } x \in B^c \\ &\iff x \in A^c \cap B^c. \end{align*}

Since the two sets have exactly the same elements, they are equal. ■\blacksquare

Basically, De Morgan says that negating an "or" gives an "and", and vice versa. In probability language: "neither AA nor BB happened" is the same event as "not AA and not BB", and "not both" is the same as "at least one failed". Getting this backwards is a very common source of wrong answers in "at least one" problems.

Note

Definition
If AA is a finite set, ∣A∣|A| denotes the number of elements in AA.

Note

Theorem (inclusion–exclusion for two sets)
For finite sets AA and BB,  ∣A∪B∣=∣A∣+∣B∣−∣A∩B∣\ |A \cup B| = |A| + |B| - |A \cap B|.

Proof. Adding ∣A∣|A| and ∣B∣|B| counts every element of A∩BA \cap B exactly twice — once as a member of AA and once as a member of BB — while every other element of A∪BA \cup B is counted once. Subtracting ∣A∩B∣|A\cap B| removes the double count. ■\blacksquare

Example. Of 2020 music students, 77 play guitar, 88 play piano, and 33 play both. How many play at least one of the two, and how many play neither?
Let GG and PP be the sets of guitarists and pianists, so ∣G∣=7|G| = 7, ∣P∣=8|P| = 8 and ∣G∩P∣=3|G \cap P| = 3. Then

∣G∪P∣=∣G∣+∣P∣−∣G∩P∣=7+8−3=12,\begin{align*} |G \cup P| &= |G| + |P| - |G\cap P| \\ &= 7 + 8 - 3 \\ &= 12, \end{align*}

and by De Morgan the students who play neither form (G∪P)c(G\cup P)^c, of size 20−12=820 - 12 = 8. Therefore 1212 students play at least one instrument and 88 play neither. Notice that 7+8=15≠127 + 8 = 15 \neq 12; forgetting to subtract the overlap is the classic error, and it is exactly the error the addition rule for probabilities is designed to prevent.

9.2 Probability#

9.2.1 Sample space and probability axioms#

An experiment is any procedure with an uncertain outcome.

Note

Definition
The sample space SS of an experiment is the set of all its possible outcomes. An event is any subset A⊆SA \subseteq S.

Example. Tossing a coin is an experiment with sample space S={H,T}S = \{H, T\}.

Example. Tossing a coin three times and recording the ordered sequence has sample space

S={HHH,HHT,HTH,HTT,THH,THT,TTH,TTT},S = \{HHH, HHT, HTH, HTT, THH, THT, TTH, TTT\},

with ∣S∣=23=8|S| = 2^3 = 8. The event AA that exactly two heads are tossed is the subset

A={HHT,HTH,THH}.A = \{HHT, HTH, THH\}.

The choice of sample space is part of modelling, not something handed to you. If we only cared about the number of heads we could have taken S={0,1,2,3}S = \{0,1,2,3\} — but then the outcomes would not be equally likely, which usually makes life harder. Choose a sample space in which the outcomes are equally likely whenever you can.

Note

Definition
A probability PP on a sample space SS is a function assigning a real number to each event, such that

(a)P(A)≥0 for every event A;(b)P(S)=1;(c)P(∅)=0;(d)P(A∪B)=P(A)+P(B) whenever A and B are disjoint.\begin{align*} \text{(a)} \quad &P(A) \geq 0 \text{ for every event } A; \\ \text{(b)} \quad &P(S) = 1; \\ \text{(c)} \quad &P(\varnothing) = 0; \\ \text{(d)} \quad &P(A \cup B) = P(A) + P(B) \text{ whenever } A \text{ and } B \text{ are disjoint}. \end{align*}

Basically, probability is a way of measuring how big an event is compared with the whole sample space: never negative, the whole space has measure 11, and non-overlapping pieces add. Condition (d) requires disjointness and is worthless without it — this is where almost every incorrect probability calculation goes wrong.

Note

Theorem
Let PP be a probability on a sample space SS and AA an event.
1. If SS is finite or countable, then P(A)=∑a∈AP({a})P(A) = \sum_{a \in A}P(\{a\}).
2. If SS is finite and P({a})P(\{a\}) is the same for every outcome a∈Sa \in S, then P(A)=∣A∣∣S∣P(A) = \dfrac{|A|}{|S|}.
3. If SS is finite or countable, then ∑a∈SP({a})=1\sum_{a \in S}P(\{a\}) = 1.

Proof. Statement 1 follows for finite SS by induction on ∣A∣|A| using axiom (d), since the singletons {a}\{a\} for a∈Aa \in A are pairwise disjoint and partition AA. Statement 3 is statement 1 applied to A=SA = S, together with P(S)=1P(S) = 1. For statement 2, suppose P({a})=pP(\{a\}) = p for every a∈Sa \in S. By statement 3,

1=∑a∈SP({a})=∑a∈Sp=∣S∣ p,1 = \sum_{a\in S}P(\{a\}) = \sum_{a \in S}p = |S|\,p,

so p=1∣S∣p = \frac{1}{|S|}. Then by statement 1,

P(A)=∑a∈AP({a})=∑a∈A1∣S∣=∣A∣∣S∣.■P(A) = \sum_{a\in A}P(\{a\}) = \sum_{a\in A}\frac{1}{|S|} = \frac{|A|}{|S|}. \qquad \blacksquare

Statement 2 is the formula you have used since school — probability equals favourable outcomes over total outcomes. But it is only valid when the outcomes are equally likely, and it is a theorem, not a definition. Applying it to a sample space with unequally likely outcomes is a serious error.

Example. Pick a ball at random from a bag containing 33 red and 77 blue balls. Taking SS to be the ten individual balls, each equally likely, the event RR that the ball is red has ∣R∣=3|R| = 3, so

P(R)=∣R∣∣S∣=310.P(R) = \frac{|R|}{|S|} = \frac{3}{10}.

Therefore the probability of drawing red is 0.30.3. Notice that had we taken S={red,blue}S = \{\text{red}, \text{blue}\} we could not have used the formula, since those two outcomes are not equally likely.

9.2.2 Rules for probabilities#

Note

Theorem
Let AA and BB be events of a sample space SS.
1. P(A∪B)=P(A)+P(B)−P(A∩B)P(A\cup B) = P(A) + P(B) - P(A\cap B) (the addition rule).
2. P(Ac)=1−P(A)P(A^c) = 1 - P(A) (the complement rule).
3. If A⊆BA \subseteq B then P(A)≤P(B)P(A) \leq P(B).

Proof.

  1. The sets A−BA - B, A∩BA\cap B and B−AB - A partition A∪BA \cup B, so by axiom (d),

P(A∪B)=P(A−B)+P(A∩B)+P(B−A).P(A\cup B) = P(A-B) + P(A\cap B) + P(B-A).

Also A−BA - B and A∩BA \cap B partition AA, giving P(A−B)=P(A)−P(A∩B)P(A-B) = P(A) - P(A\cap B), and similarly P(B−A)=P(B)−P(A∩B)P(B-A) = P(B) - P(A\cap B). Substituting,

P(A∪B)=(P(A)−P(A∩B))+P(A∩B)+(P(B)−P(A∩B))=P(A)+P(B)−P(A∩B).P(A\cup B) = \big(P(A) - P(A\cap B)\big) + P(A\cap B) + \big(P(B) - P(A\cap B)\big) = P(A)+P(B)-P(A\cap B).

  1. The sets AA and AcA^c partition SS, so P(A)+P(Ac)=P(S)=1P(A) + P(A^c) = P(S) = 1.
  2. If A⊆BA \subseteq B then AA and B−AB - A partition BB, so P(B)=P(A)+P(B−A)≥P(A)P(B) = P(A) + P(B-A) \geq P(A) by axiom (a). ■\blacksquare

Notice that this is exactly the inclusion–exclusion principle with ∣⋅∣|\cdot| replaced by PP, and for the same reason: the overlap gets counted twice and must be removed once.

The complement rule is the most useful computational tool in the whole chapter. Whenever a question asks for the probability of "at least one" of something, computing P(none)P(\text{none}) and subtracting from 11 is almost always dramatically easier, because "none" is a single intersection while "at least one" is a messy union.

Example. What is the probability that at least two of nn people share a birthday?
Ignoring leap years, let YY be the 365365 days of the year. The sample space of all possible birthday lists is

Sn={(b1,…,bn):b1,…,bn∈Y},∣Sn∣=365n,S_n = \{(b_1,\dots,b_n) : b_1,\dots,b_n \in Y\}, \qquad |S_n| = 365^n,

with all outcomes equally likely. Let AA be the event that at least two people share a birthday. Computing P(A)P(A) directly would require a horrible union over all pairs, so use the complement: AcA^c is the event that all nn birthdays are distinct, and such a list is an ordered selection of nn different days,

∣Ac∣=365×364×⋯×(365−n+1).|A^c| = 365 \times 364 \times \cdots \times (365 - n + 1).

Therefore

P(A)=1−P(Ac)=1−365×364×⋯×(365−n+1)365n.P(A) = 1 - P(A^c) = 1 - \frac{365 \times 364 \times \cdots \times (365-n+1)}{365^n}.

Evaluating this gives famously counter-intuitive numbers:

nn 2323 3030 5050 7070
P(A)P(A) 0.5070.507 0.7060.706 0.9700.970 0.9990.999

So with just 2323 people the odds are already better than even. Therefore the answer is 1−365!(365−n)! 365n1 - \frac{365!}{(365-n)!\,365^n}. The intuition that misleads people here is comparing themselves to everyone else (n−1n-1 comparisons) rather than counting all (n2)\binom{n}{2} pairs, which for n=23n=23 is 253253 pairs — plenty of chances for a match.

Example. In a town, 80%80\% of the population has comprehensive car insurance, 60%60\% has house insurance, and 50%50\% has both. What proportion has at least one, and what proportion has neither?
Let CC and HH be the events of having car and house cover, so P(C)=0.8P(C) = 0.8, P(H)=0.6P(H) = 0.6 and P(C∩H)=0.5P(C\cap H) = 0.5. By the addition rule,

P(C∪H)=P(C)+P(H)−P(C∩H)=0.8+0.6−0.5=0.9,\begin{align*} P(C\cup H) &= P(C) + P(H) - P(C\cap H) \\ &= 0.8 + 0.6 - 0.5 \\ &= 0.9, \end{align*}

and by De Morgan together with the complement rule,

P(Cc∩Hc)=P((C∪H)c)=1−0.9=0.1.P(C^c\cap H^c) = P\big((C\cup H)^c\big) = 1 - 0.9 = 0.1.

Therefore 90%90\% have at least one policy and 10%10\% have neither.

9.2.3 Conditional probabilities#

Often we learn something partway through, and want to update a probability in the light of it. If we know that BB has occurred, then BB effectively becomes the new sample space, and we should measure AA only by the part of it lying inside BB.

Note

Definition
For events AA and BB with P(B)>0P(B) > 0, the conditional probability of AA given BB is

P(A∣B)=P(A∩B)P(B).P(A \mid B) = \frac{P(A\cap B)}{P(B)}.

Basically, we restrict attention to BB and rescale so that BB has total probability 11; dividing by P(B)P(B) is exactly that rescaling. The condition P(B)>0P(B) > 0 is not a technicality — conditioning on an impossible event is meaningless.

Rearranging gives the form used most often in practice:

P(A∩B)=P(A∣B) P(B)=P(B∣A) P(A).\boxed{P(A\cap B) = P(A\mid B)\,P(B) = P(B\mid A)\,P(A).}

Example. We roll a die. Let AA be the event of rolling a six and BB the event of rolling an even number. Find P(A∣B)P(A\mid B) and P(B∣A)P(B\mid A).
Here A={6}A = \{6\} and B={2,4,6}B = \{2,4,6\}, so A∩B={6}A\cap B = \{6\}, and

P(A∣B)=P(A∩B)P(B)=1/63/6=13,P(B∣A)=P(A∩B)P(A)=1/61/6=1.\begin{align*} P(A\mid B) &= \frac{P(A\cap B)}{P(B)} = \frac{1/6}{3/6} = \frac13, \\ P(B\mid A) &= \frac{P(A\cap B)}{P(A)} = \frac{1/6}{1/6} = 1. \end{align*}

Therefore P(A∣B)=13P(A\mid B) = \frac13 and P(B∣A)=1P(B\mid A) = 1. Note how different the two are: P(A∣B)≠P(B∣A)P(A\mid B) \neq P(B\mid A) in general, and confusing the two is so common in law and medicine that it has a name — the prosecutor's fallacy. Knowing the roll is even makes a six more likely than before; knowing it is a six makes it certainly even.

Example. A bag contains 33 red and 33 blue balls. Two balls are drawn without replacement. Find the probability that both are red.
Let R1R_1 and R2R_2 be the events that the first and second balls are red. Before any draw, P(R1)=36=12P(R_1) = \frac36 = \frac12. Given the first ball was red, only 55 balls remain of which 22 are red, so P(R2∣R1)=25P(R_2\mid R_1) = \frac25. Therefore

P(R1∩R2)=P(R2∣R1)P(R1)=25×12=15.P(R_1\cap R_2) = P(R_2\mid R_1)P(R_1) = \frac25\times\frac12 = \frac15.

Drawing without replacement is what makes the draws dependent; with replacement, P(R2∣R1)P(R_2\mid R_1) would still be 12\frac12 and the answer would be 14\frac14. Always check which of the two a question intends.

Conditioning becomes really powerful when combined with a partition.

Note

Theorem (the law of total probability)
Suppose the events B1,B2,…,BnB_1, B_2, \dots, B_n partition the sample space SS, each with P(Bi)>0P(B_i) > 0. Then for any event AA,

P(A)=∑i=1nP(A∣Bi) P(Bi).P(A) = \sum_{i=1}^{n}P(A\mid B_i)\,P(B_i).

Proof. Since the BiB_i partition SS, the sets A∩B1,…,A∩BnA\cap B_1, \dots, A\cap B_n are pairwise disjoint and their union is AA. By axiom (d) applied repeatedly and then the definition of conditional probability,

P(A)=∑i=1nP(A∩Bi)=∑i=1nP(A∣Bi)P(Bi).■P(A) = \sum_{i=1}^{n}P(A\cap B_i) = \sum_{i=1}^{n}P(A\mid B_i)P(B_i). \qquad \blacksquare

Basically, to find the probability of AA, split the world into cases, work out the chance of AA in each case, and average those weighted by how likely each case is. This is exactly what a tree diagram does: each path through the tree is one term of the sum, and the probability of a path is the product of the probabilities along its edges.

Note

Theorem (Bayes' theorem)
Suppose B1,…,BnB_1,\dots,B_n partition SS with each P(Bi)>0P(B_i) > 0, and AA is an event with P(A)>0P(A) > 0. Then

P(Bj∣A)=P(A∣Bj)P(Bj)∑i=1nP(A∣Bi)P(Bi).P(B_j\mid A) = \frac{P(A\mid B_j)P(B_j)}{\sum_{i=1}^{n}P(A\mid B_i)P(B_i)}.

Proof. By the definition of conditional probability twice over,

P(Bj∣A)=P(A∩Bj)P(A)=P(A∣Bj)P(Bj)P(A),P(B_j\mid A) = \frac{P(A\cap B_j)}{P(A)} = \frac{P(A\mid B_j)P(B_j)}{P(A)},

and expanding the denominator by the law of total probability gives the result. ■\blacksquare

Basically, Bayes' theorem reverses a conditional probability: it converts "the chance of the evidence given the cause" into "the chance of the cause given the evidence", which is almost always the direction you actually want. The following example is the standard illustration of how badly intuition fails here, and it is worth studying carefully.

Example. A diagnostic test for disease XX correctly indicates the disease 99%99\% of the time in people who have it, and correctly returns negative 98%98\% of the time in people who do not. Suppose 2%2\% of the population has the disease. Find the probability that a person who tests positive does not have the disease.
Let DD be the event of having the disease and TpT_p the event of testing positive, with Tn=TpcT_n = T_p^c. We are given

P(D)=0.02,P(Tp∣D)=0.99,P(Tn∣Dc)=0.98,so P(Tp∣Dc)=0.02.P(D) = 0.02, \quad P(T_p\mid D) = 0.99, \quad P(T_n\mid D^c) = 0.98, \quad \text{so } P(T_p\mid D^c) = 0.02.

The events DD and DcD^c partition SS, so by Bayes' theorem,

P(Dc∣Tp)=P(Tp∣Dc)P(Dc)P(Tp∣Dc)P(Dc)+P(Tp∣D)P(D)=0.02×0.980.02×0.98+0.99×0.02=0.01960.0196+0.0198=0.01960.0394≈0.497.\begin{align*} P(D^c\mid T_p) &= \frac{P(T_p\mid D^c)P(D^c)}{P(T_p\mid D^c)P(D^c) + P(T_p\mid D)P(D)} \\ &= \frac{0.02\times 0.98}{0.02\times 0.98 + 0.99\times 0.02} \\ &= \frac{0.0196}{0.0196 + 0.0198} \\ &= \frac{0.0196}{0.0394} \\ &\approx 0.497. \end{align*}

Therefore almost half of all positive results are false, despite the test being 9898–99%99\% accurate. The reason is that the disease is rare: out of 10 00010\,000 people, about 200200 have it and 198198 of those test positive, while of the 98009800 healthy people 2%2\% — that is, 196196 people — also test positive. The two groups are almost the same size, so a positive result is nearly a coin flip.

By contrast, a negative result is extremely informative:

P(D∣Tn)=P(Tn∣D)P(D)P(Tn∣D)P(D)+P(Tn∣Dc)P(Dc)=0.01×0.020.01×0.02+0.98×0.98=0.00020.9606≈0.000208.\begin{align*} P(D\mid T_n) &= \frac{P(T_n\mid D)P(D)}{P(T_n\mid D)P(D) + P(T_n\mid D^c)P(D^c)} \\ &= \frac{0.01\times 0.02}{0.01\times0.02 + 0.98\times0.98} \\ &= \frac{0.0002}{0.9606} \\ &\approx 0.000208. \end{align*}

So someone testing negative almost certainly does not have the disease. The lesson is that the accuracy of a test tells you very little on its own; the base rate P(D)P(D) matters just as much, and ignoring it is called the base rate fallacy. In practice this is why a cheap screening test is followed up by a more expensive confirmatory one.

Example. Modify the previous example by supposing a fraction xx of the population has disease XX, and that 5%5\% of a large random sample tests positive. What percentage of the population has the disease?
By the law of total probability,

0.05=P(Tp)=P(Tp∣D)P(D)+P(Tp∣Dc)P(Dc)=0.99x+0.02(1−x)=0.97x+0.02.\begin{align*} 0.05 = P(T_p) &= P(T_p\mid D)P(D) + P(T_p\mid D^c)P(D^c) \\ &= 0.99x + 0.02(1-x) \\ &= 0.97x + 0.02. \end{align*}

Solving, 0.97x=0.030.97x = 0.03, so

x=0.030.97≈0.031.x = \frac{0.03}{0.97} \approx 0.031.

Therefore about 3.1%3.1\% of the population has the disease. Notice that the observed positive rate of 5%5\% substantially overstates the true prevalence of 3.1%3.1\%, again because of false positives among the large healthy majority.

9.2.4 Statistical independence#

Sometimes learning that BB occurred tells us nothing at all about AA. That deserves a definition.

Note

Definition
Events AA and BB are (statistically) independent if

P(A∩B)=P(A)P(B).P(A\cap B) = P(A)P(B).

When P(B)>0P(B) > 0 this is equivalent to P(A∣B)=P(A)P(A\mid B) = P(A), which is the more intuitive form: conditioning on BB leaves the probability of AA unchanged. The product form is preferred as the definition because it is symmetric in AA and BB and does not require P(B)>0P(B) > 0.

Independent and mutually exclusive are completely different things, and confusing them is the most common conceptual error in this chapter. In fact they are almost opposites: if AA and BB are mutually exclusive with P(A),P(B)>0P(A), P(B) > 0, then P(A∩B)=P(∅)=0P(A\cap B) = P(\varnothing) = 0 but P(A)P(B)>0P(A)P(B) > 0, so they are necessarily dependent. Intuitively, if AA and BB cannot both happen, then learning AA occurred tells you a great deal about BB — namely that it did not occur.

Example. Roll a die; let AA be rolling a six and BB rolling an even number. Are AA and BB independent?
Here P(A)=16P(A) = \frac16, P(B)=12P(B) = \frac12 and P(A∩B)=P({6})=16P(A\cap B) = P(\{6\}) = \frac16. But

P(A)P(B)=16×12=112≠16=P(A∩B),P(A)P(B) = \frac16\times\frac12 = \frac{1}{12} \neq \frac16 = P(A\cap B),

so they are not independent. This matches the earlier calculation that P(A∣B)=13≠16=P(A)P(A\mid B) = \frac13 \neq \frac16 = P(A).

Example. Roll a die twice. For each ii, let XiX_i be the event that the first roll gives ii and YiY_i that the second roll gives ii. Then P(Xi)=P(Yj)=16P(X_i) = P(Y_j) = \frac16, and since each of the 3636 ordered pairs is equally likely,

P(Xi∩Yj)=136=16×16=P(Xi)P(Yj),P(X_i\cap Y_j) = \frac{1}{36} = \frac16\times\frac16 = P(X_i)P(Y_j),

so XiX_i and YjY_j are independent for all ii and jj — as they should be, since the die has no memory.

For more than two events, independence has to be required for every sub-collection, not just for pairs.

Note

Definition
Events A1,…,AnA_1,\dots,A_n are mutually independent if for every sub-collection Ai1,…,AikA_{i_1},\dots,A_{i_k},

P(Ai1∩⋯∩Aik)=P(Ai1)⋯P(Aik).P(A_{i_1}\cap\cdots\cap A_{i_k}) = P(A_{i_1})\cdots P(A_{i_k}).

So AA, BB, CC are mutually independent exactly when all four of these hold:

P(A∩B)=P(A)P(B),P(A∩C)=P(A)P(C),P(B∩C)=P(B)P(C),P(A∩B∩C)=P(A)P(B)P(C).\begin{align*} P(A\cap B) &= P(A)P(B), \\ P(A\cap C) &= P(A)P(C), \\ P(B\cap C) &= P(B)P(C), \\ P(A\cap B\cap C) &= P(A)P(B)P(C). \end{align*}

Pairwise independence does not imply mutual independence. The standard counterexample: draw a ball at random from a bag of four balls marked 0,1,2,30, 1, 2, 3, and let AA, BB, CC be the events that the ball is 11 or 33, that it is 22 or 33, and that it is 00 or 33. Each has probability 12\frac12, and each pairwise intersection is {3}\{3\} with probability 14=12×12\frac14 = \frac12\times\frac12, so the events are pairwise independent. But A∩B∩C={3}A\cap B\cap C = \{3\} also has probability 14\frac14, whereas P(A)P(B)P(C)=18P(A)P(B)P(C) = \frac18. So the three are pairwise but not mutually independent.

Note

Theorem
If A1,…,AnA_1,\dots,A_n are mutually independent and each BiB_i is either AiA_i or AicA_i^c, then B1,…,BnB_1,\dots,B_n are mutually independent.

Basically, independence survives complementing — if AA tells you nothing about BB, then AA tells you nothing about "not BB" either. This is what lets us compute "at least one" probabilities by complementing each event separately.

Example. (Reliability.) A system has three components that fail independently, with failure probabilities 0.10.1, 0.20.2 and 0.30.3. Find the probability the system works, if it needs (a) all three components working, and (b) at least one component working.
Let F1,F2,F3F_1, F_2, F_3 be the failure events, so the working events FicF_i^c have probabilities 0.90.9, 0.80.8 and 0.70.7 and are mutually independent by the theorem above.

(a) All three working:

P(F1c∩F2c∩F3c)=0.9×0.8×0.7=0.504.P(F_1^c\cap F_2^c\cap F_3^c) = 0.9\times0.8\times0.7 = 0.504.

(b) At least one working is the complement of all three failing, so

P(at least one works)=1−P(F1∩F2∩F3)=1−0.1×0.2×0.3=1−0.006=0.994.\begin{align*} P(\text{at least one works}) &= 1 - P(F_1\cap F_2\cap F_3) \\ &= 1 - 0.1\times0.2\times0.3 \\ &= 1 - 0.006 \\ &= 0.994. \end{align*}

Therefore the series system works with probability 0.5040.504 and the parallel system with probability 0.9940.994. Notice how much redundancy buys you — and notice that (b) was computed via the complement, because "at least one" as a union would have needed the full three-set inclusion–exclusion formula.

9.3 Random variables#

Describing events as subsets of SS gets clumsy fast, and it gives us nothing to do arithmetic with. The fix is to attach a number to each outcome.

Note

Definition
A random variable is a real-valued function defined on a sample space.

A random variable is neither random nor a variable; it is a function. The name is entirely historical and actively misleading, so it is worth fixing the correct picture early: XX takes an outcome s∈Ss \in S and returns a number X(s)X(s). The randomness lives in which ss occurs, not in XX.

Example. Toss a coin, with S={H,T}S = \{H, T\}. Define X(H)=1X(H) = 1 and X(T)=0X(T) = 0. Then XX counts the heads.

Example. Toss a coin three times and let XX count the number of heads. Then X(HHT)=2X(HHT) = 2, X(TTT)=0X(TTT) = 0, and so on. The event "exactly two heads" is now written compactly as {X=2}\{X = 2\}, meaning {s∈S:X(s)=2}={HHT,HTH,THH}\{s \in S : X(s) = 2\} = \{HHT, HTH, THH\}, and

P(X=2)=38.P(X = 2) = \frac{3}{8}.

This notation is the whole point: {X=2}\{X = 2\}, {X≤2}\{X \leq 2\} and {X>0}\{X > 0\} are all just events, described far more conveniently than by listing outcomes.

9.3.1 Discrete random variables#

Note

Definition
A random variable is discrete if it takes only countably many values. Its probability distribution is the list of values xkx_k it takes together with the probabilities

pk=P(X=xk).p_k = P(X = x_k).

By the summation theorem, the pkp_k are non-negative and sum to 11; these two conditions characterise a valid probability distribution, and checking that your probabilities sum to 11 is the cheapest error check available.

Example. Toss a coin three times and let XX count the heads. Since all eight outcomes are equally likely and the numbers of outcomes giving 0,1,2,30,1,2,3 heads are 1,3,3,11,3,3,1:

xkx_k 00 11 22 33
pkp_k 18\frac18 38\frac38 38\frac38 18\frac18

and indeed 18+38+38+18=1\frac18+\frac38+\frac38+\frac18 = 1.

Example. Roll a die twice and let XX be the sum of the two rolls. The sample space has 3636 equally likely ordered pairs, and counting those summing to each value:

xkx_k 22 33 44 55 66 77 88 99 1010 1111 1212
pkp_k 136\frac1{36} 236\frac2{36} 336\frac3{36} 436\frac4{36} 536\frac5{36} 636\frac6{36} 536\frac5{36} 436\frac4{36} 336\frac3{36} 236\frac2{36} 136\frac1{36}

The numerators sum to 3636, as they must. Therefore 77 is the most likely total, with probability 16\frac16.

Note

Definition
The cumulative distribution function of a random variable XX is F(x)=P(X≤x)F(x) = P(X \leq x).

9.3.2 The mean and variance of a discrete random variable#

Random variables let us do arithmetic on outcomes. The two numbers we care about most are the weighted average of the values, and a measure of how spread out they are.

Note

Definition
The expected value (or mean) of a discrete random variable XX with distribution pk=P(X=xk)p_k = P(X = x_k) is

E(X)=∑all kxkpk,E(X) = \sum_{\text{all } k}x_kp_k,

often written μ\mu or μX\mu_X.

Basically, E(X)E(X) is the long-run average value of XX if the experiment were repeated many times, with each possible value weighted by how often it occurs. The expected value need not be a value XX can actually take — the expected number of heads in three tosses is 32\frac32, which is impossible on any single trial.

Example. Toss a coin three times and let XX count the heads. Using the table above,

E(X)=0×18+1×38+2×38+3×18=0+3+6+38=128=32.\begin{align*} E(X) &= 0\times\tfrac18 + 1\times\tfrac38 + 2\times\tfrac38 + 3\times\tfrac18 \\ &= \tfrac{0+3+6+3}{8} = \tfrac{12}{8} = \tfrac32. \end{align*}

Therefore E(X)=32E(X) = \frac32, which is exactly what symmetry suggests: half of three tosses.

Note

Theorem
Let XX be a discrete random variable with pk=P(X=xk)p_k = P(X = x_k). Then for any real function gg, the expected value of Y=g(X)Y = g(X) is

E(Y)=E(g(X))=∑kg(xk)pk.E(Y) = E\big(g(X)\big) = \sum_k g(x_k)p_k.

This is more useful than it looks: it says you can compute E(g(X))E(g(X)) without first finding the distribution of g(X)g(X), which would usually be a nuisance. In particular E(X2)=∑kxk2pkE(X^2) = \sum_k x_k^2p_k.

A warning that is worth a lot of marks: in general E(g(X))≠g(E(X))E(g(X)) \neq g(E(X)). For instance with the coin example, E(X2)=0+3+12+98=3E(X^2) = \frac{0+3+12+9}{8} = 3 while (E(X))2=94\big(E(X)\big)^2 = \frac94. The two are only equal when gg is linear.

Note

Definition
The variance of a discrete random variable XX with mean μ\mu is

Var⁡(X)=E((X−μ)2)=∑k(xk−μ)2pk,\operatorname{Var}(X) = E\big((X-\mu)^2\big) = \sum_k(x_k-\mu)^2p_k,

and the standard deviation is SD⁡(X)=Var⁡(X)\operatorname{SD}(X) = \sqrt{\operatorname{Var}(X)}, often written σ\sigma.

Basically, the variance is the average squared distance from the mean; squaring stops positive and negative deviations from cancelling. The standard deviation then undoes the squaring so that the answer is back in the original units, which is why it is the more interpretable of the two.

Note

Theorem
Var⁡(X)=E(X2)−(E(X))2\operatorname{Var}(X) = E(X^2) - \big(E(X)\big)^2.

Proof. Write μ=E(X)\mu = E(X). Expanding the square,

Var⁡(X)=∑k(xk−μ)2pk=∑k(xk2−2xkμ+μ2)pk=∑kxk2pk−2μ∑kxkpk+μ2∑kpk=E(X2)−2μ⋅μ+μ2⋅1=E(X2)−(E(X))2.■\begin{align*} \operatorname{Var}(X) &= \sum_k(x_k-\mu)^2p_k \\ &= \sum_k(x_k^2 - 2x_k\mu + \mu^2)p_k \\ &= \sum_kx_k^2p_k - 2\mu\sum_kx_kp_k + \mu^2\sum_kp_k \\ &= E(X^2) - 2\mu\cdot\mu + \mu^2\cdot 1 \\ &= E(X^2) - \big(E(X)\big)^2. \qquad \blacksquare \end{align*}

This is the formula to use for hand calculation, since it needs only two sums rather than a subtraction inside every term. As a free sanity check, Var⁡(X)≥0\operatorname{Var}(X) \geq 0 always; if your computation gives a negative variance you have made an arithmetic error, most often by squaring the mean before rather than after averaging.

Example. Toss a coin three times and let XX count the heads. We found E(X)=32E(X) = \frac32 and E(X2)=3E(X^2) = 3, so

Var⁡(X)=3−(32)2=3−94=34.\operatorname{Var}(X) = 3 - \left(\frac32\right)^2 = 3 - \frac94 = \frac34.

Therefore Var⁡(X)=34\operatorname{Var}(X) = \frac34 and SD⁡(X)=32≈0.866\operatorname{SD}(X) = \frac{\sqrt3}{2} \approx 0.866.

Example. A random variable XX has the distribution

xkx_k 00 11 22 44 77
pkp_k 0.20.2 0.060.06 0.20.2 0.040.04 0.50.5

Find its mean, variance and standard deviation.
First check the probabilities sum to 0.2+0.06+0.2+0.04+0.5=10.2+0.06+0.2+0.04+0.5 = 1. Then

E(X)=0(0.2)+1(0.06)+2(0.2)+4(0.04)+7(0.5)=0+0.06+0.4+0.16+3.5=4.12,E(X2)=02(0.2)+12(0.06)+22(0.2)+42(0.04)+72(0.5)=0+0.06+0.8+0.64+24.5=26.0,\begin{align*} E(X) &= 0(0.2) + 1(0.06) + 2(0.2) + 4(0.04) + 7(0.5) \\ &= 0 + 0.06 + 0.4 + 0.16 + 3.5 \\ &= 4.12, \\ E(X^2) &= 0^2(0.2) + 1^2(0.06) + 2^2(0.2) + 4^2(0.04) + 7^2(0.5) \\ &= 0 + 0.06 + 0.8 + 0.64 + 24.5 \\ &= 26.0, \end{align*}

so

Var⁡(X)=26.0−(4.12)2=26.0−16.9744=9.0256,\operatorname{Var}(X) = 26.0 - (4.12)^2 = 26.0 - 16.9744 = 9.0256,

and SD⁡(X)=9.0256≈3.004\operatorname{SD}(X) = \sqrt{9.0256} \approx 3.004. Therefore the values of XX sit on average about 33 away from the mean of 4.124.12. (The printed course notes give the variance here as 9.2569.256; that is a typo — 26−4.122=9.025626 - 4.12^2 = 9.0256, and their own conclusion that the standard deviation is roughly 33 matches the corrected figure.)

Note

Theorem
If aa and bb are constants, then

E(aX+b)=aE(X)+b,Var⁡(aX+b)=a2Var⁡(X),SD⁡(aX+b)=∣a∣SD⁡(X).E(aX+b) = aE(X)+b, \qquad \operatorname{Var}(aX+b) = a^2\operatorname{Var}(X), \qquad \operatorname{SD}(aX+b) = |a|\operatorname{SD}(X).

Proof. For the mean,

E(aX+b)=∑k(axk+b)pk=a∑kxkpk+b∑kpk=aE(X)+b.E(aX+b) = \sum_k(ax_k+b)p_k = a\sum_kx_kp_k + b\sum_kp_k = aE(X)+b.

For the variance, using the mean result,

Var⁡(aX+b)=E((aX+b−E(aX+b))2)=E((aX+b−aE(X)−b)2)=E(a2(X−E(X))2)=a2Var⁡(X),\begin{align*} \operatorname{Var}(aX+b) &= E\Big(\big(aX+b - E(aX+b)\big)^2\Big) \\ &= E\Big(\big(aX + b - aE(X) - b\big)^2\Big) \\ &= E\Big(a^2\big(X-E(X)\big)^2\Big) \\ &= a^2\operatorname{Var}(X), \end{align*}

and the third statement follows by taking square roots. ■\blacksquare

Notice that bb disappears from the variance entirely. Shifting every value by a constant moves the mean but does not change the spread, which is exactly what a measure of dispersion ought to do. The a2a^2 rather than aa is because variance is measured in squared units.

9.4 Special distributions#

A handful of distributions turn up so often that they are worth studying once and for all. Both in this section arise from the same setup.

Note

Definition
A Bernoulli process is a sequence of trials such that

  • each trial has exactly two outcomes, "success" and "failure";
  • the probability of success pp is the same on every trial; and
  • the trials are mutually independent.

We write q=1−pq = 1 - p for the failure probability.

All three conditions matter, and questions are frequently designed so that one of them fails. Drawing balls without replacement, for instance, is not a Bernoulli process, because pp changes from trial to trial and the draws are not independent.

9.4.1 The binomial distribution#

Note

Theorem
If XX counts the successes in a Bernoulli process of nn trials with success probability pp, then

P(X=k)=(nk)pkqn−k,k=0,1,…,n.P(X = k) = \binom{n}{k}p^kq^{n-k}, \qquad k = 0, 1, \dots, n.

We write X∼B(n,p)X \sim B(n,p) and call this the binomial distribution.

Proof. Any particular sequence of nn trials containing exactly kk successes and n−kn-k failures has probability pkqn−kp^kq^{n-k}, since the trials are independent and so the probabilities multiply. The number of such sequences is the number of ways to choose which kk of the nn trials are the successes, namely (nk)\binom{n}{k}. Summing over these disjoint outcomes gives the result. ■\blacksquare

Basically, the pkqn−kp^kq^{n-k} is the probability of one particular pattern and the (nk)\binom{n}{k} counts how many patterns there are. The binomial coefficient is exactly what students leave out — if the question asks for a specific order ("heads then tails then heads") there is no coefficient, but if it asks for a count ("two heads in three tosses") there is.

As a check that this is a distribution at all, the binomial theorem gives

∑k=0n(nk)pkqn−k=(p+q)n=1n=1,\sum_{k=0}^{n}\binom{n}{k}p^kq^{n-k} = (p+q)^n = 1^n = 1,

which is also where the name comes from.

Example. Roll a die 1212 times and let XX be the number of sixes. The rolls are identical and independent with p=16p = \frac16, so X∼B(12,16)X \sim B(12, \frac16) and

P(X=2)=(122)(16)2(56)10=66×136×(56)10≈0.296.P(X = 2) = \binom{12}{2}\left(\frac16\right)^2\left(\frac56\right)^{10} = 66 \times \frac{1}{36}\times\left(\frac56\right)^{10} \approx 0.296.

Therefore about a 29.6%29.6\% chance of exactly two sixes.

Example. Ask n=40n = 40 people whether today is their birthday, and let XX count the yes answers. Ignoring leap years and twins, X∼B(40,1365)X \sim B(40, \frac{1}{365}), and the probability that nobody says yes is

P(X=0)=(400)(1365)0(364365)40=(364365)40≈0.896.P(X=0) = \binom{40}{0}\left(\frac{1}{365}\right)^0\left(\frac{364}{365}\right)^{40} = \left(\frac{364}{365}\right)^{40} \approx 0.896.

Therefore there is about a 10.4%10.4\% chance that at least one person has a birthday today. Compare this with the birthday problem of Section 9.2.2, which asked a completely different question — there we compared people with each other (many pairs), here we compare each person against one fixed day. The two are constantly confused.

Note

Theorem
If X∼B(n,p)X \sim B(n,p), then E(X)=npE(X) = np and Var⁡(X)=npq\operatorname{Var}(X) = npq.

Basically, each trial contributes on average pp successes and variance pqpq, and with nn independent trials these simply add. The mean npnp is exactly the intuitive answer: rolling a die 1212 times, you expect 12×16=212\times\frac16 = 2 sixes.

Example. Toss a coin n=3n = 3 times and let XX count the heads. Then X∼B(3,12)X \sim B(3, \frac12), so

E(X)=3×12=32,Var⁡(X)=3×12×12=34,E(X) = 3\times\tfrac12 = \tfrac32, \qquad \operatorname{Var}(X) = 3\times\tfrac12\times\tfrac12 = \tfrac34,

which agrees exactly with the values we computed by hand from the distribution table earlier — a satisfying check on both.

Example. Roll a die 1212 times. Then E(X)=12×16=2E(X) = 12\times\frac16 = 2 sixes, with Var⁡(X)=12×16×56=53\operatorname{Var}(X) = 12\times\frac16\times\frac56 = \frac53 and SD⁡(X)≈1.29\operatorname{SD}(X) \approx 1.29. So two sixes is typical, and anything from about 11 to 33 would be unremarkable.

9.4.2 Geometric distribution#

Instead of fixing the number of trials and counting successes, we can fix the number of successes at one and count the trials.

Note

Theorem
Consider an infinite Bernoulli process with success probability pp. If XX is the number of trials until the first success, then

P(X=k)=qk−1p,k=1,2,3,…P(X = k) = q^{k-1}p, \qquad k = 1, 2, 3, \dots

We write X∼G(p)X \sim G(p) and call this the geometric distribution.

Proof. The event {X=k}\{X = k\} consists of exactly one outcome: the first k−1k-1 trials all fail and the kkth succeeds. By independence its probability is qk−1pq^{k-1}p. ■\blacksquare

Note that in principle a success might never occur, but that event has probability lim⁡k→∞qk=0\lim_{k\to\infty}q^k = 0 when p>0p > 0, so we discard it. As a check the probabilities sum correctly, using the geometric series from Section 4.4:

∑k=1∞qk−1p=p∑j=0∞qj=p1−q=pp=1.\sum_{k=1}^{\infty}q^{k-1}p = p\sum_{j=0}^{\infty}q^j = \frac{p}{1-q} = \frac{p}{p} = 1.

Be careful about the convention. Here XX counts all trials including the successful one, so X≥1X \geq 1. Some texts instead count only the failures before the success, giving P(X=k)=qkpP(X=k) = q^kp for k≥0k \geq 0 and a mean one smaller. Always check which convention a question uses; this off-by-one is the standard trap.

Example. Toss a coin until a head appears and let XX count the tosses. Then X∼G(12)X \sim G(\frac12), and the probability of needing exactly seven tosses is

P(X=7)=(12)6×12=127=1128≈0.8%.P(X = 7) = \left(\frac12\right)^{6}\times\frac12 = \frac{1}{2^7} = \frac{1}{128} \approx 0.8\%.

Note

Theorem
If X∼G(p)X \sim G(p) and nn is a positive integer, then P(X>n)=qnP(X > n) = q^n. Consequently the cumulative distribution function is F(x)=P(X≤x)=1−q⌊x⌋F(x) = P(X \leq x) = 1 - q^{\lfloor x\rfloor}.

Proof. The event {X>n}\{X > n\} says the first nn trials all failed, which by independence has probability qnq^n. The rest follows by complementing. ■\blacksquare

That is a genuinely useful shortcut — the tail probability requires no summation at all, so questions about "how many trials might it take" should always be attacked through P(X>n)P(X>n) rather than by summing the distribution.

Example. Roll a die until a six appears, and let XX count the rolls. Then X∼G(16)X \sim G(\frac16), and

P(X≤4)=1−(56)4=1−0.4823≈52%,P(X≤6)=1−(56)6=1−0.3349≈67%,P(X>7)=(56)7≈28%.\begin{align*} P(X \leq 4) &= 1 - \left(\tfrac56\right)^4 = 1 - 0.4823 \approx 52\%, \\ P(X \leq 6) &= 1 - \left(\tfrac56\right)^6 = 1 - 0.3349 \approx 67\%, \\ P(X > 7) &= \left(\tfrac56\right)^7 \approx 28\%. \end{align*}

Therefore you have roughly an even chance of rolling a six within four attempts, and a better-than-a-quarter chance of still waiting after seven.

Note

Theorem
If X∼G(p)X \sim G(p), then E(X)=1pE(X) = \dfrac1p and Var⁡(X)=1−pp2\operatorname{Var}(X) = \dfrac{1-p}{p^2}.

Proof (of the mean). Using term-by-term differentiation of power series,

∑k=1∞kqk−1=ddq∑k=0∞qk=ddq(11−q)=1(1−q)2,\sum_{k=1}^{\infty}kq^{k-1} = \frac{d}{dq}\sum_{k=0}^{\infty}q^k = \frac{d}{dq}\left(\frac{1}{1-q}\right) = \frac{1}{(1-q)^2},

valid for ∣q∣<1|q|<1. Hence

E(X)=∑k=1∞k qk−1p=p(1−q)2=pp2=1p.■E(X) = \sum_{k=1}^{\infty}k\,q^{k-1}p = \frac{p}{(1-q)^2} = \frac{p}{p^2} = \frac1p. \qquad \blacksquare

This is a nice illustration of Chapter 4 doing real work in a statistics problem.

The mean is extremely intuitive: if a success happens one time in pp, you expect to wait 1p\frac1p trials for it.

Example. Toss a coin until a head appears: p=12p = \frac12, so E(X)=2E(X) = 2 tosses on average. Roll a die until a six appears: p=16p = \frac16, so E(X)=6E(X) = 6 rolls on average, with Var⁡(X)=5/61/36=30\operatorname{Var}(X) = \frac{5/6}{1/36} = 30 and SD⁡(X)≈5.5\operatorname{SD}(X) \approx 5.5. Note how enormous that standard deviation is relative to the mean — waiting times are highly variable, which is why "it took me 20 rolls" is not really surprising.

9.4.3 Sign tests#

We finish the discrete material with a first taste of statistical inference. Suppose we have independent observations of some quantity and want to know whether they differ systematically from a fixed reference value. A sign test answers this using nothing but the binomial distribution:

  • Count the observations strictly greater than the target value ("+").
  • Count the total that are strictly greater or strictly smaller (discarding exact ties).
  • Compute the probability of seeing at least as many "+" as observed, assuming "+" and "−" were equally likely. This is the tail probability.

The logic is a proof by contradiction in probabilistic clothing. We assume there is no real effect (so p=12p = \frac12), compute how surprising our data would be under that assumption, and reject the assumption if the data are too surprising. In this course we call a tail probability below 5%5\% significant.

Example. A new variety of corn yields, in bushels per acre, on 1515 plots:

138.0138.0 139.1139.1 113.0113.0 132.5132.5 140.7140.7
109.7109.7 118.0118.0 134.8134.8 109.6109.6 127.3127.3
115.6115.6 130.4130.4 130.2130.2 117.7117.7 105.5105.5

The variety currently in use yields 110110 bushels per acre. Is the new variety better?
Counting, the yield exceeds 110110 on 1212 plots and falls below on 33; there are no ties, so all 1515 observations count. Under the assumption that the new variety is no better, a plot is equally likely to exceed or fall short of 110110, so the number XX of "+" plots satisfies X∼B(15,12)X \sim B(15, \frac12). The tail probability is

P(X≥12)=∑k=1215(15k)(12)k(12)15−k=1215[(1512)+(1513)+(1514)+(1515)]=455+105+15+132768=57632768≈1.76%.\begin{align*} P(X \geq 12) &= \sum_{k=12}^{15}\binom{15}{k}\left(\frac12\right)^k\left(\frac12\right)^{15-k} \\ &= \frac{1}{2^{15}}\left[\binom{15}{12}+\binom{15}{13}+\binom{15}{14}+\binom{15}{15}\right] \\ &= \frac{455 + 105 + 15 + 1}{32768} \\ &= \frac{576}{32768} \\ &\approx 1.76\%. \end{align*}

Since 1.76%<5%1.76\% < 5\%, it is quite unlikely we would have seen 1212 of 1515 plots above the old yield if the new variety were no better. Therefore we conclude the new variety has improved yield.

Two warnings about what this does and does not show. First, the test uses only the signs, throwing away the magnitudes — which makes it robust but not very powerful. Second, a tail probability of 1.76%1.76\% is not "the probability that the new variety is no better"; it is the probability of data this extreme given that it is no better. Reversing those two is exactly the conditional-probability error from Section 9.2.3.

9.5 Continuous random variables#

So far our random variables have taken countably many values. But many quantities — a time, a height, a wage — vary continuously, and for these the whole framework needs adjusting.

The essential difficulty: if XX is the exact time of day, uniformly distributed over 2424 hours, what is P(X=3pm exactly)P(X = \text{3pm exactly})? There are uncountably many instants, all equally likely, so any positive probability would sum to infinity. The only consistent answer is 00.

Note

Definition
A random variable XX is continuous with probability density function ff if

P(a≤X≤b)=∫abf(x) dxP(a \leq X \leq b) = \int_a^b f(x)\,dx

for all a≤ba \leq b, where ff satisfies f(x)≥0f(x) \geq 0 for all xx and ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty}f(x)\,dx = 1.

The density f(x)f(x) is NOT a probability, and it may exceed 11. Probability is area under the density, not the height of it. This is the single biggest conceptual jump in the chapter, and the two defining conditions are the exact analogues of "probabilities are non-negative" and "probabilities sum to 11".

An immediate consequence is that

P(X=a)=∫aaf(x) dx=0P(X = a) = \int_a^af(x)\,dx = 0

for every single value aa. So for a continuous random variable, P(X≤a)=P(X<a)P(X \leq a) = P(X < a), and endpoints never matter — which is a relief, because it means you never have to fret about << versus ≤\leq. For a discrete variable that distinction is critical, so keep track of which kind you are dealing with.

Note

Definition
The cumulative distribution function of a continuous random variable is F(x)=P(X≤x)=∫−∞xf(t) dtF(x) = P(X\leq x) = \int_{-\infty}^{x}f(t)\,dt.

Note

Theorem
F′(x)=f(x)F'(x) = f(x) wherever ff is continuous, and P(a≤X≤b)=F(b)−F(a)P(a\leq X\leq b) = F(b)-F(a).

Proof. The first is the fundamental theorem of calculus applied to F(x)=∫−∞xf(t) dtF(x) = \int_{-\infty}^xf(t)\,dt; the second is additivity of the integral. ■\blacksquare

Example. At a random point during the day we note the time, ignoring the date. Model this by XX uniformly distributed on [0,24)[0,24), so the density is constant: f(x)=cf(x) = c on [0,24)[0,24) and 00 elsewhere. Find cc, the cdf, and the probability the time is between 9am and 5pm.
The density must integrate to 11:

∫024c dx=24c=1  ⟹  c=124.\int_0^{24}c\,dx = 24c = 1 \implies c = \frac{1}{24}.

The cdf is F(x)=∫0x124dt=x24F(x) = \int_0^x\frac{1}{24}dt = \frac{x}{24} for 0≤x<240\leq x < 24 (and 00 below, 11 above). Hence

P(9≤X≤17)=F(17)−F(9)=1724−924=824=13.P(9 \leq X \leq 17) = F(17)-F(9) = \frac{17}{24}-\frac{9}{24} = \frac{8}{24} = \frac13.

Therefore the probability is 13\frac13 — as it must be, since that window is 88 of the 2424 hours.

Example. A continuous random variable has density f(x)=cx2f(x) = cx^2 on [0,3][0,3] and 00 elsewhere. Find cc and P(1≤X≤2)P(1 \leq X \leq 2).
Normalising,

∫03cx2 dx=c[x33]03=9c=1  ⟹  c=19.\int_0^3cx^2\,dx = c\left[\frac{x^3}{3}\right]_0^3 = 9c = 1 \implies c = \frac19.

Then

P(1≤X≤2)=∫12x29 dx=19[x33]12=127(8−1)=727.P(1\leq X\leq 2) = \int_1^2\frac{x^2}{9}\,dx = \frac19\left[\frac{x^3}{3}\right]_1^2 = \frac{1}{27}(8-1) = \frac{7}{27}.

Therefore c=19c = \frac19 and the probability is 727≈0.259\frac{7}{27} \approx 0.259. Finding the normalising constant first is always step one; a density that does not integrate to 11 is not a density.

9.5.1 The mean and variance of a continuous random variable#

Every formula from the discrete case carries over with sums replaced by integrals and pkp_k replaced by f(x) dxf(x)\,dx.

Note

Definition
For a continuous random variable XX with density ff,

E(X)=∫−∞∞xf(x) dx,Var⁡(X)=E(X2)−(E(X))2=∫−∞∞x2f(x) dx−(E(X))2.E(X) = \int_{-\infty}^{\infty}xf(x)\,dx, \qquad \operatorname{Var}(X) = E(X^2)-\big(E(X)\big)^2 = \int_{-\infty}^{\infty}x^2f(x)\,dx - \big(E(X)\big)^2.

Note

Theorem
If XX is continuous with density ff and gg is a real function, then E(g(X))=∫−∞∞g(x)f(x) dxE\big(g(X)\big) = \int_{-\infty}^{\infty}g(x)f(x)\,dx. Moreover E(aX+b)=aE(X)+bE(aX+b) = aE(X)+b and Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX+b)=a^2\operatorname{Var}(X), exactly as in the discrete case.

Example. Find the mean and variance of the uniform distribution on [a,b][a,b].
The density is f(x)=1b−af(x) = \frac{1}{b-a} on [a,b][a,b]. Then

E(X)=∫abxb−a dx=1b−a[x22]ab=b2−a22(b−a)=a+b2,\begin{align*} E(X) &= \int_a^b\frac{x}{b-a}\,dx = \frac{1}{b-a}\left[\frac{x^2}{2}\right]_a^b = \frac{b^2-a^2}{2(b-a)} = \frac{a+b}{2}, \end{align*}

using the difference of two squares. Similarly

E(X2)=1b−a[x33]ab=b3−a33(b−a)=a2+ab+b23,E(X^2) = \frac{1}{b-a}\left[\frac{x^3}{3}\right]_a^b = \frac{b^3-a^3}{3(b-a)} = \frac{a^2+ab+b^2}{3},

so

Var⁡(X)=a2+ab+b23−(a+b)24=4a2+4ab+4b2−3a2−6ab−3b212=a2−2ab+b212=(b−a)212.\begin{align*} \operatorname{Var}(X) &= \frac{a^2+ab+b^2}{3} - \frac{(a+b)^2}{4} \\ &= \frac{4a^2+4ab+4b^2 - 3a^2-6ab-3b^2}{12} \\ &= \frac{a^2-2ab+b^2}{12} \\ &= \frac{(b-a)^2}{12}. \end{align*}

Therefore the uniform distribution on [a,b][a,b] has mean a+b2\frac{a+b}{2} — the midpoint, as symmetry demands — and variance (b−a)212\frac{(b-a)^2}{12}. Applying this to the time-of-day example, the mean is 1212 (noon) and the standard deviation is 2412≈6.93\frac{24}{\sqrt{12}} \approx 6.93 hours.

Note

Theorem (standardisation)
If E(X)=μE(X) = \mu and Var⁡(X)=σ2\operatorname{Var}(X) = \sigma^2, and Z=X−μσZ = \dfrac{X-\mu}{\sigma}, then E(Z)=0E(Z) = 0 and Var⁡(Z)=1\operatorname{Var}(Z) = 1.

Proof. By the linearity results with a=1σa = \frac1\sigma and b=−μσb = -\frac{\mu}{\sigma},

E(Z)=1σE(X)−μσ=μ−μσ=0,Var⁡(Z)=1σ2Var⁡(X)=σ2σ2=1.■E(Z) = \frac{1}{\sigma}E(X) - \frac{\mu}{\sigma} = \frac{\mu-\mu}{\sigma} = 0, \qquad \operatorname{Var}(Z) = \frac{1}{\sigma^2}\operatorname{Var}(X) = \frac{\sigma^2}{\sigma^2} = 1. \qquad \blacksquare

Basically, standardising re-expresses a measurement as "how many standard deviations from the mean", stripping away the units and the scale. This is what makes a single table of normal probabilities usable for every normal distribution, as we see next.

9.6 Special continuous distributions#

9.6.1 The normal distribution#

Note

Definition
A continuous random variable XX has the normal distribution N(μ,σ2)N(\mu,\sigma^2) if its density is

f(x)=1σ2π e−(x−μ)22σ2,x∈R.f(x) = \frac{1}{\sigma\sqrt{2\pi}}\,e^{-\frac{(x-\mu)^2}{2\sigma^2}}, \qquad x \in \mathbb{R}.

The case N(0,1)N(0,1) is the standard normal distribution, whose variable is written ZZ.

Note

Theorem
If X∼N(μ,σ2)X\sim N(\mu,\sigma^2) then E(X)=μE(X) = \mu and Var⁡(X)=σ2\operatorname{Var}(X) = \sigma^2. Moreover Z=X−μσ∼N(0,1)Z = \dfrac{X-\mu}{\sigma}\sim N(0,1).

The density is the familiar symmetric bell curve centred at μ\mu, with σ\sigma controlling its width. It matters enormously in practice because of the central limit theorem (a second-year topic): sums and averages of many independent random quantities are approximately normal regardless of the distribution they came from, which is why the normal turns up in measurement errors, heights, exam marks and almost everything else.

Now a genuine obstacle. To compute P(a≤X≤b)P(a\leq X\leq b) we must integrate the density — but e−x2/2e^{-x^2/2} has no elementary antiderivative, so none of the techniques of Integration Techniques can produce a formula. This is not a failure of ingenuity; it is a theorem. The practical response is to evaluate the integral numerically once and for all (by the power-series method of Section 4.8, in fact) and publish the results as a table of standard normal probabilities. Standardising is what lets one table serve every μ\mu and σ\sigma.

Tables give Φ(z)=P(Z≤z)\Phi(z) = P(Z\leq z). A small excerpt:

zz 0.000.00 1.001.00 1.331.33 1.961.96 2.002.00 3.003.00
Φ(z)\Phi(z) 0.50000.5000 0.84130.8413 0.90820.9082 0.97500.9750 0.97720.9772 0.99870.9987

Two facts make the table go further than it looks. Since the density is symmetric about 00,

Φ(−z)=1−Φ(z),P(Z>z)=1−Φ(z),P(a≤Z≤b)=Φ(b)−Φ(a).\boxed{\Phi(-z) = 1-\Phi(z), \qquad P(Z > z) = 1 - \Phi(z), \qquad P(a\leq Z\leq b) = \Phi(b)-\Phi(a).}

Sketch the bell curve and shade the region you want before reaching for the table. Nearly every error in this section is a region error rather than an arithmetic one, and a five-second sketch prevents it.

Example. Suppose XX is normally distributed with mean μ=20\mu = 20 and standard deviation σ=3\sigma = 3. Find P(X≤24)P(X\leq 24).
Standardising,

P(X≤24)=P(X−μσ≤24−203)=P(Z≤1.33)≈0.9082\begin{align*} P(X\leq 24) &= P\left(\frac{X-\mu}{\sigma}\leq\frac{24-20}{3}\right) \\ &= P(Z \leq 1.33) \\ &\approx 0.9082 \end{align*}

from the table. Therefore P(X≤24)≈0.908P(X\leq 24)\approx 0.908.

Example. Weekly wages of secretaries are normally distributed with mean $800\$800 and standard deviation $50\$50. What is the probability that a secretary earns more than $900\$900, and how many of 20002000 randomly chosen secretaries would you expect to?
Let XX be a weekly wage, so X∼N(800,502)X\sim N(800, 50^2) and Z=X−80050∼N(0,1)Z = \frac{X-800}{50}\sim N(0,1). When X=900X = 900 we have Z=900−80050=2Z = \frac{900-800}{50} = 2, so

P(X>900)=P(Z>2)=1−P(Z≤2)=1−0.9772=0.0228.\begin{align*} P(X > 900) &= P(Z>2) \\ &= 1 - P(Z\leq 2) \\ &= 1 - 0.9772 \\ &= 0.0228. \end{align*}

Out of 20002000 secretaries we would expect 0.0228×2000=45.60.0228\times2000 = 45.6, that is about 4646, to earn more than $900\$900. Therefore the probability is 0.02280.0228 and the expected number is roughly 4646.

Example. With the same wage distribution, find the wage ww exceeded by only the top 10%10\% of secretaries.
This is an inverse problem: we know the probability and want the value. We need P(X>w)=0.10P(X>w) = 0.10, i.e. P(X≤w)=0.90P(X\leq w) = 0.90. Reading the table backwards, Φ(z)=0.90\Phi(z) = 0.90 at about z=1.28z = 1.28. Un-standardising,

w−80050=1.28,w=800+1.28×50=864.\begin{align*} \frac{w-800}{50} &= 1.28, \\ w &= 800 + 1.28\times50 \\ &= 864. \end{align*}

Therefore about 10%10\% of secretaries earn more than $864\$864. For inverse problems you read the table in the opposite direction and then reverse the standardisation; setting up w−μσ=z\frac{w-\mu}{\sigma} = z and solving for ww keeps this straight.

The symmetry and the table together give the rule of thumb worth committing to memory:

P(∣X−μ∣≤σ)≈68%,P(∣X−μ∣≤2σ)≈95%,P(∣X−μ∣≤3σ)≈99.7%.\boxed{P(|X-\mu|\leq\sigma) \approx 68\%, \quad P(|X-\mu|\leq2\sigma)\approx95\%, \quad P(|X-\mu|\leq3\sigma)\approx99.7\%.}

So a normal variable is almost never more than three standard deviations from its mean, which is why a "three sigma" event is treated as evidence that something unusual is going on.

9.6.2 [X] The exponential distribution#

The geometric distribution modelled waiting for a success in discrete trials. Its continuous counterpart models waiting for an event in continuous time.

Note

Definition
A continuous random variable TT has the exponential distribution Exp⁡(λ)\operatorname{Exp}(\lambda), for λ>0\lambda > 0, if its density is

f(t)={λe−λt,t≥0,0,t<0.f(t) = \begin{cases}\lambda e^{-\lambda t}, & t\geq 0, \\ 0, & t<0.\end{cases}

Check that this is a density: f≥0f\geq0 everywhere, and

∫0∞λe−λt dt=[−e−λt]0∞=0−(−1)=1.\int_0^{\infty}\lambda e^{-\lambda t}\,dt = \left[-e^{-\lambda t}\right]_0^{\infty} = 0-(-1) = 1.

Integrating gives the cdf immediately:

F(t)=P(T≤t)=1−e−λt(t≥0),P(T>t)=e−λt.F(t) = P(T\leq t) = 1 - e^{-\lambda t} \quad (t \geq 0), \qquad P(T>t) = e^{-\lambda t}.

Compare this with the geometric tail probability P(X>n)=qnP(X>n) = q^n — the exponential is the continuous limit of the geometric, and e−λte^{-\lambda t} plays exactly the role that qnq^n did.

Note

Theorem
If T∼Exp⁡(λ)T\sim\operatorname{Exp}(\lambda), then E(T)=1λE(T) = \dfrac1\lambda and Var⁡(T)=1λ2\operatorname{Var}(T) = \dfrac{1}{\lambda^2}.

Proof (of the mean). Integrating by parts with u=tu = t and dv=λe−λtdtdv = \lambda e^{-\lambda t}dt,

E(T)=∫0∞t λe−λt dt=[−te−λt]0∞+∫0∞e−λt dt=0+[−1λe−λt]0∞=1λ,\begin{align*} E(T) &= \int_0^{\infty}t\,\lambda e^{-\lambda t}\,dt \\ &= \Big[-te^{-\lambda t}\Big]_0^{\infty} + \int_0^{\infty}e^{-\lambda t}\,dt \\ &= 0 + \left[-\frac{1}{\lambda}e^{-\lambda t}\right]_0^{\infty} \\ &= \frac1\lambda, \end{align*}

where the boundary term vanishes because te−λt→0te^{-\lambda t}\to0 as t→∞t\to\infty (the exponential beats the linear term). ■\blacksquare

So λ\lambda is a rate — events per unit time — and 1λ\frac1\lambda is the mean waiting time between them, mirroring the geometric mean 1p\frac1p exactly.

Note

Theorem (memorylessness)
If T∼Exp⁡(λ)T\sim\operatorname{Exp}(\lambda) then for all s,t≥0s,t\geq0,  P(T>s+t∣T>s)=P(T>t)\ P(T > s+t \mid T > s) = P(T>t).

Proof. Since {T>s+t}⊆{T>s}\{T>s+t\}\subseteq\{T>s\}, the intersection of the two events is just {T>s+t}\{T>s+t\}, so

P(T>s+t∣T>s)=P(T>s+t)P(T>s)=e−λ(s+t)e−λs=e−λt=P(T>t).■\begin{align*} P(T>s+t\mid T>s) &= \frac{P(T>s+t)}{P(T>s)} \\ &= \frac{e^{-\lambda(s+t)}}{e^{-\lambda s}} \\ &= e^{-\lambda t} \\ &= P(T>t). \qquad \blacksquare \end{align*}

Basically, an exponentially distributed component does not age: given that it has already survived ss hours, its chance of surviving another tt is the same as a brand-new one's. This is exactly the continuous version of the fact that a coin has no memory of previous tosses, and the exponential is the only continuous distribution with this property.

It is also why the exponential is often the wrong model for things that genuinely wear out. A light bulb that has been burning for a year really is more likely to fail soon than a new one, so its lifetime is not exponential.

Example. The time in years between claims on a certain insurance policy is exponentially distributed with mean 44 years. Find the probability that the next claim comes within 22 years, and the probability that a policy with no claim for 33 years goes a further 22 without one.
Since the mean is 1λ=4\frac1\lambda = 4, we have λ=14\lambda = \frac14. Then

P(T≤2)=1−e−2/4=1−e−0.5≈1−0.6065=0.3935.P(T\leq 2) = 1 - e^{-2/4} = 1 - e^{-0.5} \approx 1 - 0.6065 = 0.3935.

For the second part, memorylessness makes the three claim-free years irrelevant:

P(T>5∣T>3)=P(T>2)=e−0.5≈0.6065.P(T > 5\mid T>3) = P(T>2) = e^{-0.5}\approx 0.6065.

Therefore there is about a 39.4%39.4\% chance of a claim within two years, and a 60.7%60.7\% chance that a three-year-clean policy stays clean for two more — exactly the same as for a brand-new policy.

To summarise the chapter: probability is a measure on subsets of a sample space, and every rule for combining probabilities comes from partitioning a set into disjoint pieces and adding. Conditioning restricts attention to a sub-space and rescales, which gives the law of total probability and Bayes' theorem — and Bayes is where intuition fails most badly, because base rates matter as much as accuracy. Random variables attach numbers to outcomes so we can average them, giving the mean and variance, and a few standard distributions cover most situations: the binomial for counting successes in a fixed number of trials, the geometric for waiting for the first success, the normal for anything built from many small independent contributions, and the exponential for waiting times in continuous time. The discrete and continuous cases are the same theory throughout, with ∑kpk\sum_kp_k replaced by ∫f(x) dx\int f(x)\,dx; the one genuine difference is that a continuous variable assigns probability zero to every individual value, so probability lives in areas rather than in points.