Structure preview

Probability and statistics: a course in 20 chapters

From random experiments to Bayes, from distributions to confidence intervals: one course with an index and twenty expandable chapters.

Articles /probability-statistics-twenty-chapter-course
Probability and statistics: a course in 20 chapters

105 min

Probability begins with random experiments and leads to a critical reading of data. This course follows each step: events, counting, conditional probability, random variables, distributions and statistical inference. All twenty chapters belong to one article: choose a title in the index to open its box, or read them in order.

Chapter 1 – Random experiments, events and sample space

1. Why probability arises

In many mathematical problems, by knowing the initial conditions, we can determine the result with certainty.

For example:

\[ 7+5=12 \]

or, if a car travels at a constant speed of 80 km/h for 2 hours, it travels:

\[ 80\cdot2=160\text{ km} \]

However, there are situations in which we are not able to predict with certainty what will happen.

If we flip a coin, we don't know in advance whether it will come up heads or tails.

If we roll a die, we know that the result will be one of the numbers:

\[ 1,2,3,4,5,6 \]

but we don't know which one.

Probability was created precisely to mathematically study situations of this type.

It doesn't necessarily try to predict the individual outcome with certainty, but it tries to measure how plausible it is for a given outcome to occur.


2. Deterministic experiment and random experiment

Deterministic experiment

An experiment is said to be deterministic when, under the same conditions, it always produces the same result.

For example:

  • calculate \(5+7\);
  • dropping an object in ideal conditions;
  • calculate the area of ​​a square with a side of 4 cm.

The outcome is completely determined by the initial conditions.

Random experiment

A random experiment, also called random experiment, is an experiment for which we know the possible outcomes, but we cannot know for sure which of them will occur before carrying out the experiment.

Examples:

  • flip a coin;
  • roll a dice;
  • extract a card from a deck;
  • randomly choose a person from a population;
  • observe the number of customers who enter a store in an hour;
  • measure the time it takes for an electronic component to fail.

The word random does not mean we know nothing.

On the contrary, we can often know perfectly the set of possible outcomes.

What we don't know is what outcome will happen.


3. Result of an experiment

Each possible outcome of a random experiment is called an outcome.

Consider rolling a die.

The possible outcomes are:

\[ 1,\ 2,\ 3,\ 4,\ 5,\ 6 \]

If the die shows 4, we say that the outcome has occurred:

\[ 4 \]

When tossing a coin the outcomes are:

\[ T,\ C \]

where:

  • \(T\) = head;
  • \(C\) = tails.

4. The sample space

The set of all possible outcomes of a random experiment is called sample space.

It is generally indicated with the Greek letter:

\[ \Omega \]

or, in some texts, with \(S\).

Example 1 – Tossing a coin

\[ \Omega=\{T,C\} \]

Example 2 – Rolling a dice

\[ \Omega=\{1,2,3,4,5,6\} \]

Example 3 – Drawing a card

Suppose we draw a card from a normal Italian deck of 40 cards.

The sample space contains all 40 cards.

We can write:

\[ \Omega=\{\text{tutte le 40 carte del mazzo}\} \]


5. The sample space depends on what we observe

Consider rolling a die.

If we are interested in the number obtained:

\[ \Omega=\{1,2,3,4,5,6\} \]

If, however, we are only interested in knowing whether the result is even or odd:

\[ \Omega=\{\text{pari},\text{dispari}\} \]

The physical experiment is the same.

The information we want to observe has changed.


6. Toss of two coins

We flip two coins at the same time.

The correct sample space is:

\[ \Omega= \{TT,TC,CT,CC\} \]

We have four possible outcomes.

It is important to distinguish:

\[ TC \]

from:

\[ CT \]

because they represent two different results if we distinguish the first coin from the second.


7. Roll two dice

We roll two dice.

The result can be represented by an ordered pair:

\[ (i,j) \]

where:

  • \(i\) is the result of the first die;
  • \(j\) is the result of the second.

For example:

\[ (2,5) \]

means:

  • first die: 2;
  • second die: 5.

The sample space contains:

\[ 6\cdot6=36 \]

outcomes.

We can write:

\[ \Omega= \{(i,j):i,j\in\{1,2,3,4,5,6\}\} \]


8. Event

An event is a set of one or more sample space outcomes.

Mathematically:

\[ A\subseteq\Omega \]

that is, an event is a subset of the sample space.

Example

Let's roll a dice.

\[ \Omega=\{1,2,3,4,5,6\} \]

Let's consider the event:

an even number comes out.

Then:

\[ A=\{2,4,6\} \]

The \(A\) event occurs when the roll result belongs to the set \(A\).


9. Elementary event

An event consisting of a single outcome is called an elementary event.

For example:

\[ A=\{3\} \]

represents:

exits 3.


10. Compound event

An event consisting of multiple outcomes is called a compound event.

For example:

\[ A=\{2,4,6\} \]

represents:

an even number comes out.


11. Certain event

An event that coincides with the entire sample space is called a certain event.

When rolling a die:

a number between 1 and 6 comes out

corresponds to:

\[ \Omega \]

Its probability will be:

\[ P(\Omega)=1 \]


12. Impossible event

An event that contains no outcome is called an impossible event.

It is represented by the empty set:

\[ \varnothing \]

For example:

rolling a die comes out 8

it's impossible.

So:

\[ P(\varnothing)=0 \]


13. Event and phrase

Let's consider:

\[ \Omega=\{1,2,3,4,5,6\} \]

Event:

a number greater than 4 comes out.

\[ A=\{5,6\} \]

Event:

an odd number comes out.

\[ B=\{1,3,5\} \]

Event:

a multiple of 3 comes out.

\[ C=\{3,6\} \]

Correctly translating language into sets is critical.


14. Example with two dice

Let's roll two dice and consider:

the sum of the results is 7.

Favorable outcomes are:

\[ (1,6),(2,5),(3,4),(4,3),(5,2),(6,1) \]

So:

\[ A= \{(1,6),(2,5),(3,4),(4,3),(5,2),(6,1)\} \]


15. Finite and infinite sample spaces

Finite sample space

For a dice:

\[ \Omega=\{1,2,3,4,5,6\} \]

and:

\[ |\Omega|=6 \]

Discrete infinite sample space

We flip a coin until heads appear.

The number of launches can be:

\[ 1,2,3,4,\ldots \]

therefore:

\[ \Omega=\{1,2,3,4,\ldots\} \]

Continuous sample space

Let's consider the lifespan of a light bulb.

Ideally:

\[ \Omega=[0,+\infty) \]


16. Outcome and event are not the same thing

Suppose we roll a die and get:

\[ 4 \]

The number 4 is an outcome.

The whole:

\[ \{2,4,6\} \]

instead it is the event:

an even number comes out.

Because:

\[ 4\in\{2,4,6\} \]

the event occurred.


17. Complete example

Let's consider:

\[ \Omega= \{1,2,3,4,5,6,7,8,9,10\} \]

Let's define:

\[ A=\{\text{numero pari}\} \]

\[ B=\{\text{numero maggiore di 7}\} \]

\[ C=\{\text{multiplo di 5}\} \]

Then:

\[ A=\{2,4,6,8,10\} \]

\[ B=\{8,9,10\} \]

\[ C=\{5,10\} \]

If it is extracted:

\[ 10 \]

occur simultaneously:

\[ A,\quad B,\quad C \]


18. Conceptual diagram

\[ \boxed{\text{Esperimento casuale}} \]

produces one of the possible:

\[ \boxed{\text{Esiti}} \]

The set of all outcomes forms:

\[ \boxed{\Omega=\text{spazio campionario}} \]

A set of outcomes constitutes:

\[ \boxed{\text{Evento}} \]

Mathematically:

\[ \boxed{A\subseteq\Omega} \]


19. Frequent errors

Error 1

Confusing event and outcome.

Error 2

Forget the order.

When throwing two dice:

\[ (2,5)\neq(5,2) \]

Error 3

Constructing the sample space poorly.

For two coins:

\[ \{TT,TC,CT,CC\} \]

Error 4

Thinking that random means “without rules”.

A random phenomenon can be studied mathematically.


20. Exercises

Exercise 1

A dice is rolled.

Write the sample space.

Solution

\[ \Omega=\{1,2,3,4,5,6\} \]

Exercise 2

Write the event:

a number less than 4 comes out.

Solution

\[ A=\{1,2,3\} \]

Exercise 3

A number from 1 to 12 is drawn.

Event:

multiple of 4.

Solution

\[ A=\{4,8,12\} \]

Exercise 4

Two coins are tossed.

Solution

\[ \Omega=\{TT,TC,CT,CC\} \]

Exercise 5

Two dice are rolled.

Event:

the sum is 4.

Solution

\[ A=\{(1,3),(2,2),(3,1)\} \]


21. More thoughtful exercise

We roll two dice.

Event:

at least one of the two dice shows 6.

The outcomes are:

\[ (6,1),(6,2),(6,3),(6,4),(6,5),(6,6) \]

and:

\[ (1,6),(2,6),(3,6),(4,6),(5,6) \]

In total:

\[ 6+6-1=11 \]

outcomes.


22. Final idea

The starting point of every probability problem is:

  1. What is the experiment?
  2. What are the possible outcomes?
  3. What is the sample space?
  4. which event are we interested in?
  5. What outcomes belong to the event?

The fundamental relationship is:

\[ \boxed{ \text{evento}=\text{sottoinsieme dello spazio campionario} } \]


Chapter 2 – Definition of probability

1. From sample space to probability

We want to measure how likely an event is to occur.

We introduce:

\[ P(A) \]

with:

\[ 0\leq P(A)\leq1 \]

The closer \(P(A)\) is to 1, the more likely the event is.

The closer it is to 0, the less likely it is.


2. Probability 0 and probability 1

Impossible event:

\[ P(A)=0 \]

Certain event:

\[ P(A)=1 \]

Here we are thinking about a finite sample space with positive probability outcomes. In general, especially in continuous models, an event can have probability 0 without being impossible, or probability 1 without coinciding with the entire sample space.


3. Classical definition

If all outcomes are equally probable:

\[ \boxed{ P(A)= \frac{\text{casi favorevoli}} {\text{casi possibili}} } \]

In symbols:

\[ P(A)=\frac{|A|}{|\Omega|} \]


4. Example: dice

Event:

even number.

\[ A=\{2,4,6\} \]

So:

\[ P(A)=\frac36=\frac12 \]


5. Number greater than 4

\[ B=\{5,6\} \]

So:

\[ P(B)=\frac26=\frac13 \]


6. Elementary event

Event:

exits 3.

\[ P(3)=\frac16 \]


7. The key word: equiprobable

The formula:

\[ P(A)=\frac{\text{casi favorevoli}}{\text{casi possibili}} \]

holds when the outcomes are equally probable.


8. Tricked coin

Suppose:

\[ P(T)=0,7 \]

\[ P(C)=0,3 \]

We cannot write:

\[ P(T)=\frac12 \]

just because we have two outcomes.

The two outcomes are not equally probable.


9. Uneven urn

An urn contains:

  • 9 red balls;
  • 1 blue.

The probability of red is:

\[ P(R)=\frac9{10} \]

not:

\[ \frac12 \]


10. Construct the sample space correctly

We can distinguish:

\[ R_1,R_2,\ldots,R_9,B \]

We now have 10 equally probable outcomes.

The red event contains 9 outcomes:

\[ P(R)=\frac9{10} \]


11. Correct procedure

  1. build the sample space;
  2. verify equiprobability;
  3. count the possible cases;
  4. count the favorable cases;
  5. make the report.

12. Two dice

We roll two dice.

The outcomes are:

\[ 36 \]

Event:

sum 7.

Favorable cases:

\[ (1,6),(2,5),(3,4),(4,3),(5,2),(6,1) \]

So:

\[ P(\text{somma}=7) = \frac6{36} = \frac16 \]


13. Common mistake: the sums are not equiprobable

Possible sums range from:

\[ 2 \]

to:

\[ 12 \]

but they do not all have the same probability.

For example:

\[ P(\text{somma}=2)=\frac1{36} \]

while:

\[ P(\text{somma}=7)=\frac6{36} \]


14. Distribution of sums

\[ 2\rightarrow1 \]

\[ 3\rightarrow2 \]

\[ 4\rightarrow3 \]

\[ 5\rightarrow4 \]

\[ 6\rightarrow5 \]

\[ 7\rightarrow6 \]

\[ 8\rightarrow5 \]

\[ 9\rightarrow4 \]

\[ 10\rightarrow3 \]

\[ 11\rightarrow2 \]

\[ 12\rightarrow1 \]


15. Fraction, decimal and percentage

\[ \frac14=0,25=25\% \]

They are three different ways of expressing the same probability.


16. Frequentist probability

Suppose we test:

\[ 10000 \]

light bulbs.

If:

\[ 8200 \]

last more than 1000 hours:

\[ \frac{8200}{10000}=0,82 \]

We can estimate:

\[ P(\text{durata}>1000)\approx0,82 \]


17. Relative frequency

If a \(A\) event occurs \(n_A\) times in \(n\) trials:

\[ \boxed{ f_A=\frac{n_A}{n} } \]


18. Probability and frequency are not the same thing

For a balanced currency:

\[ P(T)=0,5 \]

But in 10 throws we could get 6 heads:

\[ f_T=0,6 \]

The probability is theoretical.

Frequency is observed.


19. Frequency stabilization

With many launches:

\[ f_T \]

tends to get closer to:

\[ 0,5 \]

This will be formalized by the Law of Large Numbers.


20. Axiomatic definition

Probability satisfies three fundamental axioms.

First

\[ P(A)\geq0 \]

Second

\[ P(\Omega)=1 \]

Third

If \(A\) and \(B\) are incompatible:

\[ A\cap B=\varnothing \]

then:

\[ P(A\cup B)=P(A)+P(B) \]


21. Three ways of looking at probability

Classic

\[ P(A)=\frac{|A|}{|\Omega|} \]

when the cases are equiprobable.

Frequentist

\[ P(A)\approx\frac{n_A}{n} \]

for many tests.

Axiomatic

Probability is a function that satisfies precise mathematical properties.


22. Subjective probability

Probability can also represent a degree of confidence based on available information.

This interpretation is important in:

  • economy;
  • medicine;
  • forecasts;
  • artificial intelligence;
  • decision theory.

23. Probable does not mean certain

If:

\[ P(A)=0,9 \]

the event is very probable, but may not occur.

If:

\[ P(A)=0,01 \]

It's unlikely, but it can happen.


24. Example with urn

Urn with:

  • 5 red ones;
  • 3 blue;
  • 2 greens.

Total:

\[ 10 \]

So:

\[ P(R)=\frac5{10}=\frac12 \]

\[ P(B)=\frac3{10} \]

\[ P(V)=\frac2{10}=\frac15 \]

and:

\[ \frac5{10}+\frac3{10}+\frac2{10}=1 \]


25. Deck of cards

From a deck of 52 cards:

\[ P(\text{asso})=\frac4{52}=\frac1{13} \]

\[ P(\text{cuori})=\frac{13}{52}=\frac14 \]


26. Figures

The figures are:

  • Jack;
  • Woman;
  • King.

They are:

\[ 3\cdot4=12 \]

So:

\[ P(\text{figura}) = \frac{12}{52} = \frac3{13} \]


27. Counterintuitive problem: at least one head

Let's flip two coins.

\[ \Omega=\{TT,TC,CT,CC\} \]

Event:

at least one head.

\[ A=\{TT,TC,CT\} \]

So:

\[ P(A)=\frac34 \]


28. Complementary method

The only headless case is:

\[ CC \]

So:

\[ P(\text{almeno una testa}) = 1-\frac14 = \frac34 \]


29. Exercises

Exercise 1

Die: probability of odd number.

\[ P=\frac36=\frac12 \]

Exercise 2

Number from 1 to 20: probability of a multiple of 5.

Multiple:

\[ 5,10,15,20 \]

So:

\[ P=\frac4{20}=\frac15 \]

Exercise 3

Two dice: sum 5.

Cases:

\[ (1,4),(2,3),(3,2),(4,1) \]

So:

\[ P=\frac4{36}=\frac19 \]

Exercise 4

Urn with 7 white and 3 black.

\[ P(N)=\frac3{10} \]

Exercise 5

Two coins - exactly one head.

\[ \{TC,CT\} \]

So:

\[ P=\frac24=\frac12 \]

Exercise 6

Two coins: at least one cross.

Complementary:

\[ TT \]

So:

\[ P=1-\frac14=\frac34 \]


30. Conceptual question

If:

\[ P(A)=0,2 \]

it does not mean that the event certainly occurs once every 5 trials.

It means that over many repetitions the relative frequency tends to approach:

\[ 0,2 \]


31. Gambler's fallacy

If no heads have ever come up for 10 tosses, the probability of heads on the next toss, for a balanced coin, remains:

\[ \frac12 \]

The currency does not “have to recover”.


32. Summary

Probability is a number:

\[ 0\leq P(A)\leq1 \]

The classic definition is:

\[ \boxed{ P(A)=\frac{|A|}{|\Omega|} } \]

when the outcomes are equally probable.

The fundamental distinction is:

\[ \boxed{ \text{probabilità}\neq\text{frequenza osservata} } \]

although, by increasing the number of trials, the frequency often tends to approach probability.

Chapter 3 – Operations between events

1. Events as sets

We have seen that an event is a subset of the sample space.

If:

\[ \Omega \]

is the sample space and \(A\) and \(B\) are two events, we can use normal set operations.

The most important are:

  • union;
  • intersection;
  • complementary;
  • difference;
  • incompatibility.

2. Union of two events

The union of two events \(A\) and \(B\) is indicated by:

\[ A\cup B \]

and it means:

\(A\) or \(B\), or both, occurs.

Example

Let's roll a dice.

\[ \Omega=\{1,2,3,4,5,6\} \]

Either:

\[ A=\{\text{numero pari}\}=\{2,4,6\} \]

\[ B=\{\text{numero maggiore di 3}\}=\{4,5,6\} \]

Then:

\[ A\cup B=\{2,4,5,6\} \]


3. Intersection of two events

The intersection is indicated with:

\[ A\cap B \]

and it means:

\(A\) and \(B\) occur simultaneously.

With previous events:

\[ A\cap B=\{4,6\} \]


4. “Or” and “and”

In probability:

\(A\) or \(B\)

translates:

\[ A\cup B \]

while:

\(A\) and \(B\)

translates:

\[ A\cap B \]


5. Pay attention to the meaning of “or”

The union:

\[ A\cup B \]

normally has an inclusive meaning:

\(A\), or \(B\), or both.


6. Complementary of an event

The complement of \(A\) is the event:

\(A\) does not occur.

It is indicated with:

\[ A^c \]

or:

\[ \overline A \]

Example

If:

\[ A=\{2,4,6\} \]

then:

\[ A^c=\{1,3,5\} \]


7. Properties of the complementary

An event and its complement satisfy:

\[ A\cup A^c=\Omega \]

and:

\[ A\cap A^c=\varnothing \]


8. Complementary formula

Because:

\[ P(\Omega)=1 \]

we have:

\[ P(A)+P(A^c)=1 \]

therefore:

\[ \boxed{ P(A^c)=1-P(A) } \]


9. Example

If:

\[ P(A)=0,7 \]

then:

\[ P(A^c)=0,3 \]


10. At least one and complementary

We flip three coins.

What is the probability of getting at least one head?

It's easier to use the complementary.

The opposite event is:

no head.

This means:

\[ CCC \]

So:

\[ P(\text{nessuna testa}) = \left(\frac12\right)^3 = \frac18 \]

Therefore:

\[ P(\text{almeno una testa}) = 1-\frac18 = \frac78 \]


11. Difference between events

The difference:

\[ A\setminus B \]

means:

\(A\) occurs, but not \(B\).

Equivalently:

\[ A\setminus B=A\cap B^c \]

Example

If:

\[ A=\{2,4,6\} \]

\[ B=\{4,5,6\} \]

then:

\[ A\setminus B=\{2\} \]


12. Incompatible events

Two events are incompatible when they cannot occur at the same time.

Mathematically:

\[ A\cap B=\varnothing \]

Example

Rolling a dice:

\[ A=\{\text{numero pari}\} \]

\[ B=\{\text{numero dispari}\} \]

They are incompatible.


13. Compatible events

They are compatible if:

\[ A\cap B\neq\varnothing \]

For example:

\[ A=\{2,4,6\} \]

\[ B=\{4,5,6\} \]

have:

\[ A\cap B=\{4,6\} \]


14. Probability of union for incompatible events

If:

\[ A\cap B=\varnothing \]

then:

\[ \boxed{ P(A\cup B)=P(A)+P(B) } \]


15. Example

Dice:

\[ A=\{1\} \]

\[ B=\{6\} \]

So:

\[ P(A\cup B) = \frac16+\frac16 = \frac13 \]


16. Because we can't always add up

Suppose:

\[ A=\{2,4,6\} \]

\[ B=\{4,5,6\} \]

If we did:

\[ P(A)+P(B) = \frac36+\frac36 = 1 \]

we would be wrong because:

\[ 4,6 \]

they were counted twice.


17. General formula of the union

The correct formula is:

\[ \boxed{ P(A\cup B) = P(A)+P(B)-P(A\cap B) } \]


18. Example

\[ P(A)=\frac36 \]

\[ P(B)=\frac36 \]

\[ P(A\cap B)=\frac26 \]

So:

\[ P(A\cup B) = \frac36+\frac36-\frac26 = \frac46 = \frac23 \]


19. Principle of inclusion-exclusion

The formula:

\[ P(A\cup B) = P(A)+P(B)-P(A\cap B) \]

it is a case of the inclusion-exclusion principle.

The idea is:

  1. we count \(A\);
  2. we count \(B\);
  3. the intersection was counted twice;
  4. we subtract it once.

20. Venn Diagrams

In Venn diagrams:

  • the rectangle represents \(\Omega\);
  • the internal regions represent the events;
  • the overlap represents \(A\cap B\);
  • the set of the two regions represents \(A\cup B\);
  • the part outside \(A\) represents \(A^c\).

21. Laws of De Morgan

\[ \boxed{ (A\cup B)^c=A^c\cap B^c } \]

and:

\[ \boxed{ (A\cap B)^c=A^c\cup B^c } \]


22. Interpretation

The first means:

neither \(A\) nor \(B\) occurs.

The second means:

\(A\) and \(B\) do not occur at the same time.


23. “At least one”

The expression:

at least one

often leads to union.

The opposite is:

none.

So often:

\[ P(\text{almeno uno}) = 1-P(\text{nessuno}) \]


24. “Exactly one”

It does not coincide with "at least one".

With two coins:

\[ \Omega=\{TT,TC,CT,CC\} \]

“at least one head”:

\[ \{TT,TC,CT\} \]

“exactly one head”:

\[ \{TC,CT\} \]

So:

\[ P(\text{almeno una testa})=\frac34 \]

\[ P(\text{esattamente una testa})=\frac12 \]


25. “At most one”

It means:

\[ 0\text{ oppure }1 \]

For example:

at most one head

includes:

  • no head;
  • exactly one head.

26. Example with cards

From a deck of 52 cards we draw a card.

Either:

\[ A=\{\text{asso}\} \]

\[ B=\{\text{cuori}\} \]

We have:

\[ P(A)=\frac4{52} \]

\[ P(B)=\frac{13}{52} \]

\[ P(A\cap B)=\frac1{52} \]

So:

\[ P(A\cup B) = \frac4{52} + \frac{13}{52} - \frac1{52} \]

\[ = \frac{16}{52} = \frac4{13} \]


27. Probability of three events

For three events:

\[ P(A\cup B\cup C) \]

is valid:

\[ P(A)+P(B)+P(C) \]

\[ -P(A\cap B) -P(A\cap C) -P(B\cap C) \]

\[ +P(A\cap B\cap C) \]


28. Counting example

In a classroom:

  • 18 study English;
  • 12 French;
  • 7 both.

How many study at least one language?

\[ 18+12-7=23 \]


29. Exercises

Exercise 1

\[ A=\{1,2,3\} \]

\[ B=\{3,4,5\} \]

Find:

\[ A\cup B \]

and:

\[ A\cap B \]

Solution

\[ A\cup B=\{1,2,3,4,5\} \]

\[ A\cap B=\{3\} \]

Exercise 2

Dice:

\[ A=\{\text{numero maggiore di 2}\} \]

Then:

\[ A=\{3,4,5,6\} \]

\[ A^c=\{1,2\} \]

Exercise 3

If:

\[ P(A)=0,35 \]

then:

\[ P(A^c)=0,65 \]

Exercise 4

If:

\[ P(A)=0,4 \]

\[ P(B)=0,3 \]

and they are incompatible:

\[ P(A\cup B)=0,7 \]

Exercise 5

If:

\[ P(A)=0,6 \]

\[ P(B)=0,5 \]

\[ P(A\cap B)=0,2 \]

then:

\[ P(A\cup B)=0,9 \]


30. Summary of Chapter 3

\[ \boxed{ A\cup B } \]

means:

\(A\) or \(B\).

\[ \boxed{ A\cap B } \]

means:

\(A\) and \(B\).

\[ \boxed{ A^c } \]

means:

not \(A\).

Complementary formula:

\[ \boxed{ P(A^c)=1-P(A) } \]

Union formula:

\[ \boxed{ P(A\cup B) = P(A)+P(B)-P(A\cap B) } \]


Chapter 4 – Combinatorial calculus applied to probability

1. Why combinatorics are needed

In classical probability:

\[ P(A)=\frac{\text{casi favorevoli}}{\text{casi possibili}} \]

Often the real problem is counting these cases correctly.

For example:

  • how many passwords can we form?
  • how many groups can we choose?
  • how many arrangements are possible?
  • how many hands of cards are there?

Combinatorial calculus allows you to answer without listing all the cases.


2. Fundamental principle of counting

If a procedure has multiple phases and:

  • the first can be carried out in \(n_1\) ways;
  • the second in \(n_2\);
  • the third in \(n_3\);

then the total number of possibilities is:

\[ \boxed{ n_1n_2n_3 } \]


3. Example

4 shirts and 3 trousers.

The combinations are:

\[ 4\cdot3=12 \]


4. Menu

3 first courses, 4 second courses, 2 desserts.

\[ 3\cdot4\cdot2=24 \]

menu.


5. Repeat password

4-digit password.

Each location has:

\[ 10 \]

possibility.

So:

\[ 10^4=10000 \]


6. Password without repetition

If the digits cannot repeat:

\[ 10\cdot9\cdot8\cdot7 = 5040 \]


7. Factorial

For a positive integer:

\[ \boxed{ n!=n(n-1)(n-2)\cdots2\cdot1 } \]

Examples:

\[ 5!=120 \]

\[ 4!=24 \]

\[ 3!=6 \]

By convention:

\[ \boxed{0!=1} \]


8. Simple permutations

If we have \(n\) distinct objects and we want to sort them all:

\[ \boxed{ P_n=n! } \]


9. Example

With:

\[ A,B,C \]

we have:

\[ ABC,\ ACB,\ BAC,\ BCA,\ CAB,\ CBA \]

So:

\[ 3!=6 \]


10. People in line

Five people:

\[ 5!=120 \]

orders.


11. When to use permutations

We use permutations when:

  1. we use all the elements;
  2. the elements are distinct;
  3. order matters.

12. Simple provisions

If we have \(n\) elements and we choose \(k\), with important order and without repetition:

\[ \boxed{ D_{n,k} = \frac{n!}{(n-k)!} } \]


13. Example

Among 10 people we choose:

  • president;
  • vice president;
  • secretary.

Order matters.

\[ D_{10,3} = 10\cdot9\cdot8 = 720 \]


14. Why order matters

Marco president, Luca vice president

is different from:

Luca president, Marco vice president.


15. Provisions with repetition

If we have \(n\) symbols and we construct sequences of length \(k\), with repetition:

\[ \boxed{ n^k } \]


16. Example

5-digit password:

\[ 10^5=100000 \]


17. Simple combinations

If we choose \(k\) items between \(n\) and the order does not count:

\[ \boxed{ \binom nk = \frac{n!}{k!(n-k)!} } \]


18. Example

We choose a commission of 3 people out of 10.

\[ \binom{10}{3} = \frac{10!}{3!7!} = 120 \]


19. Why do we divide by \(k!\)

The provisions would count each group:

\[ k! \]

times.

So:

\[ \binom nk = \frac{D_{n,k}}{k!} \]


20. Decisive question

We need to ask ourselves:

do I get a different result by changing the order?

If yes:

\[ \text{disposizioni o permutazioni} \]

If not:

\[ \text{combinazioni} \]


21. Conceptual scheme

All items, important order

\[ n! \]

\(k\) items on \(n\), important order

\[ \frac{n!}{(n-k)!} \]

\(k\) items on \(n\), order not important

\[ \binom nk \]

Repetitions allowed and important order

\[ n^k \]


22. Permutations with repeated elements

Word:

\[ MAMMA \]

We have:

  • 5 letters;
  • \(M\) repeated 3 times;
  • \(A\) repeated 2 times.

The distinct anagrams are:

\[ \frac{5!}{3!2!} = 10 \]


23. General formula

If:

\[ n_1+n_2+\cdots+n_r=n \]

then:

\[ \boxed{ \frac{n!}{n_1!n_2!\cdots n_r!} } \]


24. Example: MATHEMATICS

The word has 10 letters.

Repetitions:

  • \(A\): 3;
  • \(M\): 2;
  • \(T\): 2.

So:

\[ \frac{10!}{3!2!2!} \]


25. Cards and melds

From 52 cards we choose 5.

If the order is not of interest:

\[ \boxed{ \binom{52}{5} } \]


26. Chance of 5 hearts

There are 13 heart cards.

Favorable cases:

\[ \binom{13}{5} \]

Possible cases:

\[ \binom{52}{5} \]

So:

\[ \boxed{ P= \frac{\binom{13}{5}}{\binom{52}{5}} } \]


27. Exactly 2 aces

In a hand of 5 cards:

  • we choose 2 aces among 4;
  • we choose 3 non-aces among 48.

So:

\[ \boxed{ P= \frac{ \binom42\binom{48}{3} }{ \binom{52}{5} } } \]


28. Lottery

If you draw 6 numbers from 90 and the order does not count:

\[ \binom{90}{6} \]

possible sextuplets.

The probability of a specific sextuplet is:

\[ \frac1{\binom{90}{6}} \]


29. Common mistake

For a 5 card hand we don't have to use dispositions if the order doesn't matter.

The hand:

\[ A,B,C,D,E \]

is the same as:

\[ E,D,C,B,A \]


30. When the order counts instead

If three numbers form a code:

\[ 2,5,7 \]

is different from:

\[ 7,5,2 \]

So we use dispositions.


31. Code with letters and numbers

Code consisting of:

  • 2 letters;
  • 3 digits.

With repetition:

\[ 26^2\cdot10^3 \]


32. Password that does not start with zero

4-digit numeric password.

First digit:

\[ 9 \]

possibility.

Three more:

\[ 10 \]

each.

Total:

\[ 9\cdot10^3=9000 \]


33. Numbers with 4 different digits

First digit:

\[ 9 \]

Second:

\[ 9 \]

Third:

\[ 8 \]

Fourth:

\[ 7 \]

Total:

\[ 9\cdot9\cdot8\cdot7 = 4536 \]


34. Properties of combinations

\[ \boxed{ \binom nk=\binom n{n-k} } \]

Because choosing \(k\) elements is equivalent to choosing which \(n-k\) to leave out.


35. Pascal relation

\[ \boxed{ \binom nk = \binom{n-1}{k} + \binom{n-1}{k-1} } \]


36. Pascal's triangle

\[ 1 \]

\[ 1\quad1 \]

\[ 1\quad2\quad1 \]

\[ 1\quad3\quad3\quad1 \]

\[ 1\quad4\quad6\quad4\quad1 \]

Each internal number is the sum of the two above.


37. Connection with Newton's binomial

\[ (a+b)^4 = a^4+4a^3b+6a^2b^2+4ab^3+b^4 \]

The coefficients:

\[ 1,4,6,4,1 \]

they are binomial coefficients.


38. Complete problem

A class has:

  • 7 girls;
  • 5 guys.

4 students are chosen.

Probability of exactly 3 girls:

\[ \boxed{ P= \frac{ \binom73\binom51 }{ \binom{12}{4} } } \]


39. Another complete problem

Urn with:

  • 6 red ones;
  • 4 blue.

3 balls are drawn.

Probability of exactly 2 reds:

\[ P= \frac{ \binom62\binom41 }{ \binom{10}{3} } \]

Calculating:

\[ \binom62=15 \]

\[ \binom41=4 \]

\[ \binom{10}{3}=120 \]

So:

\[ P=\frac{60}{120}=\frac12 \]


40. Exercises

Exercise 1

Arrange 6 people in a row.

\[ 6!=720 \]

Exercise 2

President and vice president among 8 people.

\[ D_{8,2}=8\cdot7=56 \]

Exercise 3

Commission of 2 people among 8.

\[ \binom82=28 \]

Exercise 4

6-digit sequences with repetition.

\[ 10^6 \]

Exercise 5

6-digit sequences without repetition.

\[ 10\cdot9\cdot8\cdot7\cdot6\cdot5 \]

Exercise 6

Anagrams of:

\[ ANNA \]

\[ \frac{4!}{2!2!}=6 \]


41. Probability of a person being chosen

From 8 people we choose 3.

Probability of Marco being chosen:

\[ \frac{\binom72}{\binom83} \]

\[ = \frac{21}{56} = \frac38 \]


42. At least one ace

From 52 cards we choose 5.

Probability of at least one ace:

\[ 1- \frac{ \binom{48}{5} }{ \binom{52}{5} } \]

The complementary simplifies the calculation.


43. The method before the formula

Before choosing a formula, ask ourselves:

  1. what am I building or choosing?
  2. do I use all the elements?
  3. does the order count?
  4. Are repetitions allowed?
  5. Are there the same elements?
  6. Are there any constraints?

44. Quick Start Guide

All elements, important order:

\[ \boxed{n!} \]

\(k\) to \(n\), important order:

\[ \boxed{ \frac{n!}{(n-k)!} } \]

\(k\) on \(n\), important order with repetition:

\[ \boxed{n^k} \]

\(k\) over \(n\), order not important:

\[ \boxed{ \binom nk } \]

Repeating elements:

\[ \boxed{ \frac{n!}{n_1!n_2!\cdots n_r!} } \]


45. Summary of Chapter 4

Combinatorial calculus allows you to count without listing.

The basic formulas are:

\[ \boxed{P_n=n!} \]

\[ \boxed{ D_{n,k}=\frac{n!}{(n-k)!} } \]

\[ \boxed{ D'_{n,k}=n^k } \]

\[ \boxed{ \binom nk= \frac{n!}{k!(n-k)!} } \]

and, with repeated elements:

\[ \boxed{ \frac{n!}{n_1!n_2!\cdots n_r!} } \]

The most important rule remains:

\[ \boxed{ \text{prima ragionare, poi scegliere la formula} } \]

Chapter 5 – Conditional probability

1. The idea of conditional probability

So far we have calculated probabilities knowing only the random experiment.

But often we receive new information.

For example:

we roll a die and we know that an even number is rolled.

At this point the possible outcomes are no longer:

\[ \{1,2,3,4,5,6\} \]

but only:

\[ \{2,4,6\} \]

The new information has narrowed the space of possibilities.

If we now ask:

what is the probability that the number is greater than 3?

between:

\[ 2,4,6 \]

favorable outcomes are:

\[ 4,6 \]

therefore:

\[ \frac23 \]

This is a conditional probability.


2. Notation

The probability of the \(A\) event, knowing that \(B\) occurred, is indicated by:

\[ \boxed{P(A\mid B)} \]

and it reads:

probability of \(A\) given \(B\).


3. First example

Let's roll a dice.

Either:

\[ A=\{\text{numero maggiore di 3}\} \]

therefore:

\[ A=\{4,5,6\} \]

and:

\[ B=\{\text{numero pari}\} \]

therefore:

\[ B=\{2,4,6\} \]

We want:

\[ P(A\mid B) \]

Knowing that \(B\) occurred, we only consider:

\[ \{2,4,6\} \]

The outcomes also belonging to \(A\) are:

\[ \{4,6\} \]

So:

\[ \boxed{ P(A\mid B)=\frac23 } \]


4. Conditional probability formula

In general:

\[ \boxed{ P(A\mid B) = \frac{P(A\cap B)}{P(B)} } \]

provided:

\[ P(B)>0 \]


5. Why does the intersection appear

Knowing that \(B\) has occurred, \(B\) becomes the new possibility space.

Among the outcomes of \(B\), those favorable to \(A\) are:

\[ A\cap B \]

So:

\[ P(A\mid B) = \frac{\text{parte di }B\text{ favorevole ad }A} {\text{tutto }B} \]


6. Check with the dice

We have:

\[ P(A\cap B)=\frac26=\frac13 \]

and:

\[ P(B)=\frac36=\frac12 \]

So:

\[ P(A\mid B) = \frac{\frac13}{\frac12} = \frac23 \]


7. Example with cards

From a deck of 52 cards we draw a card.

We know it's a figure.

What is the probability that he is a king?

The figures are:

  • 4 Jacks;
  • 4 Women;
  • 4 Kings.

In total:

\[ 12 \]

The kings are:

\[ 4 \]

So:

\[ \boxed{ P(\text{Re}\mid\text{Figura}) = \frac4{12} = \frac13 } \]


8. Order matters

In general:

\[ \boxed{ P(A\mid B)\neq P(B\mid A) } \]

For example:

\[ P(\text{Re}\mid\text{Figura})=\frac13 \]

but:

\[ P(\text{Figura}\mid\text{Re})=1 \]

because all kings are figures.


9. Frequent error

Confuse:

\[ P(A\mid B) \]

with:

\[ P(B\mid A) \]

it is one of the most common errors.

The part after the slash is the information we know to be true.


10. Example with students

In a classroom:

  • 18 girls;
  • 12 boys;
  • 10 girls wear glasses;
  • 4 boys wear glasses.

Probability of glasses knowing it's a girl:

\[ P(O\mid R) = \frac{10}{18} = \frac59 \]


11. Let's reverse the question

Probability of being a girl knowing that she wears glasses:

Students with glasses are:

\[ 10+4=14 \]

therefore:

\[ P(R\mid O) = \frac{10}{14} = \frac57 \]

We observe:

\[ \frac59\neq\frac57 \]


12. Tree diagram

Conditional probability is well represented with a tree.

Suppose:

\[ P(A)=0,6 \]

\[ P(A^c)=0,4 \]

and:

\[ P(B\mid A)=0,7 \]

\[ P(B\mid A^c)=0,2 \]

Each branch represents a conditional probability.


13. Product rule

Let's start from:

\[ P(A\mid B) = \frac{P(A\cap B)}{P(B)} \]

Multiplying by \(P(B)\):

\[ \boxed{ P(A\cap B) = P(B)P(A\mid B) } \]

Similarly:

\[ \boxed{ P(A\cap B) = P(A)P(B\mid A) } \]


14. Interpretation

The probability of both events occurring can be calculated as:

\[ P(A) \]

multiplied by the probability of \(B\) knowing that \(A\) has already occurred.


15. Urn without reinsertion

An urn contains:

  • 3 red balls;
  • 2 blue.

We extract two balls without reinserting them.

Probability that they are both red:

\[ P(R_1)=\frac35 \]

After a red they remain:

  • 2 red;
  • 2 blue.

So:

\[ P(R_2\mid R_1)=\frac24 \]

and:

\[ P(R_1\cap R_2) = \frac35\cdot\frac24 \]

\[ = \frac3{10} \]


16. With reinstatement

If we put the first ball back in the urn:

\[ P(R_2\mid R_1)=\frac35 \]

So:

\[ P(R_1\cap R_2) = \frac35\cdot\frac35 = \frac9{25} \]


17. Dependence and independence

If knowing that \(A\) occurred changes the probability of \(B\), the events are dependent.

If instead:

\[ P(B\mid A)=P(B) \]

the events are independent.

This concept will be explored in more depth in Chapter 6.


18. Example of independence

We toss a coin and a dice.

Either:

\[ A=\{\text{dado mostra 6}\} \]

\[ B=\{\text{moneta mostra testa}\} \]

Knowing that heads does not change the probability of 6.

So:

\[ P(A\mid B)=P(A)=\frac16 \]


19. Dependency example

We extract two cards without reinserting.

Either:

\[ A=\{\text{prima carta è un asso}\} \]

\[ B=\{\text{seconda carta è un asso}\} \]

Before:

\[ P(B)=\frac4{52} \]

Knowing that the first one is an ace:

\[ P(B\mid A)=\frac3{51} \]

The two probabilities are different.


20. Two dice

We roll two dice.

We know that the sum is 8.

What is the probability that at least one shows a 5?

The outcomes with a sum of 8 are:

\[ (2,6),(3,5),(4,4),(5,3),(6,2) \]

Those with at least a 5:

\[ (3,5),(5,3) \]

So:

\[ \boxed{ P(\text{almeno un 5}\mid\text{somma}=8) = \frac25 } \]


21. Two-child problem

Simplified model:

\[ \{MM,MF,FM,FF\} \]

Knowing that at least one child is male, we eliminate:

\[ FF \]

Remaining:

\[ MM,MF,FM \]

So:

\[ \boxed{ P(MM\mid\text{almeno un maschio}) = \frac13 } \]


22. Different information produces different probabilities

If we know instead:

the first child is a boy

remain:

\[ MM,MF \]

So:

\[ P(MM\mid\text{primo figlio maschio}) = \frac12 \]

The conditioning information is decisive.


23. Three events

The product rule extends:

\[ \boxed{ P(A\cap B\cap C) = P(A) P(B\mid A) P(C\mid A\cap B) } \]


24. Example with three cards

From a deck of 52 cards we extract three cards without replacing them.

Probability that they are all aces:

\[ \frac4{52} \cdot \frac3{51} \cdot \frac2{50} \]


25. Alternative method

We can also write:

\[ \frac{\binom43}{\binom{52}{3}} \]

The two methods lead to the same result.


26. Mistake to avoid

We can't always write:

\[ P(A\cap B)=P(A)P(B) \]

The general correct formula is:

\[ \boxed{ P(A\cap B)=P(A)P(B\mid A) } \]

Only in case of independence:

\[ P(B\mid A)=P(B) \]

and therefore:

\[ P(A\cap B)=P(A)P(B) \]


27. Exercises

Exercise 1

Die: probability of 6 knowing that an even number is rolled.

\[ P(6\mid\text{pari})=\frac13 \]

Exercise 2

Card: probability of ace knowing that it is of hearts.

\[ P(\text{asso}\mid\text{cuori}) = \frac1{13} \]

Exercise 3

Urn with 5 red and 3 blue.

Two extractions without reinsertion.

Probability of two blues:

\[ \frac38\cdot\frac27 = \frac3{28} \]

Exercise 4

Probability first red and then blue:

\[ \frac58\cdot\frac37 = \frac{15}{56} \]

Exercise 5

If:

\[ P(A)=0,4 \]

and:

\[ P(B\mid A)=0,3 \]

then:

\[ P(A\cap B)=0,12 \]

Exercise 6

If:

\[ P(A\cap B)=0,18 \]

and:

\[ P(B)=0,3 \]

then:

\[ P(A\mid B)=0,6 \]


28. Summary of Chapter 5

The conditional probability is:

\[ \boxed{ P(A\mid B) = \frac{P(A\cap B)}{P(B)} } \]

The product rule is:

\[ \boxed{ P(A\cap B) = P(A)P(B\mid A) } \]

Fundamental idea:

\[ \boxed{ \text{una nuova informazione restringe lo spazio delle possibilità} } \]


Chapter 6 – Independence and dependence of events

1. The idea of independence

Two events are independent when the occurrence of one does not change the probability of the other.

In other words, knowing the result of \(A\) does not provide us with useful information about \(B\).


2. Definition via conditional probability

If:

\[ P(A)>0 \]

and:

\[ P(B)>0 \]

events are independent when:

\[ \boxed{ P(A\mid B)=P(A) } \]

and therefore also:

\[ \boxed{ P(B\mid A)=P(B) } \]


3. Equivalent definition

Using the product rule:

\[ P(A\cap B) = P(A)P(B\mid A) \]

If the events are independent:

\[ P(B\mid A)=P(B) \]

therefore:

\[ \boxed{ P(A\cap B)=P(A)P(B) } \]

This is the most used form.


4. Example with dice and coin

We toss a coin and a dice.

Either:

\[ A=\{\text{testa}\} \]

\[ B=\{\text{esce 6}\} \]

We have:

\[ P(A)=\frac12 \]

\[ P(B)=\frac16 \]

The probability of both occurring is:

\[ P(A\cap B)=\frac1{12} \]

Because:

\[ \frac12\cdot\frac16=\frac1{12} \]

the events are independent.


5. Example with two coin flips

Let's flip a fair coin twice.

Either:

\[ A=\{\text{prima testa}\} \]

\[ B=\{\text{seconda testa}\} \]

We have:

\[ P(A)=\frac12 \]

\[ P(B)=\frac12 \]

and:

\[ P(A\cap B)=\frac14 \]

Because:

\[ \frac12\cdot\frac12=\frac14 \]

the events are independent.


6. Addiction

Two events are dependent when:

\[ P(A\mid B)\neq P(A) \]

that is, knowing that \(B\) occurred changes the probability of \(A\).


7. Example with cards

From a deck of 52 cards we extract two cards without replacing them.

Either:

\[ A=\{\text{prima carta asso}\} \]

\[ B=\{\text{seconda carta asso}\} \]

We have:

\[ P(B)=\frac4{52} \]

but:

\[ P(B\mid A)=\frac3{51} \]

So \(A\) and \(B\) are dependent.


8. Reintegration and independence

If after the first draw we put the card back in the deck and shuffle:

\[ P(B\mid A)=\frac4{52} \]

equal to:

\[ P(B) \]

In this case the two extractions are independent.


9. Warning: incompatibility and independence are not the same thing

This is a key distinction.

Incompatible events

They cannot occur together:

\[ A\cap B=\varnothing \]

Independent events

The occurrence of one does not change the probability of the other:

\[ P(A\cap B)=P(A)P(B) \]

They are completely different concepts.


10. Example of incompatibility

Let's roll a dice.

Either:

\[ A=\{\text{esce 2}\} \]

\[ B=\{\text{esce 5}\} \]

They cannot occur together.

So:

\[ A\cap B=\varnothing \]

They are incompatible.


11. Are they also independent?

We have:

\[ P(A)=\frac16 \]

\[ P(B)=\frac16 \]

If they were independent it should hold:

\[ P(A\cap B) = \frac1{36} \]

But in reality:

\[ P(A\cap B)=0 \]

So they are not independent.


12. Important rule

Two incompatible events with positive probability cannot be independent.

In fact, if:

\[ P(A)>0 \]

and:

\[ P(B)>0 \]

for independence it would take:

\[ P(A\cap B)=P(A)P(B)>0 \]

but due to incompatibility:

\[ P(A\cap B)=0 \]

contradiction.


13. Intuitive example

If we know that 2 is out:

\[ P(\text{esce 5}\mid\text{esce 2})=0 \]

while before:

\[ P(\text{esce 5})=\frac16 \]

Knowing that \(A\) occurred completely changes the probability of \(B\).


14. Independent events can occur together

In coin and dice tossing:

\[ A=\{\text{testa}\} \]

\[ B=\{\text{6}\} \]

can occur simultaneously.

In fact:

\[ A\cap B \]

it is not empty.

Independence does not mean the impossibility of occurring together.


15. Product formula for independent events

If \(A\) and \(B\) are independent:

\[ \boxed{ P(A\cap B)=P(A)P(B) } \]

This formula is very important.


16. Three independent events

If three events:

\[ A,B,C \]

are mutually independent, then:

\[ P(A\cap B)=P(A)P(B) \]

\[ P(A\cap C)=P(A)P(C) \]

\[ P(B\cap C)=P(B)P(C) \]

and also:

\[ \boxed{ P(A\cap B\cap C) = P(A)P(B)P(C) } \]


17. Independence in pairs is not enough

It is possible for three events to be pairwise independent but not mutually independent.

This is a more advanced but important point.


18. Classic example

We flip two balanced coins.

Let's define:

\[ A=\{\text{prima moneta testa}\} \]

\[ B=\{\text{seconda moneta testa}\} \]

\[ C=\{\text{le due monete danno lo stesso risultato}\} \]

The space is:

\[ \{TT,TC,CT,CC\} \]

We have:

\[ P(A)=P(B)=P(C)=\frac12 \]


19. Independence in pairs

\[ A\cap B=\{TT\} \]

therefore:

\[ P(A\cap B)=\frac14 = \frac12\cdot\frac12 \]

Similarly:

\[ P(A\cap C)=\frac14 \]

and:

\[ P(B\cap C)=\frac14 \]

So the events are independent in pairs.


20. But not mutually independent

The triple intersection is:

\[ A\cap B\cap C=\{TT\} \]

therefore:

\[ P(A\cap B\cap C)=\frac14 \]

But:

\[ P(A)P(B)P(C) = \frac18 \]

Because:

\[ \frac14\neq\frac18 \]

the three events are not mutually independent.


21. Independence and complementary

If \(A\) and \(B\) are independent, then also:

\[ A^c \]

and:

\[ B \]

they are independent.

Also applies to:

\[ A \]

and:

\[ B^c \]

and for:

\[ A^c \]

and:

\[ B^c \]


22. Demonstration

If \(A\) and \(B\) are independent:

\[ P(A\cap B)=P(A)P(B) \]

Time:

\[ P(A^c\cap B) = P(B)-P(A\cap B) \]

\[ = P(B)-P(A)P(B) \]

\[ = P(B)(1-P(A)) \]

\[ = P(B)P(A^c) \]

So \(A^c\) and \(B\) are independent.


23. At least one success in independent tests

Suppose we have \(n\) independent evidence.

The probability of success in each is:

\[ p \]

The probability of no success is:

\[ (1-p)^n \]

So:

\[ \boxed{ P(\text{almeno un successo}) = 1-(1-p)^n } \]


24. Example with a dice

We roll a die 4 times.

What is the probability of rolling at least a 6?

Probability of not rolling 6 in one roll:

\[ \frac56 \]

Probability of no 6s in 4 rolls:

\[ \left(\frac56\right)^4 \]

So:

\[ \boxed{ P(\text{almeno un 6}) = 1-\left(\frac56\right)^4 } \]


25. Probability of a sequence

If the rolls are independent, the probability of a specific sequence is obtained by multiplying.

For example, with a balanced currency:

\[ P(TCTT) = \frac12\cdot\frac12\cdot\frac12\cdot\frac12 = \frac1{16} \]


26. Unbalanced currency

If:

\[ P(T)=p \]

and:

\[ P(C)=1-p \]

then:

\[ P(TCTT) = p(1-p)p^2 \]

\[ = p^3(1-p) \]


27. Connection with the binomial

This reasoning leads directly to the binomial distribution.

If in \(n\) independent trials we want exactly \(k\) successes, every sequence with \(k\) successes has probability:

\[ p^k(1-p)^{n-k} \]

The number of possible sequences is:

\[ \binom nk \]

So:

\[ P(X=k) = \binom nk p^k (1-p)^{n-k} \]


28. Example with given probabilities

Suppose:

\[ P(A)=0,4 \]

\[ P(B)=0,5 \]

and \(A\), \(B\) independent.

Then:

\[ P(A\cap B) = 0,4\cdot0,5 = 0,2 \]


29. Probability of union

We use:

\[ P(A\cup B) = P(A)+P(B)-P(A\cap B) \]

So:

\[ P(A\cup B) = 0,4+0,5-0,2 = 0,7 \]


30. Reverse example

Suppose:

\[ P(A)=0,3 \]

\[ P(B)=0,4 \]

\[ P(A\cap B)=0,12 \]

Because:

\[ 0,3\cdot0,4=0,12 \]

the events are independent.


31. Another example

If:

\[ P(A)=0,3 \]

\[ P(B)=0,4 \]

but:

\[ P(A\cap B)=0,20 \]

then:

\[ 0,20\neq0,12 \]

therefore events are dependent.


32. Independence does not mean equal probabilities

Two independent events can have very different probabilities.

For example:

\[ P(A)=0,9 \]

\[ P(B)=0,1 \]

they can be perfectly independent.

Independence concerns the relationship between events, not the value of their probabilities.


33. Example with repeated tests

A machine produces a defective part with probability:

\[ 0,02 \]

We assume that the pieces are independent.

What is the probability that 3 consecutive pieces are all good?

Probability that a piece is good:

\[ 0,98 \]

So:

\[ P(\text{3 buoni}) = 0,98^3 \]


34. Probability of at least one defective

The complementary is:

all good.

So:

\[ P(\text{almeno un difettoso}) = 1-0,98^3 \]


35. Exercises

Exercise 1

If:

\[ P(A)=0,5 \]

\[ P(B)=0,2 \]

and the events are independent, calculate:

\[ P(A\cap B) \]

Solution

\[ P(A\cap B)=0,5\cdot0,2=0,1 \]


Exercise 2

With the same data, calculate:

\[ P(A\cup B) \]

Solution

\[ P(A\cup B) = 0,5+0,2-0,1 = 0,6 \]


Exercise 3

If:

\[ P(A)=0,6 \]

\[ P(B)=0,5 \]

\[ P(A\cap B)=0,3 \]

are they independent?

Solution

\[ 0,6\cdot0,5=0,3 \]

Yes.


Exercise 4

Dice rolled 5 times.

Odds of no 6:

\[ \boxed{ \left(\frac56\right)^5 } \]

Odds of at least a 6:

\[ \boxed{ 1-\left(\frac56\right)^5 } \]


Exercise 5

Balanced coin flipped 8 times.

Probability of always getting heads:

\[ \boxed{ \left(\frac12\right)^8 } \]


36. Typical conceptual error

“The events are incompatible, therefore they are independent.”

It's false.

If two incompatible events have positive probability, the occurrence of one makes the other impossible.

So they are highly dependent.


37. Another error

“The events are independent, so they cannot happen together.”

It's false.

Independence requires precisely:

\[ P(A\cap B)=P(A)P(B) \]

which, if both probabilities are positive, is positive.


38. Comparison scheme

Incompatible

\[ \boxed{ P(A\cap B)=0 } \]

Independent

\[ \boxed{ P(A\cap B)=P(A)P(B) } \]

They are two distinct concepts.


39. Summary of Chapter 6

Two events are independent if:

\[ \boxed{ P(A\mid B)=P(A) } \]

equivalently:

\[ \boxed{ P(A\cap B)=P(A)P(B) } \]

They are dependent if the occurrence of one changes the probability of the other.

The fundamental distinction is:

\[ \boxed{ \text{indipendenza}\neq\text{incompatibilità} } \]

For incompatible events:

\[ P(A\cap B)=0 \]

For independent events:

\[ P(A\cap B)=P(A)P(B) \]

This distinction will be fundamental to understanding total probability, Bayes and binomial distribution.

Chapter 7 – Total probability theorem

1. The starting problem

Suppose that an event can occur through multiple different routes.

For example, a product can come from:

  • from machine 1;
  • from machine 2;
  • from the car 3.

If we want to know the probability that the product is defective, we must take into account:

  • the probability that it comes from each machine;
  • the probability that it is defective knowing which machine it comes from.

The total probability theorem allows you to combine this information.


2. An intuitive example

A factory uses two machines.

The \(M_1\) machine produces the:

\[ 60\% \]

of the pieces.

The \(M_2\) machine produces the:

\[ 40\% \]

of the pieces.

So:

\[ P(M_1)=0,6 \]

\[ P(M_2)=0,4 \]

Let's also assume that:

\[ P(D\mid M_1)=0,02 \]

and:

\[ P(D\mid M_2)=0,05 \]

where \(D\) is the event:

the part is defective.

We want to calculate:

\[ P(D) \]


3. We divide the problem into two paths

A defective part can come from:

  • from \(M_1\);
  • or from \(M_2\).

So:

\[ D=(D\cap M_1)\cup(D\cap M_2) \]

The two events are incompatible.

Therefore:

\[ P(D) = P(D\cap M_1)+P(D\cap M_2) \]


4. Product rule

We know that:

\[ P(D\cap M_1) = P(M_1)P(D\mid M_1) \]

and:

\[ P(D\cap M_2) = P(M_2)P(D\mid M_2) \]

So:

\[ P(D) = 0,6\cdot0,02 + 0,4\cdot0,05 \]

\[ = 0,012+0,020 \]

\[ = 0,032 \]

So:

\[ \boxed{ P(D)=3,2\% } \]


5. Fundamental idea

The overall probability is the sum of the probabilities of the different paths.

Intuitively:

\[ \boxed{ \text{probabilità totale} = \text{somma delle probabilità dei diversi percorsi} } \]


6. Partition of the sample space

The events:

\[ B_1,B_2,\ldots,B_n \]

form a partition of the sample space if:

  1. they are incompatible two by two:

\[ B_i\cap B_j=\varnothing \]

for:

\[ i\neq j \]

  1. they cover the entire sample space:

\[ B_1\cup B_2\cup\cdots\cup B_n=\Omega \]


7. Partition example

Suppose we divide a population into age groups:

\[ B_1=\{\text{meno di 30 anni}\} \]

\[ B_2=\{\text{tra 30 e 60 anni}\} \]

\[ B_3=\{\text{più di 60 anni}\} \]

If the bands do not overlap and cover the entire population, they constitute a partition.


8. Total probability theorem

If:

\[ B_1,B_2,\ldots,B_n \]

form a partition, then for each event \(A\):

\[ \boxed{ P(A) = \sum_{i=1}^{n} P(B_i)P(A\mid B_i) } \]

that is:

\[ P(A) = P(B_1)P(A\mid B_1) +\cdots+ P(B_n)P(A\mid B_n) \]


9. Because it works

The \(A\) event can be broken down into:

\[ A\cap B_1,\quad A\cap B_2,\quad \ldots,\quad A\cap B_n \]

These events are incompatible.

So:

\[ P(A) = \sum_i P(A\cap B_i) \]

But:

\[ P(A\cap B_i) = P(B_i)P(A\mid B_i) \]

hence the theorem.


10. Tree diagram

The theorem can be interpreted well with a tree.

Branches start from the root:

\[ B_1,\ B_2,\ B_3,\ldots \]

The \(A\) event then starts from each branch.

The probability of a path is:

\[ P(B_i)P(A\mid B_i) \]


11. Tree rule of thumb

We remember:

\[ \boxed{\text{lungo un percorso si moltiplica}} \]

and:

\[ \boxed{\text{tra percorsi alternativi si somma}} \]


12. Example with three factories

A company produces in three factories:

\[ P(S_1)=0,5 \]

\[ P(S_2)=0,3 \]

\[ P(S_3)=0,2 \]

The defect rates are:

\[ P(D\mid S_1)=0,01 \]

\[ P(D\mid S_2)=0,03 \]

\[ P(D\mid S_3)=0,04 \]

Let's calculate:

\[ P(D) \]


13. Solution

\[ P(D) = 0,5\cdot0,01 + 0,3\cdot0,03 + 0,2\cdot0,04 \]

\[ = 0,005+0,009+0,008 \]

\[ = 0,022 \]

So:

\[ \boxed{ P(D)=2,2\% } \]


14. Weighted average

The formula:

\[ P(D) = 0,5\cdot0,01 + 0,3\cdot0,03 + 0,2\cdot0,04 \]

can be seen as a weighted average of the conditional probabilities.

The weights are:

\[ 0,5,\quad0,3,\quad0,2 \]


15. Important consequence

Since the weights sum to 1:

\[ P(B_1)+\cdots+P(B_n)=1 \]

the total probability must be between the minimum and maximum of the conditional probabilities:

\[ \min_i P(A\mid B_i) \leq P(A) \leq \max_i P(A\mid B_i) \]


16. Case \(B\) and \(B^c\)

An event and its complement form a partition.

So:

\[ \boxed{ P(A) = P(B)P(A\mid B) + P(B^c)P(A\mid B^c) } \]


17. Example

30% of a population uses a service:

\[ P(S)=0,3 \]

therefore:

\[ P(S^c)=0,7 \]

Suppose:

\[ P(A\mid S)=0,8 \]

\[ P(A\mid S^c)=0,2 \]

Then:

\[ P(A) = 0,3\cdot0,8 + 0,7\cdot0,2 \]

\[ = 0,24+0,14 = 0,38 \]

So:

\[ \boxed{ P(A)=38\% } \]


18. Example with urns

Let's choose an urn:

\[ P(U_1)=\frac13 \]

\[ P(U_2)=\frac23 \]

The \(U_1\) urn contains:

  • 4 red ones;
  • 6 blue.

The \(U_2\) urn contains:

  • 7 red;
  • 3 blue.

We want the probability of drawing a red ball.


19. Solution

\[ P(R\mid U_1)=\frac4{10} \]

\[ P(R\mid U_2)=\frac7{10} \]

So:

\[ P(R) = \frac13\cdot\frac4{10} + \frac23\cdot\frac7{10} \]

\[ = \frac4{30} + \frac{14}{30} = \frac{18}{30} = \frac35 \]

So:

\[ \boxed{ P(R)=\frac35 } \]


20. Example with students

In a school:

  • 40% scientific;
  • 35% classic;
  • 25% linguistic.

The probability of participating in an activity is:

\[ 0,60 \]

for the scientific,

\[ 0,50 \]

for the classic,

\[ 0,40 \]

for linguistics.

So:

\[ P(A) = 0,40\cdot0,60 + 0,35\cdot0,50 + 0,25\cdot0,40 \]

\[ = 0,24+0,175+0,10 \]

\[ = 0,515 \]

So:

\[ \boxed{ P(A)=51,5\% } \]


21. Example with diagnostic test

Suppose:

\[ P(M)=0,02 \]

\[ P(M^c)=0,98 \]

and:

\[ P(+\mid M)=0,95 \]

\[ P(+\mid M^c)=0,04 \]

What is the overall probability of a positive test?


22. Solution

\[ P(+) = P(M)P(+\mid M) + P(M^c)P(+\mid M^c) \]

\[ = 0,02\cdot0,95 + 0,98\cdot0,04 \]

\[ = 0,019+0,0392 \]

\[ = 0,0582 \]

So:

\[ \boxed{ P(+)=5,82\% } \]


23. A natural question

Now we can ask:

knowing that the test is positive, what is the probability that the person has the condition?

We want:

\[ P(M\mid+) \]

Instead we know:

\[ P(+\mid M) \]

The two probabilities are not the same.

This leads to Bayes' Theorem.


24. Absolute frequencies

We can make the problem more intuitive by imagining 10,000 people.

With:

\[ P(M)=0,02 \]

we have:

\[ 200 \]

people with the condition.

95% test positive:

\[ 190 \]

People without the condition are:

\[ 9800 \]

4% test positive:

\[ 392 \]

Total positives:

\[ 190+392=582 \]

and:

\[ \frac{582}{10000}=0,0582 \]


25. Example with means of transport

A person goes to work:

  • car 50%;
  • train 30%;
  • bike 20%.

Probability of delay:

\[ 0,10 \]

with car,

\[ 0,20 \]

by train,

\[ 0,05 \]

with bike.

So:

\[ P(R) = 0,50\cdot0,10 + 0,30\cdot0,20 + 0,20\cdot0,05 \]

\[ = 0,05+0,06+0,01 = 0,12 \]

So:

\[ \boxed{12\%} \]


26. General procedure

Step 1

Identify the causes:

\[ B_1,\ldots,B_n \]

Step 2

Verify that they form a partition.

Step 3

Locate the final event \(A\).

Step 4

Write:

\[ P(B_i) \]

and:

\[ P(A\mid B_i) \]

Step 5

Multiply along each path:

\[ P(B_i)P(A\mid B_i) \]

Step 6

Add:

\[ P(A) = \sum_i P(B_i)P(A\mid B_i) \]


27. Exercises

Exercise 1

Two machines produce:

\[ 70\% \]

and:

\[ 30\% \]

of the pieces.

Defect Rates:

\[ 2\% \]

and:

\[ 5\% \]

So:

\[ P(D) = 0,7\cdot0,02 + 0,3\cdot0,05 \]

\[ = 0,029 \]

\[ \boxed{2,9\%} \]


Exercise 2

Two urns chosen with equal probability.

Urn A:

  • 8 white;
  • 2 black.

Urn B:

  • 3 white;
  • 7 black.

So:

\[ P(Bianca) = \frac12\cdot\frac8{10} + \frac12\cdot\frac3{10} \]

\[ = \frac{11}{20} = 0,55 \]


Exercise 3

25% purchases online, 75% in store.

Returns:

\[ 10\% \]

online,

\[ 4\% \]

shop.

\[ P(R) = 0,25\cdot0,10 + 0,75\cdot0,04 \]

\[ = 0,055 \]

\[ \boxed{5,5\%} \]


28. Summary of Chapter 7

If:

\[ B_1,\ldots,B_n \]

form a partition:

\[ \boxed{ P(A) = \sum_iP(B_i)P(A\mid B_i) } \]

Rule of thumb:

\[ \boxed{\text{moltiplica lungo i percorsi}} \]

\[ \boxed{\text{somma i percorsi alternativi}} \]


Chapter 8 – Bayes' theorem

1. The inverse problem

In the previous chapter we calculated:

\[ P(A) \]

knowing:

\[ P(B_i) \]

and:

\[ P(A\mid B_i) \]

Now we want to do the reverse path.

We look at \(A\) and want to know:

what is the probability that \(B_i\) was the cause?

That is, we want:

\[ P(B_i\mid A) \]


2. Where the formula comes from

Let's start from:

\[ P(B\mid A) = \frac{P(A\cap B)}{P(A)} \]

But:

\[ P(A\cap B) = P(B)P(A\mid B) \]

So:

\[ \boxed{ P(B\mid A) = \frac{ P(B)P(A\mid B) }{ P(A) } } \]

This is the basic form of Bayes' Theorem.


3. General formula

If:

\[ B_1,\ldots,B_n \]

form a partition:

\[ \boxed{ P(B_i\mid A) = \frac{ P(B_i)P(A\mid B_i) }{ \sum_j P(B_j)P(A\mid B_j) } } \]


4. Intuitive interpretation

Bayes updates an initial probability after receiving new information.

Scheme:

\[ \boxed{ \text{probabilità iniziale} \rightarrow \text{nuova evidenza} \rightarrow \text{probabilità aggiornata} } \]


5. A priori and a posteriori

The initial probability:

\[ P(B) \]

It's called a priori probability.

The updated probability:

\[ P(B\mid A) \]

It's called posterior probability.


6. Example with urns

We choose:

\[ P(U_1)=\frac13 \]

\[ P(U_2)=\frac23 \]

Urn \(U_1\):

  • 4 red ones;
  • 6 blue.

Urn \(U_2\):

  • 7 red;
  • 3 blue.

We observe a red ball.

We want:

\[ P(U_1\mid R) \]


7. Total probability of red

\[ P(R) = \frac13\cdot\frac4{10} + \frac23\cdot\frac7{10} \]

\[ = \frac4{30} + \frac{14}{30} = \frac{18}{30} = \frac35 \]


8. Let's apply Bayes

\[ P(U_1\mid R) = \frac{ \frac13\cdot\frac4{10} }{ \frac35 } \]

\[ = \frac29 \]

So:

\[ \boxed{ P(U_1\mid R)=\frac29 } \]


9. Second urn

\[ P(U_2\mid R) = 1-\frac29 = \frac79 \]


10. Meaning

Before extraction:

\[ P(U_1)=\frac13 \]

After looking at red:

\[ P(U_1\mid R)=\frac29 \]

The evidence changed the initial probability.


11. Example with factories

Two factories produce:

\[ 70\% \]

and:

\[ 30\% \]

of the pieces.

Defect Rates:

\[ 2\% \]

and:

\[ 5\% \]

One piece is defective.

What is the probability that it comes from the second plant?


12. Probability of defect

\[ P(D) = 0,7\cdot0,02 + 0,3\cdot0,05 \]

\[ = 0,014+0,015 = 0,029 \]


13. Bayes

\[ P(S_2\mid D) = \frac{ 0,3\cdot0,05 }{ 0,029 } \]

\[ = \frac{0,015}{0,029} \]

\[ \approx0,5172 \]

So:

\[ \boxed{ P(S_2\mid D)\approx51,7\% } \]


14. Interpretation

The second factory produces only 30% of the parts.

But among the defective ones, the probability of originating from \(S_2\) rises above 50%.

This is because it has a higher defect rate.


15. Diagnostic test

Suppose:

\[ P(M)=0,02 \]

\[ P(M^c)=0,98 \]

and:

\[ P(+\mid M)=0,95 \]

\[ P(+\mid M^c)=0,04 \]

One person tests positive.

We want:

\[ P(M\mid+) \]


16. Let's calculate \(P(+)\)

\[ P(+) = 0,02\cdot0,95 + 0,98\cdot0,04 \]

\[ = 0,019+0,0392 \]

\[ = 0,0582 \]


17. Let's apply Bayes

\[ P(M\mid+) = \frac{ 0,02\cdot0,95 }{ 0,0582 } \]

\[ = \frac{0,019}{0,0582} \]

\[ \approx0,3265 \]

So:

\[ \boxed{ P(M\mid+)\approx32,6\% } \]


18. Counterintuitive result

The test recognizes 95% of people with the condition.

Yet a positive has about a 32.6% chance of actually having the condition.

The reason is that the condition is rare.


19. Natural frequencies

Let's imagine:

\[ 10000 \]

people.

With the condition:

\[ 200 \]

True positives:

\[ 200\cdot0,95=190 \]

Without condition:

\[ 9800 \]

False positives:

\[ 9800\cdot0,04=392 \]

Total positives:

\[ 190+392=582 \]

So:

\[ \frac{190}{582} \approx0,3265 \]


20. The two probabilities are not equal

\[ P(+\mid M)=95\% \]

but:

\[ P(M\mid+)\approx32,6\% \]

In general:

\[ \boxed{ P(A\mid B)\neq P(B\mid A) } \]


21. Bayes as inversion

Bayes allows you to go from:

\[ P(A\mid B) \]

to:

\[ P(B\mid A) \]

correcting for basic probabilities.


22. Interpretative form

We can remember:

\[ \boxed{ \text{posteriori} = \frac{ \text{priori}\times\text{verosimiglianza} }{ \text{evidenza} } } \]

that is:

\[ P(B\mid A) = \frac{ P(B)P(A\mid B) }{ P(A) } \]


23. Example with students

In a school:

  • 60% for two years;
  • 40% three-year period.

Take a course:

  • 20% of the two-year period;
  • 50% of the three-year period.

We want the probability that a participating student is a three-year student.


24. Probability of participation

\[ P(C) = 0,6\cdot0,2 + 0,4\cdot0,5 \]

\[ = 0,12+0,20 = 0,32 \]


25. Bayes

\[ P(T\mid C) = \frac{ 0,4\cdot0,5 }{ 0,32 } \]

\[ = 0,625 \]

So:

\[ \boxed{ P(T\mid C)=62,5\% } \]


26. Example with three causes

Three machines produce:

\[ 50\%,\quad30\%,\quad20\% \]

of the pieces.

Defect Rates:

\[ 1\%,\quad3\%,\quad4\% \]

One piece is defective.

What is the probability that it comes from the third car?


27. Total defect probability

\[ P(D) = 0,5\cdot0,01 + 0,3\cdot0,03 + 0,2\cdot0,04 \]

\[ = 0,005+0,009+0,008 = 0,022 \]


28. Bayes

\[ P(M_3\mid D) = \frac{ 0,2\cdot0,04 }{ 0,022 } \]

\[ = \frac{0,008}{0,022} \]

\[ \approx0,3636 \]

So:

\[ \boxed{ P(M_3\mid D)\approx36,4\% } \]


29. Check

\[ P(M_1\mid D) = \frac{0,005}{0,022} \approx22,7\% \]

\[ P(M_2\mid D) = \frac{0,009}{0,022} \approx40,9\% \]

\[ P(M_3\mid D) \approx36,4\% \]

The sum is approximately:

\[ 100\% \]


30. Bayes as path ratio

We can remember:

\[ \boxed{ P(B_i\mid A) = \frac{ \text{probabilità del percorso }B_i\to A }{ \text{probabilità totale di arrivare ad }A } } \]


31. Base rate error

A common mistake is ignoring:

\[ P(B) \]

that is, the initial probability.

If an event is very rare, even an accurate test can produce many false positives compared to true positives.


32. When to use Bayes

Bayes is useful when:

  1. there are several possible causes;
  2. we know the initial probability of the causes;
  3. we know the probability of the evidence given each cause;
  4. we look at the evidence;
  5. we want to trace the cause.

33. Exercise 1

Two urns:

\[ P(U_1)=0,4 \]

\[ P(U_2)=0,6 \]

Probability of white:

\[ P(B\mid U_1)=0,3 \]

\[ P(B\mid U_2)=0,7 \]

Let's calculate:

\[ P(B) = 0,4\cdot0,3 + 0,6\cdot0,7 \]

\[ = 0,54 \]

So:

\[ P(U_1\mid B) = \frac{0,12}{0,54} = \frac29 \]


34. Exercise 2

Line 1:

\[ 75\% \]

production, defects:

\[ 2\% \]

Line 2:

\[ 25\% \]

production, defects:

\[ 6\% \]

Total probability:

\[ P(D) = 0,75\cdot0,02 + 0,25\cdot0,06 \]

\[ = 0,03 \]

So:

\[ P(L_2\mid D) = \frac{0,015}{0,03} = 0,5 \]

\[ \boxed{50\%} \]


35. Exercise 3

Characteristic present in 10% of the population.

Positive test:

  • 90% if present;
  • 20% if absent.

\[ P(+) = 0,1\cdot0,9 + 0,9\cdot0,2 \]

\[ = 0,27 \]

So:

\[ P(C\mid+) = \frac{0,09}{0,27} = \frac13 \]

\[ \boxed{33,3\%} \]


36. Counterintuitive exercise

One condition affects:

\[ 1 \]

person on:

\[ 1000 \]

So:

\[ P(M)=0,001 \]

Tests:

\[ P(+\mid M)=0,99 \]

\[ P(+\mid M^c)=0,01 \]

Let's calculate:

\[ P(+) = 0,001\cdot0,99 + 0,999\cdot0,01 \]

\[ = 0,01098 \]

So:

\[ P(M\mid+) = \frac{0,00099}{0,01098} \]

\[ \approx0,0902 \]

that is, approximately:

\[ \boxed{9\%} \]


37. Bayes does not create information

The theorem requires initial data:

  • a priori probability;
  • conditional probabilities;
  • observed evidence.

If this information is uncertain, the outcome will also be uncertain.


38. General procedure

Step 1

Identify the causes:

\[ B_1,\ldots,B_n \]

Step 2

Write:

\[ P(B_i) \]

Step 3

Write:

\[ P(A\mid B_i) \]

Step 4

Calculate:

\[ P(A) \]

with the total probability.

Step 5

Apply:

\[ \boxed{ P(B_i\mid A) = \frac{ P(B_i)P(A\mid B_i) }{ P(A) } } \]


39. Total probability and Bayes

The two theorems work together.

Total probability:

\[ \boxed{ P(A) = \sum_iP(B_i)P(A\mid B_i) } \]

Bayes:

\[ \boxed{ P(B_i\mid A) = \frac{ P(B_i)P(A\mid B_i) }{ \sum_jP(B_j)P(A\mid B_j) } } \]


40. Final idea

Bayes can be summarized like this:

\[ \boxed{ \text{una nuova informazione modifica razionalmente le probabilità iniziali} } \]

First we have:

\[ P(B) \]

After observing \(A\):

\[ P(B\mid A) \]

It is one of the fundamental principles of statistics, diagnostics and probabilistic inference.

Chapter 9 – Discrete and continuous random variables

1. Why we introduce random variables

So far we have studied events such as:

an even number comes out

or:

the part is defective.

But often we want to associate a number with the result of the experiment.

For example:

  • number of heads in 5 throws;
  • number of customers entering a store in an hour;
  • waiting time;
  • height of a person;
  • lifespan of a light bulb.

To mathematically describe these quantities we introduce the concept of random variable.


2. Intuitive definition

A random variable is a function that associates a real number with every possible outcome of a random experiment.

In symbols:

\[ X:\Omega\rightarrow\mathbb{R} \]

At each outcome:

\[ \omega\in\Omega \]

the variable associates:

\[ X(\omega) \]


3. Example with two coins

Let's flip two coins.

\[ \Omega=\{TT,TC,CT,CC\} \]

Let's define:

\[ X=\text{numero di teste} \]

Then:

\[ X(TT)=2 \]

\[ X(TC)=1 \]

\[ X(CT)=1 \]

\[ X(CC)=0 \]

So \(X\) can take:

\[ 0,1,2 \]


4. Outcome and value of the variable

In the example:

\[ TC \]

it is an outcome.

Instead:

\[ X(TC)=1 \]

is the value of the random variable.

The sample space contains the outcomes.

The variable turns them into numbers.


5. Because it is useful

In the toss of 10 coins there are:

\[ 2^{10}=1024 \]

possible sequences.

If we are only interested in the number of heads, the variable can only take:

\[ 0,1,2,\ldots,10 \]

This greatly simplifies the problem.


6. Discrete random variable

A variable is discrete when it takes on a finite or countable number of values.

Examples:

  • number of children;
  • number of telephone calls;
  • number of defects;
  • number of successes;
  • number of times it appears 6.

7. Continuous random variable

A variable is continuous when it can take on all values in a range.

Examples:

  • height;
  • weight;
  • temperature;
  • time;
  • distance;
  • duration.

8. Discreet example

Let's roll a dice.

\[ X=\text{numero ottenuto} \]

Then:

\[ X\in\{1,2,3,4,5,6\} \]


9. Continuous example

We measure the time it takes for an activity.

We can have:

\[ 10,4 \]

\[ 10,43 \]

\[ 10,437 \]

and so on.

Ideally:

\[ X\in[0,+\infty) \]


10. Discrete probability function

If \(X\) is discrete:

\[ \boxed{ p(x)=P(X=x) } \]

It must be worth:

\[ p(x)\geq0 \]

and:

\[ \sum_xp(x)=1 \]


11. Example with two coins

For:

\[ X=\text{numero di teste} \]

we have:

\[ P(X=0)=\frac14 \]

\[ P(X=1)=\frac12 \]

\[ P(X=2)=\frac14 \]

and:

\[ \frac14+\frac12+\frac14=1 \]


12. Probability distribution

The set of possible values of \(X\) and their respective probabilities constitutes its:

\[ \boxed{\text{distribuzione di probabilità}} \]


13. Regular nut

If:

\[ X=\text{numero ottenuto} \]

then:

\[ P(X=x)=\frac16 \]

for:

\[ x=1,\ldots,6 \]


14. Indicator variable

Let's define:

\[ I_A= \begin{cases} 1 & \text{se }A\text{ si verifica}\\ 0 & \text{altrimenti} \end{cases} \]

A variable of this type is called:

\[ \boxed{\text{indicatrice}} \]


15. Example with dice

Let's define:

\[ X= \begin{cases} 1 & \text{se esce pari}\\ 0 & \text{se esce dispari} \end{cases} \]

Then:

\[ P(X=1)=\frac12 \]

\[ P(X=0)=\frac12 \]


16. Continuous variables and punctual probability

For a continuous variable:

\[ \boxed{ P(X=x)=0 } \]

for each individual \(x\) value.

This doesn't mean that value is impossible.

It means that the probability of a single point is zero.


17. Probability density

For a continuous variable we use a function:

\[ f(x) \]

called density.

It must be worth:

\[ f(x)\geq0 \]

and:

\[ \int_{-\infty}^{+\infty}f(x)\,dx=1 \]


18. Probability is an area

For a continuous variable:

\[ \boxed{ P(a\leq X\leq b) = \int_a^bf(x)\,dx } \]

Probability is therefore an area under the curve.


19. Density and probability are not the same thing

The value:

\[ f(x) \]

it is not a precise probability.

A density can even exceed 1.

What must equal 1 is the total area.


20. Uniform continued

If:

\[ X\sim U(0,10) \]

the density is:

\[ f(x)=\frac1{10} \]

for:

\[ 0\leq x\leq10 \]


21. Example

\[ P(2\leq X\leq5) = 3\cdot\frac1{10} = 0,3 \]


22. Distribution function

For any variable we define:

\[ \boxed{ F(x)=P(X\leq x) } \]

It is the probability accumulated up to \(x\).


23. Discreet example

Suppose:

\[ P(X=0)=0,2 \]

\[ P(X=1)=0,5 \]

\[ P(X=2)=0,3 \]

Then:

\[ F(0)=0,2 \]

\[ F(1)=0,7 \]

\[ F(2)=1 \]


24. Properties of \(F(x)\)

The distribution function:

  • is between 0 and 1;
  • does not decrease;
  • tends to 0 for \(x\to-\infty\);
  • curtains to 1 for \(x\to+\infty\).

25. Probability of an interval

\[ \boxed{ P(a<X\leq b) = F(b)-F(a) } \]


26. Discreet and continuous

For a discrete variable:

\[ P(X=x) \]

it can be positive.

For a continuation:

\[ P(X=x)=0 \]


27. Example: number of customers

\[ X=\text{numero di clienti in un'ora} \]

Possible values:

\[ 0,1,2,\ldots \]

So:

\[ \boxed{\text{discreta}} \]


28. Example: waiting time

\[ X=\text{tempo di attesa} \]

It can take on positive real values.

So:

\[ \boxed{\text{continua}} \]


29. Multiple variables on the same experiment

By rolling two dice we can define:

\[ X=\text{somma} \]

\[ Y=\text{massimo} \]

\[ Z=\text{numero di 6} \]

The sample space is the same, but the variables describe different aspects.


30. Sum of two dice

If:

\[ X=D_1+D_2 \]

then:

\[ X\in\{2,3,\ldots,12\} \]

with probability:

\[ \frac1{36},\frac2{36},\frac3{36},\ldots \]

The sum 7 has probability:

\[ \frac6{36} \]


31. Distribution of the sum

\[ P(X=2)=\frac1{36} \]

\[ P(X=3)=\frac2{36} \]

\[ P(X=4)=\frac3{36} \]

\[ P(X=5)=\frac4{36} \]

\[ P(X=6)=\frac5{36} \]

\[ P(X=7)=\frac6{36} \]

and then symmetrically.


32. Exercises

Exercise 1

Three coins.

\[ X=\text{numero di teste} \]

Values:

\[ \boxed{0,1,2,3} \]

Exercise 2

\[ P(X=3)=\frac18 \]

Exercise 3

\[ P(X=2)=\frac38 \]

Exercise 4

Number of emails received:

\[ \boxed{\text{discreta}} \]

Exercise 5

Temperature:

\[ \boxed{\text{continua}} \]


33. Summary of Chapter 9

A random variable is:

\[ \boxed{ X:\Omega\rightarrow\mathbb R } \]

It can be:

\[ \boxed{\text{discreta}} \]

or:

\[ \boxed{\text{continua}} \]

For a discreet:

\[ p(x)=P(X=x) \]

For a continuation:

\[ P(a\leq X\leq b) = \int_a^bf(x)\,dx \]

For both:

\[ \boxed{ F(x)=P(X\leq x) } \]


Chapter 10 – Expected value, variance and standard deviation

1. Because they are needed

A distribution can be summarized with some fundamental numbers:

\[ \boxed{\text{valore atteso}} \]

\[ \boxed{\text{varianza}} \]

\[ \boxed{\text{deviazione standard}} \]

The expected value describes the center.

Variance and standard deviation describe dispersion.


2. Expected value

For a discrete variable:

\[ \boxed{ E(X)=\sum_i x_ip_i } \]

It is a weighted average of the possible values.


3. Example with dice

\[ E(X) = 1\cdot\frac16+ 2\cdot\frac16+ \cdots+ 6\cdot\frac16 \]

\[ = \frac{21}{6} = 3,5 \]

So:

\[ \boxed{E(X)=3,5} \]


4. An important point

The die cannot show:

\[ 3,5 \]

The expected value does not necessarily have to be a possible value of the variable.

It is a theoretical long-term average.


5. Two coins

If:

\[ X=\text{numero di teste} \]

with probability:

\[ \frac14,\frac12,\frac14 \]

then:

\[ E(X) = 0\cdot\frac14+ 1\cdot\frac12+ 2\cdot\frac14 = 1 \]


6. Games and expected value

Suppose:

  • win €20 with probability 0.1;
  • you lose €2 with probability 0.9.

\[ E(X) = 20\cdot0,1 + (-2)\cdot0,9 \]

\[ = 0,2 \]

So the theoretical average gain is:

\[ 0,20€ \]


7. Fair play

A game is fair if:

\[ \boxed{ E(X)=0 } \]

If:

\[ E(X)>0 \]

is favorable to the player in terms of average.

If:

\[ E(X)<0 \]

it is unfavorable.


8. Linearity of the expected value

\[ \boxed{ E(aX+b)=aE(X)+b } \]

and:

\[ \boxed{ E(X+Y)=E(X)+E(Y) } \]

This property holds even without independence.


9. Indicator variable

For:

\[ I_A= \begin{cases} 1 & A\\ 0 & A^c \end{cases} \]

we have:

\[ E(I_A)=P(A) \]

So:

\[ \boxed{ E(I_A)=P(A) } \]


10. Because the average is not enough

Two distributions can have the same mean but very different dispersions.

Example:

A

\[ X=10 \]

always.

B

\[ Y= \begin{cases} 0 & \text{prob. }1/2\\ 20 & \text{prob. }1/2 \end{cases} \]

Both have average:

\[ 10 \]

but the second is much more dispersed.


11. Deviation from the average

Let's say:

\[ \mu=E(X) \]

The gap is:

\[ X-\mu \]

The average of the deviations is 0.

For this we consider the squared deviations.


12. Variance

\[ \boxed{ \operatorname{Var}(X) = E[(X-\mu)^2] } \]

For discrete variables:

\[ \boxed{ \operatorname{Var}(X) = \sum_i(x_i-\mu)^2P(X=x_i) } \]


13. Example with two coins

We had:

\[ E(X)=1 \]

So:

\[ \operatorname{Var}(X) = (0-1)^2\frac14+ (1-1)^2\frac12+ (2-1)^2\frac14 \]

\[ = \frac12 \]


14. Alternative formula

\[ \boxed{ \operatorname{Var}(X) = E(X^2)-[E(X)]^2 } \]

It is often more comfortable.


15. Demonstration

\[ (X-\mu)^2 = X^2-2\mu X+\mu^2 \]

Taking the expected value:

\[ E(X^2)-2\mu E(X)+\mu^2 \]

Because:

\[ E(X)=\mu \]

we get:

\[ E(X^2)-\mu^2 \]


16. Example with dice

\[ E(X)=3,5 \]

Let's calculate:

\[ E(X^2) = \frac{1+4+9+16+25+36}{6} = \frac{91}{6} \]

So:

\[ \operatorname{Var}(X) = \frac{91}{6} - \left(\frac72\right)^2 \]

\[ = \frac{35}{12} \]


17. Standard deviation

The variance is expressed in square units.

To return to the original unit we define:

\[ \boxed{ \sigma= \sqrt{\operatorname{Var}(X)} } \]


18. Example with dice

\[ \sigma = \sqrt{\frac{35}{12}} \approx1,708 \]


19. Interpretation

Small standard deviation:

\[ \text{valori concentrati} \]

Large standard deviation:

\[ \text{valori dispersi} \]


20. Same mean, different dispersion

Suppose:

\[ X= \begin{cases} 9 & 1/2\\ 11 & 1/2 \end{cases} \]

and:

\[ Y= \begin{cases} 0 & 1/2\\ 20 & 1/2 \end{cases} \]

Both have an average of 10.

But:

\[ \operatorname{Var}(X)=1 \]

while:

\[ \operatorname{Var}(Y)=100 \]


21. Linear transformations

If:

\[ Y=aX+b \]

then:

\[ \boxed{ E(Y)=aE(X)+b } \]

and:

\[ \boxed{ \operatorname{Var}(Y) = a^2\operatorname{Var}(X) } \]


22. Standard deviation

\[ \boxed{ \sigma_Y=|a|\sigma_X } \]

The term \(b\) does not change the dispersion.


23. Sum of variables

\[ E(X+Y)=E(X)+E(Y) \]

In general:

\[ \operatorname{Var}(X+Y) = \operatorname{Var}(X) + \operatorname{Var}(Y) + 2\operatorname{Cov}(X,Y) \]


24. Independent case

If \(X\) and \(Y\) are independent:

\[ \boxed{ \operatorname{Var}(X+Y) = \operatorname{Var}(X) + \operatorname{Var}(Y) } \]


25. Attention

Not valid:

\[ \sigma_{X+Y} = \sigma_X+\sigma_Y \]

In case of independence:

\[ \sigma_{X+Y} = \sqrt{ \sigma_X^2+\sigma_Y^2 } \]


26. Continuous variables

For a continuous variable:

\[ \boxed{ E(X) = \int_{-\infty}^{+\infty}xf(x)\,dx } \]

and:

\[ \boxed{ \operatorname{Var}(X) = \int_{-\infty}^{+\infty} (x-\mu)^2f(x)\,dx } \]


27. Uniform on \([0,10]\)

For:

\[ X\sim U(0,10) \]

\[ E(X)=5 \]

and:

\[ \operatorname{Var}(X) = \frac{100}{12} = \frac{25}{3} \]


28. Theoretical mean and sample mean

Theoretical average:

\[ \mu=E(X) \]

Sample average:

\[ \bar x= \frac{x_1+\cdots+x_n}{n} \]

They are different concepts.


29. Complete example

Suppose:

\[ P(X=0)=0,2 \]

\[ P(X=1)=0,5 \]

\[ P(X=3)=0,3 \]

Average:

\[ E(X) = 0+0,5+0,9 = 1,4 \]


30. Calculation of \(E(X^2)\)

\[ E(X^2) = 0+0,5+2,7 = 3,2 \]


31. Variance

\[ \operatorname{Var}(X) = 3,2-(1,4)^2 \]

\[ = 1,24 \]


32. Standard deviation

\[ \sigma=\sqrt{1,24} \approx1,11 \]


33. Exercises

Exercise 1

\[ X=1,2,3 \]

with probability:

\[ 0,2,\ 0,5,\ 0,3 \]

\[ E(X)=2,1 \]

Exercise 2

\[ E(X^2)=4,9 \]

Exercise 3

\[ \operatorname{Var}(X)=0,49 \]

Exercise 4

\[ \sigma=0,7 \]

Exercise 5

If:

\[ E(X)=8 \]

and:

\[ Y=5X-3 \]

then:

\[ E(Y)=37 \]


34. Summary of Chapter 10

\[ \boxed{ E(X)=\sum_xxP(X=x) } \]

\[ \boxed{ \operatorname{Var}(X) = E[(X-\mu)^2] } \]

\[ \boxed{ \operatorname{Var}(X) = E(X^2)-[E(X)]^2 } \]

\[ \boxed{ \sigma= \sqrt{\operatorname{Var}(X)} } \]


Chapter 11 – Discrete distributions

1. Why use standard distributions

Many different experiments share the same mathematical structure.

The main discrete distributions we study are:

\[ \boxed{\text{Bernoulli}} \]

\[ \boxed{\text{Binomiale}} \]

\[ \boxed{\text{Geometrica}} \]

\[ \boxed{\text{Poisson}} \]


2. Bernoulli

A Bernoulli proof has two outcomes:

\[ \text{successo} \]

and:

\[ \text{insuccesso} \]

with probability:

\[ p \]

and:

\[ 1-p \]

Let's define:

\[ X= \begin{cases} 1 & \text{successo}\\ 0 & \text{insuccesso} \end{cases} \]

We write:

\[ \boxed{ X\sim Bernoulli(p) } \]


3. Probability function

\[ P(X=1)=p \]

\[ P(X=0)=1-p \]

Compact shape:

\[ \boxed{ P(X=x) = p^x(1-p)^{1-x} } \]


4. Bernoulli mean and variance

\[ \boxed{ E(X)=p } \]

\[ \boxed{ \operatorname{Var}(X) = p(1-p) } \]


5. From Bernoulli to Binomial

We repeat a Bernoulli test \(n\) times.

Let's define:

\[ X=\text{numero di successi} \]

If:

  • \(n\) is fixed;
  • the tests are independent;
  • the probability \(p\) is constant;

then:

\[ \boxed{ X\sim Bin(n,p) } \]


6. Binomial formula

\[ \boxed{ P(X=k) = \binom nkp^k(1-p)^{n-k} } \]

for:

\[ k=0,\ldots,n \]


7. Why does \(\binom nk\) appear

If we want \(k\) successes in \(n\) trials, we must choose in which positions \(k\) appear.

The number of ways is:

\[ \binom nk \]


8. Example

Balanced coin flipped 5 times.

Probability of 3 heads:

\[ P(X=3) = \binom53 \left(\frac12\right)^5 \]

\[ = \frac5{16} \]


9. Binomial mean and variance

\[ \boxed{ E(X)=np } \]

\[ \boxed{ \operatorname{Var}(X) = np(1-p) } \]

\[ \boxed{ \sigma= \sqrt{np(1-p)} } \]


10. Exactly, at least, at most

Exactly \(k\):

\[ P(X=k) \]

At least \(k\):

\[ P(X\geq k) \]

At most \(k\):

\[ P(X\leq k) \]


11. At least one success

\[ P(X\geq1) = 1-P(X=0) \]

So:

\[ \boxed{ P(X\geq1) = 1-(1-p)^n } \]


12. Example

Probability of defect:

\[ 0,02 \]

20 independent components.

Probability of at least one defective:

\[ 1-(0,98)^{20} \]


13. When NOT to use the binomial

They must be valid:

  • fixed number of tests;
  • two outcomes;
  • same probability;
  • independence.

If we extract without reinsertion, these conditions can fail.


14. Geometric distribution

The question changes:

how many trials do you need to get the first success?

If the evidence is independent and has probability \(p\):

\[ \boxed{ X\sim Geom(p) } \]


15. Geometric formula

\[ \boxed{ P(X=k) = (1-p)^{k-1}p } \]

for:

\[ k=1,2,\ldots \]


16. Example

We roll a die until the first 6.

Probability of reaching the fourth throw:

\[ \left(\frac56\right)^3 \frac16 \]


17. Geometric mean and variance

\[ \boxed{ E(X)=\frac1p } \]

\[ \boxed{ \operatorname{Var}(X) = \frac{1-p}{p^2} } \]


18. Memoryless properties

\[ \boxed{ P(X>m+n\mid X>m) = P(X>n) } \]

The process does not “remember” what we have already waited for.


19. Binomial and geometric

Binomial:

how many successes in \(n\) tests?

Geometric:

how many trials until the first success?


20. Poisson distribution

It is used to count events in a range of:

  • time;
  • space;
  • length;
  • area;
  • volume.

We write:

\[ \boxed{ X\sim Poisson(\lambda) } \]


21. Poisson formula

\[ \boxed{ P(X=k) = e^{-\lambda} \frac{\lambda^k}{k!} } \]


22. Example

On average 3 calls per hour.

\[ X\sim Poisson(3) \]

Probability of 2 calls:

\[ P(X=2) = e^{-3}\frac{3^2}{2!} \]


23. No events

\[ \boxed{ P(X=0)=e^{-\lambda} } \]


24. At least one event

\[ \boxed{ P(X\geq1) = 1-e^{-\lambda} } \]


25. Poisson mean and variance

\[ \boxed{ E(X)=\lambda } \]

\[ \boxed{ \operatorname{Var}(X)=\lambda } \]

\[ \boxed{ \sigma=\sqrt\lambda } \]


26. Change interval

If 12 calls come in per hour, in 15 minutes:

\[ \lambda= 12\cdot\frac14 = 3 \]


27. When to use Poisson

When:

  • we count events;
  • the average rate is stable;
  • the events are sufficiently independent;
  • the range is well defined.

28. Poisson as an approximation of the binomial

If:

\[ n \]

it's big and:

\[ p \]

small, with:

\[ \lambda=np \]

we can use:

\[ \boxed{ Bin(n,p)\approx Poisson(np) } \]


29. Final comparison

Bernoulli

A test.

Binomial

Number of successes in \(n\) tests.

Geometric

Number of trials until first success.

Poisson

Number of events in an interval.


30. Averages

Bernoulli:

\[ p \]

Binomial:

\[ np \]

Geometric:

\[ \frac1p \]

Poisson:

\[ \lambda \]


31. Variances

Bernoulli:

\[ p(1-p) \]

Binomial:

\[ np(1-p) \]

Geometric:

\[ \frac{1-p}{p^2} \]

Poisson:

\[ \lambda \]


32. Summary of Chapter 11

\[ \boxed{ Bernoulli(p) } \]

a single test.

\[ \boxed{ Bin(n,p) } \]

successes in \(n\) tests.

\[ \boxed{ Geom(p) } \]

tests until the first success.

\[ \boxed{ Poisson(\lambda) } \]

events in an interval.


Chapter 12 – Continuous distributions: uniform, exponential and normal

1. Continuous variables

For a continuous variable:

\[ P(X=x)=0 \]

Probabilities concern ranges:

\[ \boxed{ P(a\leq X\leq b) = \int_a^bf(x)\,dx } \]


2. Continuous uniform distribution

If all values of:

\[ [a,b] \]

are equally plausible:

\[ \boxed{ X\sim U(a,b) } \]


3. Uniform density

\[ \boxed{ f(x)= \frac1{b-a} } \]

for:

\[ a\leq x\leq b \]


4. Uniform probability

\[ \boxed{ P(c\leq X\leq d) = \frac{d-c}{b-a} } \]


5. Example

\[ X\sim U(0,10) \]

\[ P(2\leq X\leq5) = \frac3{10} = 0,3 \]


6. Uniform mean and variance

\[ \boxed{ E(X)=\frac{a+b}{2} } \]

\[ \boxed{ \operatorname{Var}(X) = \frac{(b-a)^2}{12} } \]


7. Exponential distribution

The exponential distribution mainly models waiting times.

We write:

\[ \boxed{ X\sim Exp(\lambda) } \]


8. Exponential density

\[ \boxed{ f(x)=\lambda e^{-\lambda x} } \]

for:

\[ x\geq0 \]


9. Distribution function

\[ \boxed{ F(x)=1-e^{-\lambda x} } \]


10. Chance of passing \(x\)

\[ \boxed{ P(X>x)=e^{-\lambda x} } \]


11. Example

\[ X\sim Exp(0,2) \]

\[ P(X>5) = e^{-1} \approx0,3679 \]


12. Probability on an interval

\[ \boxed{ P(a<X<b) = e^{-\lambda a} - e^{-\lambda b} } \]


13. Exponential mean and variance

\[ \boxed{ E(X)=\frac1\lambda } \]

\[ \boxed{ \operatorname{Var}(X) = \frac1{\lambda^2} } \]

\[ \boxed{ \sigma=\frac1\lambda } \]


14. Connection with Poisson

Poisson:

how many events?

Exponential:

how long between events?

So:

\[ \boxed{ Poisson\leftrightarrow conteggio } \]

\[ \boxed{ Esponenziale\leftrightarrow attesa } \]


15. No memory

\[ \boxed{ P(X>s+t\mid X>s) = P(X>t) } \]


16. Normal distribution

The normal, or Gaussian, distribution is indicated:

\[ \boxed{ X\sim N(\mu,\sigma^2) } \]


17. Density of the normal

\[ \boxed{ f(x) = \frac1{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} } \]


18. Properties of the normal

The curve:

  • it is symmetrical;
  • has a bell shape;
  • is centered in \(\mu\);
  • has total area 1.

19. Mean, median and mode

For normal:

\[ \boxed{ \text{media}= \text{mediana}= \text{moda} = \mu } \]


20. Role of \(\mu\)

\(\mu\) determines the position of the curve.

By changing \(\mu\), the curve shifts.


21. Role of \(\sigma\)

\(\sigma\) determines the dispersion.

Small \(\sigma\):

\[ \text{curva stretta} \]

Large \(\sigma\):

\[ \text{curva larga} \]


22. Symmetry

\[ P(X<\mu)=\frac12 \]

\[ P(X>\mu)=\frac12 \]


23. Rule 68–95–99.7

For a normal one:

\[ P(\mu-\sigma<X<\mu+\sigma) \approx68\% \]

\[ P(\mu-2\sigma<X<\mu+2\sigma) \approx95\% \]

\[ P(\mu-3\sigma<X<\mu+3\sigma) \approx99,7\% \]


24. Standardization

Let's define:

\[ \boxed{ Z= \frac{X-\mu}{\sigma} } \]

If:

\[ X\sim N(\mu,\sigma^2) \]

then:

\[ \boxed{ Z\sim N(0,1) } \]


25. Meaning of the z-score

\[ z= \frac{x-\mu}{\sigma} \]

indicates how many standard deviations a value is above or below the mean.


26. Example

\[ \mu=100 \]

\[ \sigma=15 \]

\[ x=130 \]

\[ z= \frac{130-100}{15} = 2 \]

The value is two standard deviations above the mean.


27. Another example

\[ x=85 \]

\[ z= \frac{85-100}{15} = -1 \]

One standard deviation below the mean.


28. Function \(\Phi\)

For standard normal:

\[ \boxed{ \Phi(z)=P(Z\leq z) } \]


29. Symmetry of the standard normal

\[ \boxed{ \Phi(-z)=1-\Phi(z) } \]


30. Odds to the right

\[ \boxed{ P(Z>z)=1-\Phi(z) } \]


31. Probability between two values

\[ \boxed{ P(a<Z<b) = \Phi(b)-\Phi(a) } \]


32. Complete example

\[ X\sim N(100,15^2) \]

Calculate:

\[ P(X<115) \]

We standardize:

\[ z=1 \]

So:

\[ P(X<115) = \Phi(1) \approx0,8413 \]


33. Higher probability

\[ P(X>130) \]

\[ z=2 \]

\[ P(Z>2) = 1-\Phi(2) \]

\[ \approx0,0228 \]


34. Central probability

\[ P(85<X<115) \]

corresponds to:

\[ P(-1<Z<1) \]

\[ \approx0,6827 \]


35. Reverse lookup

If:

\[ \Phi(z)=0,95 \]

then:

\[ z\approx1,645 \]


36. Important critical values

90% central:

\[ z\approx1,645 \]

95% central:

\[ z\approx1,96 \]

99% central:

\[ z\approx2,576 \]


37. Why 1.96 and not 2

The exact central 95% is approximately:

\[ P(-1,96<Z<1,96)=0,95 \]

The value 2 is a convenient approximation.


38. Percentiles

The \(p\) percentile satisfies:

\[ P(X\leq x_p)=p \]

For a normal one:

\[ \boxed{ x_p= \mu+z_p\sigma } \]


39. Example

\[ X\sim N(100,15^2) \]

95th percentile:

\[ z_{0,95}\approx1,645 \]

\[ x_{0,95} = 100+1,645\cdot15 \]

\[ \approx124,7 \]


40. Central 95% and 95th percentile

They are not the same thing.

95% central:

\[ -1,96<Z<1,96 \]

95th percentile:

\[ Z\leq1,645 \]


41. When the normal is plausible

Normal often appears for:

  • measurement errors;
  • biological characteristics;
  • physical measurements;
  • sums of many small effects;
  • sample averages.

But not all phenomena are normal.


42. Comparison of continuous distributions

Uniform

Equiplausible values in an interval.

Exponential

Waiting times.

Normal

Symmetric bell-shaped phenomena.


43. Exercises

Exercise 1

\[ X\sim U(10,20) \]

\[ P(12<X<16) = \frac4{10} = 0,4 \]

Exercise 2

\[ E(X)=15 \]

Exercise 3

\[ X\sim Exp(0,5) \]

\[ P(X>4) = e^{-2} \approx0,1353 \]

Exercise 4

Average time:

\[ E(X)=\frac1{0,5}=2 \]

Exercise 5

\[ X\sim N(50,10^2) \]

for:

\[ x=70 \]

we have:

\[ z=2 \]


44. Summary of Chapter 12

Uniform:

\[ \boxed{ X\sim U(a,b) } \]

\[ E(X)=\frac{a+b}{2} \]

\[ \operatorname{Var}(X) = \frac{(b-a)^2}{12} \]

Exponential:

\[ \boxed{ X\sim Exp(\lambda) } \]

\[ P(X>x)=e^{-\lambda x} \]

\[ E(X)=\frac1\lambda \]

Normal:

\[ \boxed{ X\sim N(\mu,\sigma^2) } \]

and standardization:

\[ \boxed{ Z= \frac{X-\mu}{\sigma} } \]

Try the models: open the distributions laboratory.

Chapter 13 – Law of Large Numbers and Central Limit Theorem

1. Two different but related results

When we repeat a random experiment many times two fundamental things happen:

  • the observed average tends to stabilize;
  • the distribution of the mean often tends to take on a normal shape.

These phenomena are described by:

\[ \boxed{\text{Legge dei Grandi Numeri}} \]

and:

\[ \boxed{\text{Teorema Centrale del Limite}} \]

The Law of Large Numbers concerns towards which value the average tends.

The Central Limit Theorem concerns how the mean is distributed.


2. Sample mean

Let's consider:

\[ X_1,X_2,\ldots,X_n \]

independent and with the same distribution.

The sample mean is:

\[ \boxed{ \bar X= \frac{X_1+\cdots+X_n}{n} } \]


3. Expected value of the mean

If:

\[ E(X_i)=\mu \]

then:

\[ E(\bar X) = \frac1n \sum_{i=1}^nE(X_i) \]

\[ = \frac{n\mu}{n} \]

therefore:

\[ \boxed{ E(\bar X)=\mu } \]


4. Variance of the mean

If:

\[ \operatorname{Var}(X_i)=\sigma^2 \]

and the observations are independent:

\[ \operatorname{Var}(\bar X) = \frac{\sigma^2}{n} \]

therefore:

\[ \boxed{ \operatorname{Var}(\bar X) = \frac{\sigma^2}{n} } \]


5. Standard error

The standard deviation of the sample mean is:

\[ \boxed{ SE(\bar X) = \frac{\sigma}{\sqrt n} } \]

This quantity is called standard error of the mean.


6. Consequence

Increasing \(n\):

\[ \frac{\sigma}{\sqrt n} \]

decreases.

The sample averages therefore become increasingly concentrated around:

\[ \mu \]


7. Example

Suppose:

\[ \mu=100 \]

\[ \sigma=20 \]

With:

\[ n=25 \]

we have:

\[ SE=\frac{20}{5}=4 \]

With:

\[ n=100 \]

we have:

\[ SE=\frac{20}{10}=2 \]


8. Law of Large Numbers

The Law of Large Numbers states, intuitively, that:

\[ \boxed{ \bar X_n\longrightarrow\mu } \]

when:

\[ n\to\infty \]

The observed average tends to approach the expected value.


9. Example with a die

For a regular nut:

\[ E(X)=3,5 \]

After a few launches the average can be very different.

After many launches it will instead tend to stabilize close to:

\[ 3,5 \]


10. Relative frequency

Let's consider an event \(A\) and define:

\[ X_i= \begin{cases} 1 & \text{se }A\text{ si verifica}\\ 0 & \text{altrimenti} \end{cases} \]

Then:

\[ E(X_i)=P(A) \]

The average:

\[ \bar X = \frac{X_1+\cdots+X_n}{n} \]

coincides with the relative frequency.

So:

\[ \boxed{ f_n(A)\to P(A) } \]


11. What the Law of Large Numbers does NOT say

It doesn't say that:

after many tails it must come up heads.

If the launches are independent:

\[ P(T)=\frac12 \]

with every launch.

The coin does not “catch up” to previous results.


12. Probabilistic version

An important formulation is:

\[ \boxed{ P(|\bar X_n-\mu|>\varepsilon) \to0 } \]

for:

\[ n\to\infty \]


13. Central Limit Theorem

If:

\[ X_1,\ldots,X_n \]

are independent and identically distributed with:

\[ E(X_i)=\mu \]

and:

\[ \operatorname{Var}(X_i)=\sigma^2<\infty \]

then for large \(n\):

\[ \boxed{ \bar X \approx N\left( \mu, \frac{\sigma^2}{n} \right) } \]


14. Standardized form

\[ \boxed{ Z= \frac{ \bar X-\mu }{ \sigma/\sqrt n } \approx N(0,1) } \]

This formula will be critical for statistical inference.


15. The surprising result

The initial population can be:

  • uniform;
  • exponential;
  • discreet;
  • asymmetrical.

The distribution of the sample mean tends to be normal, under appropriate conditions.


16. Attention

The TCL does not say that the original data becomes normal.

It says that the distribution of:

\[ \boxed{\bar X} \]


17. Sampling distribution of the mean

The distribution of possible values of:

\[ \bar X \]

has:

\[ E(\bar X)=\mu \]

and:

\[ \operatorname{Var}(\bar X)=\frac{\sigma^2}{n} \]


18. Increasing \(n\)

Two things happen:

  1. the distribution of the mean tends towards normal;
  2. dispersion decreases.

19. Special case: normal population

If:

\[ X\sim N(\mu,\sigma^2) \]

then:

\[ \boxed{ \bar X \sim N\left( \mu, \frac{\sigma^2}{n} \right) } \]

for any \(n\).


20. Example

Suppose:

\[ X\sim N(175,10^2) \]

and:

\[ n=25 \]

Then:

\[ \bar X \sim N\left( 175, \frac{100}{25} \right) \]

therefore:

\[ \bar X\sim N(175,4) \]

and:

\[ SE=2 \]


21. Probability on average

Let's calculate:

\[ P(\bar X>178) \]

We standardize:

\[ Z= \frac{178-175}{2} = 1,5 \]

So:

\[ P(\bar X>178) = P(Z>1,5) \approx0,0668 \]


22. Single observation and average

For a single observation:

\[ \sigma_X=\sigma \]

For an average:

\[ \boxed{ \sigma_{\bar X} = \frac{\sigma}{\sqrt n} } \]

The average is therefore much less variable.


23. To halve the standard error

Because:

\[ SE\propto\frac1{\sqrt n} \]

to halve it you need to quadruple \(n\).


24. TCL for the sum

If:

\[ S_n=X_1+\cdots+X_n \]

then:

\[ E(S_n)=n\mu \]

\[ \operatorname{Var}(S_n)=n\sigma^2 \]

and:

\[ \boxed{ S_n \approx N(n\mu,n\sigma^2) } \]


25. Standardization of the sum

\[ \boxed{ Z= \frac{ S_n-n\mu }{ \sigma\sqrt n } \approx N(0,1) } \]


26. Connection with the binomial

If:

\[ X\sim Bin(n,p) \]

then:

\[ E(X)=np \]

\[ \operatorname{Var}(X)=np(1-p) \]

For \(n\) large:

\[ \boxed{ X \approx N(np,np(1-p)) } \]


27. Continuity fix

Since the binomial is discrete and the normal is continuous:

\[ P(X\leq10) \]

is approximated with:

\[ P(Y<10,5) \]

and:

\[ P(X=10) \]

with:

\[ P(9,5<Y<10,5) \]


28. Comparison of LGN and TCL

Law of Large Numbers:

\[ \boxed{ \bar X\to\mu } \]

Central Limit Theorem:

\[ \boxed{ \frac{\bar X-\mu}{\sigma/\sqrt n} \approx N(0,1) } \]


29. Final idea

We can remember:

\[ \boxed{ \text{LGN: la media va verso }\mu } \]

\[ \boxed{ \text{TCL: la media si distribuisce come una normale} } \]

These results are the basis of inferential statistics.


Chapter 14 – Sampling, point estimate and confidence intervals

1. From probability to statistics

In probability we start from a known model.

In statistics we start from the data and we want to know the population.

Scheme:

\[ \boxed{ \text{Popolazione} \rightarrow \text{Campione} \rightarrow \text{Inferenza} } \]


2. Population and sample

The population is the complete set of elements studied.

The sample is an observed subset.


3. Parameters and statistics

Population parameters:

\[ \mu,\quad\sigma^2,\quad p \]

Sample statistics:

\[ \bar X,\quad S^2,\quad\hat p \]


4. Punctual estimate

If \(\mu\) is unknown, we can estimate it with:

\[ \boxed{ \hat\mu=\bar X } \]

For example, if:

\[ \bar x=175,4 \]

we estimate:

\[ \mu\approx175,4 \]


5. Because it's not enough

An estimate calculated from 5 observations is less precise than the same estimate calculated from 500 observations.

It is therefore necessary to quantify the uncertainty.


6. Distribution of the mean

We know that:

\[ E(\bar X)=\mu \]

and:

\[ SE(\bar X) = \frac{\sigma}{\sqrt n} \]


7. Confidence interval with \(\sigma\) known

If:

\[ Z= \frac{ \bar X-\mu }{ \sigma/\sqrt n } \]

and:

\[ P(-1,96<Z<1,96)\approx0,95 \]

then:

\[ \boxed{ \mu \in \bar X \pm 1,96\frac{\sigma}{\sqrt n} } \]

at the 95% confidence level.


8. General form

\[ \boxed{ \bar X \pm z_{\alpha/2} \frac{\sigma}{\sqrt n} } \]

where:

\[ 1-\alpha \]

is the confidence level.


9. Critical values

90%:

\[ z\approx1,645 \]

95%:

\[ z\approx1,96 \]

99%:

\[ z\approx2,576 \]


10. Complete example

Suppose:

\[ \sigma^2=10,66 \]

\[ n=58 \]

\[ \bar x=175,4 \]

Then:

\[ \sigma=\sqrt{10,66}\approx3,265 \]

and:

\[ SE= \frac{3,265}{\sqrt{58}} \approx0,429 \]


11. 90% range

\[ E= 1,645\cdot0,429 \approx0,706 \]

So:

\[ 175,4\pm0,706 \]

i.e. approximately:

\[ \boxed{ [174,69,\ 176,11] } \]


12. 95% range

\[ E= 1,96\cdot0,429 \approx0,841 \]

So:

\[ \boxed{ [174,56,\ 176,24] } \]


13. 99% range

\[ E= 2,576\cdot0,429 \approx1,105 \]

So:

\[ \boxed{ [174,30,\ 176,51] } \]


14. Confidence level and breadth

As the confidence level is increased, the interval becomes wider.

The more conservative we want to be, the more we have to accept a wide range.


15. Margin for error

\[ \boxed{ E= z_{\alpha/2} \frac{\sigma}{\sqrt n} } \]

The range is:

\[ \boxed{ \bar X\pm E } \]


16. What it depends on

The margin depends on:

\[ z_{\alpha/2} \]

\[ \sigma \]

\[ n \]


17. Effect of \(n\)

If \(n\) increases, the range narrows.

To halve the margin you need to quadruple the sample.


18. 95% interpretation

Formally, the parameter:

\[ \mu \]

it is fixed.

It is the intervals that change from sample to sample.

If we repeated the sampling many times, about 95% of the constructed intervals will contain \(\mu\).


19. When \(\sigma\) is unknown

In practice \(\sigma\) is often not known.

We estimate it with:

\[ s \]

and we use the distribution:

\[ \boxed{\text{t di Student}} \]


20. Interval with t

\[ \boxed{ \bar X \pm t_{\alpha/2,n-1} \frac{s}{\sqrt n} } \]

The degrees of freedom are:

\[ \boxed{ n-1 } \]


21. Example

Suppose:

\[ n=25 \]

\[ \bar x=50 \]

\[ s=10 \]

For 95%, with 24 degrees of freedom:

\[ t\approx2,064 \]

Standard error:

\[ \frac{10}{5}=2 \]

Margin:

\[ 2,064\cdot2=4,128 \]

Range:

\[ \boxed{ [45,872,\ 54,128] } \]


22. Sample variance

\[ \boxed{ S^2= \frac{ \sum_{i=1}^{n}(X_i-\bar X)^2 }{ n-1 } } \]

The denominator \(n-1\) makes \(S^2\) an unbiased estimator of \(\sigma^2\).


23. Interval for a proportion

If:

\[ \hat p=\frac Xn \]

then, for large samples:

\[ \boxed{ \hat p \pm z_{\alpha/2} \sqrt{ \frac{\hat p(1-\hat p)}{n} } } \]


24. Example

Out of 400 people:

\[ 240 \]

they answer yes.

\[ \hat p=\frac{240}{400}=0,6 \]

Standard error:

\[ SE= \sqrt{ \frac{0,6\cdot0,4}{400} } \approx0,0245 \]

Margin at 95%:

\[ 1,96\cdot0,0245 \approx0,048 \]

Range:

\[ \boxed{ [0,552,\ 0,648] } \]


25. Correct sampling

A large but biased sample can produce a very precise estimate of the wrong value.

The quality of sampling is critical.


26. Sampling and systematic error

Sampling error:

  • due to chance;
  • decreases by increasing \(n\).

Systematic error:

  • non-representative sample;
  • measurement errors;
  • biased selection.

Increasing \(n\) does not eliminate it.


27. Sample size for the mean

From:

\[ E= z_{\alpha/2} \frac{\sigma}{\sqrt n} \]

we get:

\[ \boxed{ n= \left( \frac{ z_{\alpha/2}\sigma }{ E } \right)^2 } \]


28. Example

We want:

\[ \sigma=12 \]

95% confidence,

\[ E=2 \]

So:

\[ n= \left( \frac{1,96\cdot12}{2} \right)^2 \approx138,3 \]

You need at least:

\[ \boxed{139} \]

observations.


29. Sample size for a proportion

\[ \boxed{ n= \frac{ z_{\alpha/2}^2p(1-p) }{ E^2 } } \]

If \(p\) is unknown we often use:

\[ p=0,5 \]

because it maximizes:

\[ p(1-p) \]


30. Summary of Chapter 14

Average with \(\sigma\) note:

\[ \boxed{ \bar x \pm z_{\alpha/2} \frac{\sigma}{\sqrt n} } \]

Average with unknown \(\sigma\):

\[ \boxed{ \bar x \pm t_{\alpha/2,n-1} \frac{s}{\sqrt n} } \]

Proportion:

\[ \boxed{ \hat p \pm z_{\alpha/2} \sqrt{ \frac{\hat p(1-\hat p)}{n} } } \]


From sample to interval: follow the guided laboratory.

Chapter 15 – Hypothesis testing and p-values

1. From estimate to verification

Suppose someone states:

the population mean is 100.

We want to understand if the observed data are compatible with this statement.

This is how hypothesis tests are born.


2. Null and alternative hypotheses

Let's define:

\[ \boxed{H_0} \]

null hypothesis,

and:

\[ \boxed{H_1} \]

alternative hypothesis.

Example:

\[ H_0:\mu=100 \]

\[ H_1:\mu\neq100 \]


3. Test principle

The reasoning is:

  1. We assume true \(H_0\);
  2. let's see how plausible the data is;
  3. if they are too extreme, we reject \(H_0\).

4. Test statistics

For an average with \(\sigma\) note:

\[ \boxed{ Z= \frac{ \bar X-\mu_0 }{ \sigma/\sqrt n } } \]


5. Interpretation

\(Z\) measures how many standard error units separate:

\[ \bar X \]

from:

\[ \mu_0 \]


6. Example

\[ \mu_0=100 \]

\[ \bar x=104 \]

\[ \sigma=12 \]

\[ n=36 \]

Standard error:

\[ SE=2 \]

So:

\[ Z= \frac{104-100}{2} = 2 \]


7. Bilateral test

If:

\[ H_1:\mu\neq\mu_0 \]

the test is bilateral.

Both queues are relevant.


8. Right unilateral test

If:

\[ H_1:\mu>\mu_0 \]

we look for large values of the test statistic.


9. Left unilateral test

If:

\[ H_1:\mu<\mu_0 \]

we look for very negative values.


10. Level of significance

We choose:

\[ \alpha \]

Common values:

\[ 0,10,\quad0,05,\quad0,01 \]


11. Type I error

It consists of rejecting \(H_0\) when it is true.

The probability is:

\[ \boxed{\alpha} \]


12. Critical region

In a 5% two-sided test:

\[ |Z|>1,96 \]

leads to rejection of \(H_0\).


13. One-sided test

At 5%:

\[ Z>1,645 \]

for a right test,

or:

\[ Z<-1,645 \]

for a sinister test.


14. p-value

The:

\[ \boxed{p\text{-value}} \]

is the probability, assuming true \(H_0\), of observing a result at least as extreme as the one obtained.


15. Decision rule

If:

\[ p<\alpha \]

we reject \(H_0\).

If:

\[ p\geq\alpha \]

we do not reject \(H_0\).


16. Not rejecting does not mean accepting

Do not reject \(H_0\) only means:

the data does not provide sufficient evidence against \(H_0\).

It does not mean that you have proven that \(H_0\) is true.


17. Complete example

We have:

\[ Z=2 \]

In a 5% two-sided test:

\[ |2|>1,96 \]

therefore we reject \(H_0\).

The p-value is:

\[ p= 2P(Z\geq2) \]

\[ \approx0,0456 \]

Because:

\[ 0,0456<0,05 \]

we reject \(H_0\).


18. At the 1% level

\[ 0,0456>0,01 \]

so we don't reject \(H_0\).


19. Statistical significance

If:

\[ p<\alpha \]

let's say that the result is statistically significant.

But:

\[ \boxed{ \text{significativo} \neq \text{importante} } \]


20. Type II error

It consists of not rejecting \(H_0\) when it is false.

Its probability is indicated by:

\[ \boxed{\beta} \]


21. Test power

\[ \boxed{ 1-\beta } \]

is the probability of correctly rejecting \(H_0\) when it is false.


22. How potency increases

It generally increases when:

  • increase \(n\);
  • increases the real effect;
  • decreases variability;
  • increases \(\alpha\).

23. t-test

If \(\sigma\) is unknown:

\[ \boxed{ T= \frac{ \bar X-\mu_0 }{ s/\sqrt n } } \]

with:

\[ n-1 \]

degrees of freedom.


24. Test on a proportion

For:

\[ H_0:p=p_0 \]

we use:

\[ \boxed{ Z= \frac{ \hat p-p_0 }{ \sqrt{ p_0(1-p_0)/n } } } \]


25. Proportion example

A company declares:

\[ p=0,60 \]

On:

\[ 400 \]

satisfied customers:

\[ 260 \]

therefore:

\[ \hat p=0,65 \]

The statistics are approximately:

\[ Z\approx2,04 \]

Because:

\[ 2,04>1,96 \]

we reject \(H_0\) at 5%.


26. Binding with confidence interval

For a bilateral test at the level:

\[ \alpha \]

we refuse:

\[ H_0:\mu=\mu_0 \]

if:

\[ \mu_0 \]

does not belong to the confidence interval:

\[ 1-\alpha \]


27. Example

If the 95% interval is:

\[ [101,2,\ 106,8] \]

the value:

\[ 100 \]

does not belong to the range.

So the 5% test rejects:

\[ H_0:\mu=100 \]


28. p-value: common error

If:

\[ p=0,04 \]

does not mean:

\(H_0\) has a 4% chance of being true.

The p-value is calculated assuming \(H_0\) true.


29. Practical example

A manufacturer states:

\[ \mu=500\text{ g} \]

We want to check:

\[ H_1:\mu<500 \]

Suppose:

\[ n=64 \]

\[ \bar x=497 \]

\[ \sigma=12 \]

Then:

\[ Z= \frac{497-500}{12/8} = -2 \]

Because:

\[ -2<-1,645 \]

we reject \(H_0\) at 5%.


30. Multiple tests

If we do a lot of testing, the risk of false positives increases.

A simple fix is Bonferroni:

\[ \boxed{ \alpha^\ast= \frac{\alpha}{m} } \]

where \(m\) is the test number.


31. Effect size

A small p-value does not necessarily imply a large effect.

It is also useful to consider:

  • observed difference;
  • confidence interval;
  • measurement of the effect.

32. General procedure

  1. define the parameter;
  2. write \(H_0\) and \(H_1\);
  3. choose \(\alpha\);
  4. calculate the test statistic;
  5. calculate the p-value;
  6. make the decision;
  7. interpret in context.

33. Summary of Chapter 15

Z-test:

\[ \boxed{ Z= \frac{ \bar X-\mu_0 }{ \sigma/\sqrt n } } \]

t-test:

\[ \boxed{ T= \frac{ \bar X-\mu_0 }{ s/\sqrt n } } \]

Proportion test:

\[ \boxed{ Z= \frac{ \hat p-p_0 }{ \sqrt{ p_0(1-p_0)/n } } } \]

Type I error:

\[ \alpha \]

Type II error:

\[ \beta \]

Power:

\[ 1-\beta \]


Chapter 16 – Covariance, correlation and linear regression

1. From one variable to two variables

We often want to study two variables at the same time.

Examples:

  • height and weight;
  • hours of study and grade;
  • temperature and consumption;
  • advertising and sales.

The question is:

is there a relationship between \(X\) and \(Y\)?


2. Pairs of observations

We observe:

\[ (x_1,y_1), (x_2,y_2), \ldots, (x_n,y_n) \]

Each observation is a pair.


3. Scatter diagram

The scatter plot represents:

\[ X \]

on the horizontal axis and:

\[ Y \]

on the vertical one.

It can show:

  • positive relationship;
  • negative;
  • absence of linear relationship;
  • curved relationship;
  • outliers.

4. Covariance

The covariance is:

\[ \boxed{ \operatorname{Cov}(X,Y) = E[(X-\mu_X)(Y-\mu_Y)] } \]


5. Alternative formula

\[ \boxed{ \operatorname{Cov}(X,Y) = E(XY)-E(X)E(Y) } \]


6. Interpretation

If:

\[ \operatorname{Cov}(X,Y)>0 \]

the variables tend to move in the same direction.

If:

\[ \operatorname{Cov}(X,Y)<0 \]

they tend to move in opposite directions.


7. Sample covariance

\[ \boxed{ s_{XY} = \frac{ \sum_{i=1}^{n} (x_i-\bar x)(y_i-\bar y) }{ n-1 } } \]


8. Limit of covariance

It depends on the units of measurement.

This is why it is difficult to directly compare their size.


9. Correlation

Let's standardize the covariance:

\[ \boxed{ \rho= \frac{ \operatorname{Cov}(X,Y) }{ \sigma_X\sigma_Y } } \]

In the sample:

\[ \boxed{ r= \frac{ s_{XY} }{ s_Xs_Y } } \]


10. Interval

\[ \boxed{ -1\leq r\leq1 } \]


11. Interpretation

\[ r\approx1 \]

strong positive linear relationship.

\[ r\approx-1 \]

strong negative linear relationship.

\[ r\approx0 \]

weak or absent linear relationship.


12. Zero correlation does not mean independence

A non-linear relationship can also exist when:

\[ r=0 \]

For example:

\[ Y=X^2 \]

with \(X\) symmetric around zero.


13. Correlation does not imply causation

\[ \boxed{ \text{correlazione}\neq\text{causalità} } \]

Two variables can be correlated due to the effect of a third variable.


14. Linear regression

The simplest model is:

\[ \boxed{ Y=a+bX+\varepsilon } \]

where:

  • \(a\) = intercept;
  • \(b\) = angular coefficient;
  • \(\varepsilon\) = error.

15. Estimated tuition

\[ \boxed{ \hat Y=a+bX } \]


16. Meaning of \(b\)

The coefficient:

\[ b \]

indicates how much \(Y\) varies on average when \(X\) increases by one unit.


17. Residue

\[ \boxed{ e_i=y_i-\hat y_i } \]

It is the difference between observed value and predicted value.


18. Least squares

The straight line is chosen by minimizing:

\[ \boxed{ \sum_{i=1}^{n}(y_i-\hat y_i)^2 } \]


19. Slope

\[ \boxed{ b= \frac{ \sum(x_i-\bar x)(y_i-\bar y) }{ \sum(x_i-\bar x)^2 } } \]

equivalently:

\[ \boxed{ b= \frac{s_{XY}}{s_X^2} } \]


20. Intercept

\[ \boxed{ a= \bar y-b\bar x } \]

The line always passes through:

\[ (\bar x,\bar y) \]


21. Connection with correlation

\[ \boxed{ b= r\frac{s_Y}{s_X} } \]

The sign of \(b\) coincides with the sign of \(r\).


22. Example

Suppose:

\[ \bar x=4 \]

\[ \bar y=10 \]

\[ s_X=2 \]

\[ s_Y=6 \]

\[ r=0,5 \]

Then:

\[ b= 0,5\cdot\frac62 = 1,5 \]

and:

\[ a= 10-1,5\cdot4 = 4 \]

So:

\[ \boxed{ \hat Y=4+1,5X } \]


23. Prediction

For:

\[ X=6 \]

we have:

\[ \hat Y= 4+1,5\cdot6 = 13 \]


24. Interpolation and extrapolation

Interpolation:

forecast within the observed range.

Extrapolation:

forecast outside the observed range.

Extrapolation is riskier.


25. Coefficient of determination

In simple linear regression with intercept:

\[ \boxed{ R^2=r^2 } \]


26. Interpretation

If:

\[ R^2=0,64 \]

approximately 64% of the variability of \(Y\) is described by the linear model.


27. Outliers

An extreme value can strongly modify:

  • \(r\);
  • slope;
  • intercept.

It must be analyzed, not automatically eliminated.


28. Residuals

A good linear model should produce residuals:

  • centered around zero;
  • without pattern;
  • with fairly constant dispersion.

29. Homoscedasticity

If the error dispersion is constant:

\[ \boxed{\text{omoschedasticità}} \]

If it changes to \(X\):

\[ \boxed{\text{eteroschedasticità}} \]


30. Slope test

We can verify:

\[ H_0:b=0 \]

against:

\[ H_1:b\neq0 \]

with a statistic like:

\[ \boxed{ t= \frac{\hat b}{SE(\hat b)} } \]


31. Example of study hours

Suppose:

\[ \hat Y=18+4X \]

Slope 4 means:

every additional hour of study is associated with an average increase of 4 points.

It doesn't necessarily mean causation.


32. Confounding variables

The observed relationship can be influenced by:

  • initial preparation;
  • motivation;
  • difficulty;
  • other factors.

33. Exercises

Exercise 1

If:

\[ \operatorname{Cov}(X,Y)=12 \]

\[ \sigma_X=3 \]

\[ \sigma_Y=8 \]

then:

\[ r= \frac{12}{24} = 0,5 \]


Exercise 2

If:

\[ r=-0,7 \]

we have a fairly strong negative linear relationship.


Exercise 3

If:

\[ r=0 \]

we cannot conclude independence.


Exercise 4

If:

\[ \hat Y=12+3X \]

for:

\[ X=4 \]

we get:

\[ \hat Y=24 \]


Exercise 5

If:

\[ r=0,9 \]

then:

\[ R^2=0,81 \]

that is:

\[ 81\% \]


34. Summary of Chapter 16

Covariance:

\[ \boxed{ \operatorname{Cov}(X,Y) = E[(X-\mu_X)(Y-\mu_Y)] } \]

Correlation:

\[ \boxed{ \rho= \frac{ \operatorname{Cov}(X,Y) }{ \sigma_X\sigma_Y } } \]

Fee:

\[ \boxed{ \hat Y=a+bX } \]

Slope:

\[ \boxed{ b= \frac{s_{XY}}{s_X^2} = r\frac{s_Y}{s_X} } \]

Intercept:

\[ \boxed{ a=\bar y-b\bar x } \]

Coefficient of determination:

\[ \boxed{ R^2=r^2 } \]

Fundamental idea:

\[ \boxed{ \text{correlazione non implica causalità} } \]

Chapter 17 – Hypergeometric and negative binomial distributions

1. Two new discrete models

There are two situations that do not fit perfectly into the models already studied:

  1. drawing without replacement;
  2. repeating trials until the \(r\)-th success.

These situations lead, respectively, to:

\[ \boxed{\text{distribuzione ipergeometrica}} \]

and:

\[ \boxed{\text{distribuzione binomiale negativa}} \]


2. Why the binomial model is not enough

The binomial model requires:

  • a constant probability of success;
  • independent trials.

If we draw without replacement, the composition of the population changes.

Thus the probability of success changes too.


3. Hypergeometric distribution

Suppose we have:

\[ N \]

elements in total, of which:

\[ K \]

are successes.

We draw:

\[ n \]

elements without replacement.

Define:

\[ X=\text{numero di successi estratti} \]

Then:

\[ \boxed{ X\sim Hyp(N,K,n) } \]


4. Formula

\[ \boxed{ P(X=k) = \frac{ \binom Kk \binom{N-K}{n-k} }{ \binom Nn } } \]


5. Interpretation

The denominator:

\[ \binom Nn \]

counts all possible samples.

The numerator counts:

  • \(k\) successes among \(K\);
  • \(n-k\) failures among \(N-K\).

6. Example

An urn contains:

  • 6 red balls;
  • 4 blue balls.

We draw 3 balls without replacement.

The probability of exactly 2 red balls is:

\[ P(X=2) = \frac{ \binom62\binom41 }{ \binom{10}{3} } \]

\[ = \frac{15\cdot4}{120} = \frac12 \]


7. Expected value

\[ \boxed{ E(X)=n\frac KN } \]


8. Variance

\[ \boxed{ \operatorname{Var}(X) = n\frac KN \left(1-\frac KN\right) \frac{N-n}{N-1} } \]

The factor:

\[ \frac{N-n}{N-1} \]

is the finite-population correction.


9. Comparison with the binomial distribution

With replacement, or with a very large population:

\[ \boxed{\text{Binomiale}} \]

Without replacement from a finite population:

\[ \boxed{\text{Ipergeometrica}} \]


10. Card example

We draw 5 cards from a deck of 52.

The probability of exactly 2 aces is:

\[ \boxed{ P(X=2) = \frac{ \binom42\binom{48}{3} }{ \binom{52}{5} } } \]


11. At least one success

\[ P(X\geq1) = 1-P(X=0) \]

Therefore:

\[ \boxed{ P(X\geq1) = 1- \frac{ \binom{N-K}{n} }{ \binom Nn } } \]


12. Negative binomial distribution

The geometric distribution asks:

how many trials until the first success?

The negative binomial distribution asks:

how many trials until the \(r\)-th success?


13. Definition

Consider independent Bernoulli trials with probability of success:

\[ p \]

Define:

\[ X=\text{numero di prove necessarie per ottenere }r\text{ successi} \]

Then:

\[ \boxed{ X\sim NB(r,p) } \]


14. Formula

\[ \boxed{ P(X=k) = \binom{k-1}{r-1} p^r(1-p)^{k-r} } \]

where:

\[ k=r,r+1,\ldots \]


15. Why it works

For the \(r\)-th success to occur on trial \(k\):

  • among the first \(k-1\) trials there must be \(r-1\) successes;
  • trial \(k\) must be a success.

16. Example

A fair coin.

What is the probability that the third head occurs on the fifth toss?

\[ P(X=5) = \binom42 \left(\frac12\right)^3 \left(\frac12\right)^2 \]

\[ = \frac3{16} \]


17. Mean and variance

\[ \boxed{ E(X)=\frac rp } \]

\[ \boxed{ \operatorname{Var}(X) = \frac{r(1-p)}{p^2} } \]


18. The geometric distribution as a special case

If:

\[ r=1 \]

we obtain:

\[ \boxed{ Geom(p)=NB(1,p) } \]


19. Final comparison

Binomial:

\[ \boxed{\text{numero di prove fisso}} \]

\[ \boxed{\text{numero di successi casuale}} \]

Negative binomial:

\[ \boxed{\text{numero di successi fissato}} \]

\[ \boxed{\text{numero di prove casuale}} \]


20. Summary of Chapter 17

Hypergeometric:

\[ \boxed{ P(X=k) = \frac{ \binom Kk\binom{N-K}{n-k} }{ \binom Nn } } \]

Negative binomial:

\[ \boxed{ P(X=k) = \binom{k-1}{r-1} p^r(1-p)^{k-r} } \]


Chapter 18 – How to recognize the right distribution

1. The real problem

Often the difficulty is not the formula.

The real question is:

\[ \boxed{ \text{quale distribuzione devo usare?} } \]


2. First distinction

The variable is:

\[ \boxed{\text{discreta}} \]

or:

\[ \boxed{\text{continua}} \]


3. Discrete variables

Examples:

  • number of customers;
  • number of defects;
  • number of successes;
  • number of attempts.

Main distributions:

  • Bernoulli;
  • Binomial;
  • Geometric;
  • Negative binomial;
  • Hypergeometric;
  • Poisson.

4. Continuous variables

Examples:

  • height;
  • time;
  • temperature;
  • weight.

Main distributions:

  • Uniform;
  • Exponential;
  • Normal.

5. Bernoulli

Just one test:

\[ \boxed{\text{successo/insuccesso}} \]


6. Binomial

Fixed number of tests.

Question:

how many successes?

\[ \boxed{ X\sim Bin(n,p) } \]


7. Geometric

Question:

how many trials until the first success?

\[ \boxed{ X\sim Geom(p) } \]


8. Negative binomial

Question:

how many tests until the \(r\)-th success?

\[ \boxed{ X\sim NB(r,p) } \]


9. Hypergeometric

Extraction without replacement from a finite population.

\[ \boxed{ X\sim Hyp(N,K,n) } \]


10. Poisson

Question:

how many events in an interval?

\[ \boxed{ X\sim Poisson(\lambda) } \]


11. Uniform

All values in an interval are equally plausible.

\[ \boxed{ X\sim U(a,b) } \]


12. Exponential

Question:

how long until the next event?

\[ \boxed{ X\sim Exp(\lambda) } \]


13. Normal

Continuous, symmetrical and bell-shaped phenomenon:

\[ \boxed{ X\sim N(\mu,\sigma^2) } \]


14. Discrete decision path

Just one test?

\[ \boxed{\text{Bernoulli}} \]

Fixed number of trials, do I count successes?

\[ \boxed{\text{Binomiale}} \]

First success?

\[ \boxed{\text{Geometrica}} \]

\(r\)-th success?

\[ \boxed{\text{Binomiale negativa}} \]

Without reinstatement?

\[ \boxed{\text{Ipergeometrica}} \]

Events in an interval?

\[ \boxed{\text{Poisson}} \]


15. Continuous decision-making process

Equiplausible values?

\[ \boxed{\text{Uniforme}} \]

Waiting time?

\[ \boxed{\text{Esponenziale}} \]

Symmetrical bell?

\[ \boxed{\text{Normale}} \]


16. Example: coin 10 times

Probability of exactly 7 heads.

Fixed number of tests:

\[ n=10 \]

So:

\[ \boxed{\text{Binomiale}} \]


17. Same coin up to the first head

\[ \boxed{\text{Geometrica}} \]


18. Same coin up to the third head

\[ \boxed{\text{Binomiale negativa}} \]


19. Urn without reinsertion

\[ \boxed{\text{Ipergeometrica}} \]


20. Six calls an hour

Number of calls:

\[ \boxed{\text{Poisson}} \]

Time until next time:

\[ \boxed{\text{Esponenziale}} \]


21. Bus in a range

If every moment between 8:00 and 8:20 is equally probable:

\[ \boxed{\text{Uniforme}} \]


22. Normally distributed height

\[ \boxed{\text{Normale}} \]


23. Exact model and approximation

Example:

\[ X\sim Bin(5000,0,001) \]

It's the exact model.

Because:

\[ np=5 \]

we can use:

\[ Poisson(5) \]

as an approximation.


24. Important approximations

Binomial versus Poisson:

\[ n\text{ grande},\ p\text{ piccolo} \]

Binomial versus normal:

\[ n\text{ grande} \]

Poisson towards normal:

\[ \lambda\text{ grande} \]


25. Fundamental rule

Always ask yourself:

\[ \boxed{ \text{che cosa è fissato e che cosa è casuale?} } \]


26. Frequent errors

Confuse:

  • Poisson and exponential;
  • binomial and geometric;
  • binomial and hypergeometric;
  • distribution of data and distribution of the mean.

27. Model and reality

A distribution is a model.

We don't say:

reality is a Poisson.

We say:

Poisson is a useful model to describe the phenomenon.


28. Summary of Chapter 18

\[ \boxed{ \text{fenomeno} \rightarrow \text{variabile} \rightarrow \text{ipotesi} \rightarrow \text{distribuzione} \rightarrow \text{calcolo} } \]

First understand the phenomenon, then choose the formula.


Chapter 19 – Monte Carlo simulation and experimental verification

1. Why simulate

Many problems can be studied by computer simulating the random experiment many times.

The idea is:

\[ \boxed{ \text{probabilità} \approx \text{frequenza relativa in molte simulazioni} } \]


2. Relative frequency

If a \(A\) event occurs:

\[ N_A \]

times on:

\[ N \]

simulations:

\[ \boxed{ \hat p= \frac{N_A}{N} } \]


3. Example with coin

10 throws can give:

\[ 6 \]

heads.

So:

\[ \hat p=0,6 \]

With 10000 launches we could obtain:

\[ 0,502 \]

The frequency tends towards:

\[ 0,5 \]


4. Connection with the Law of Large Numbers

Let's define:

\[ I_i= \begin{cases} 1 & \text{se }A\text{ avviene}\\ 0 & \text{altrimenti} \end{cases} \]

Then:

\[ \hat p= \frac1N\sum_{i=1}^NI_i \]

and:

\[ \hat p\to P(A) \]


5. Two dice

Theoretical probability of sum 7:

\[ \frac16 \]

Simulating 100,000 launches and obtaining 16642 times the sum 7:

\[ \hat p= 0,16642 \]

very close to:

\[ 0,16667 \]


6. What is Monte Carlo

The Monte Carlo method uses repeated random simulations to estimate mathematical quantities.

Procedure:

  1. generate random results;
  2. calculate the amount of interest;
  3. repeat many times;
  4. make an average or frequency.

7. Monte Carlo estimate

\[ \boxed{ \hat p_N= \frac1N \sum_{i=1}^{N}I_i } \]


8. Monte Carlo standard error

\[ \boxed{ SE(\hat p) = \sqrt{ \frac{p(1-p)}{N} } } \]

In practice we use:

\[ \hat p \]

instead of \(p\).


9. Consequence

The error decreases as:

\[ \frac1{\sqrt N} \]

To halve it you need to quadruple the number of simulations.


10. Monte Carlo interval

Approximately:

\[ \boxed{ \hat p \pm 1,96 \sqrt{ \frac{\hat p(1-\hat p)}{N} } } \]


11. Estimation of \(\pi\)

Let's generate random points in the square:

\[ [-1,1]\times[-1,1] \]

A point belongs to the circle if:

\[ X^2+Y^2\leq1 \]

Because:

\[ P(\text{cerchio})=\frac{\pi}{4} \]

we have:

\[ \boxed{ \pi \approx 4\frac{N_C}{N} } \]


12. Monte Carlo for integrals

If:

\[ U\sim U(a,b) \]

then:

\[ E[f(U)] = \frac1{b-a} \int_a^bf(x)\,dx \]

So:

\[ \boxed{ \int_a^bf(x)\,dx \approx (b-a) \frac1N \sum_{i=1}^{N}f(U_i) } \]


13. Simulate the Law of Large Numbers

We can:

  1. generate data;
  2. calculate the running average;
  3. represent it graphically;
  4. compare it with \(\mu\).

14. Simulate the TCL

Procedure:

  1. choose a non-normal distribution;
  2. extract a sample of size \(n\);
  3. calculate the average;
  4. repeat thousands of times;
  5. build the histogram of the averages.

As \(n\) increases, the normal shape emerges.


15. Standardization of averages

\[ \boxed{ Z= \frac{ \bar X-\mu }{ \sigma/\sqrt n } } \]

The simulated \(Z\)s should look like:

\[ N(0,1) \]


16. Simulation of a binomial

For:

\[ X\sim Bin(20,0,3) \]

Let's simulate 20 Bernoulli trials and count the successes.

We repeat many times.

The empirical histogram will tend towards the theoretical distribution.


17. Poisson simulation

For:

\[ X\sim Poisson(4) \]

we simulate many values.

We should observe:

\[ \bar X\approx4 \]

and:

\[ s^2\approx4 \]


18. Pseudorandom generators

Computers often use numbers:

\[ \boxed{\text{pseudocasuali}} \]

The seed allows you to reproduce the same sequence.


19. Inverse method

If:

\[ U\sim U(0,1) \]

and \(F\) is invertible:

\[ \boxed{ X=F^{-1}(U) } \]

It has \(F\) distribution.


20. Exponential example

For:

\[ F(x)=1-e^{-\lambda x} \]

we get:

\[ \boxed{ X= -\frac{\ln U}{\lambda} } \]


21. Rare events

If:

\[ P(A)=10^{-8} \]

even millions of simulations may never observe \(A\).

Direct Monte Carlo becomes inefficient.


22. Bootstrap

The bootstrap uses the observed data and generates new samples with reinsertion.

Procedure:

  1. original sample;
  2. resampling with reinsertion;
  3. calculation of statistics;
  4. repetition;
  5. study of the empirical distribution.

23. Bootstrapping utility

It allows you to estimate:

  • standard error;
  • bias;
  • confidence intervals;

even for complicated statistics.


24. Online laboratory

A simulator might ask:

  • distribution;
  • parameters;
  • number of simulations;
  • event.

And show:

  • frequency;
  • theoretical probability;
  • mistake;
  • histogram;
  • mean and variance.

25. TCL Laboratory

The user chooses:

  • initial distribution;
  • sample size \(n\);
  • number of simulations.

The program displays the histogram of the averages.


26. Confidence intervals laboratory

100 95% intervals can be simulated.

About:

\[ 95 \]

they should contain the true value.

This makes the Frequentist meaning of confidence visible.


27. Hypothesis testing laboratory

With true \(H_0\):

\[ \alpha=0,05 \]

repeating many tests we should incorrectly reject \(H_0\) in about 5% of cases.


28. Theory and simulation

The theory:

\[ \boxed{\text{spiega perché}} \]

The simulation:

\[ \boxed{\text{mostra cosa accade}} \]

They are complementary tools.


29. Summary of Chapter 19

Monte Carlo procedure:

  1. define the experiment;
  2. generate random data;
  3. calculate the result;
  4. repeat many times;
  5. summarize frequencies and averages.

Try the simulations: open the interactive laboratory.

Chapter 20 – From data to distribution

1. The inverse problem

So far we have often done:

\[ \boxed{ \text{distribuzione} \rightarrow \text{probabilità} } \]

In practice we can have:

\[ x_1,x_2,\ldots,x_n \]

and want to understand:

\[ \boxed{ \text{quale distribuzione potrebbe aver generato i dati?} } \]


2. First step: discrete or continuous

Integer count data:

\[ 0,1,2,\ldots \]

suggest a discrete variable.

Actual measurements:

\[ 2,13,\ 5,78,\ 9,41 \]

suggest a continuous variable.


3. Understand the phenomenon

It's not enough to look at the graph.

We also need to understand how the data is generated.

For example:

  • counts → Poisson;
  • waiting times → exponential;
  • successes → binomial;
  • symmetrical measurements → normal.

4. Descriptive statistics

Let's calculate:

  • numerousness;
  • minimum;
  • maximum;
  • average;
  • median;
  • variance;
  • standard deviation;
  • quartiles;
  • percentiles.

5. Average

\[ \boxed{ \bar x= \frac1n \sum_{i=1}^{n}x_i } \]


6. Sample variance

\[ \boxed{ s^2= \frac{ \sum_{i=1}^{n}(x_i-\bar x)^2 }{ n-1 } } \]


7. Quartiles and IQR

\[ \boxed{ IQR=Q_3-Q_1 } \]

Measure the dispersion of the central part.


8. Histogram

The histogram allows you to observe:

  • symmetry;
  • asymmetry;
  • peaks;
  • queues;
  • outliers.

9. Symmetrical distribution

If the histogram is bell-shaped and symmetrical, a normal may be plausible.


10. Asymmetry to the right

Many small values and a few large values can suggest patterns such as exponential.


11. Uniformity

If the histogram is flat enough:

\[ \boxed{\text{Uniforme}} \]

could be a candidate.


12. Multimodality

Multiple peaks may indicate a mixture of different populations.

In this case a single simple deployment may not be sufficient.


13. Box plot

Show:

  • quartiles;
  • median;
  • dispersion;
  • extreme values.

14. Outliers

A common rule considers values lower than: suspicious.

\[ Q_1-1,5IQR \]

or greater than:

\[ Q_3+1,5IQR \]

But they should not be automatically eliminated.


15. Poisson: mean and variance

For a Poisson:

\[ E(X)=\operatorname{Var}(X)=\lambda \]

So if:

\[ \bar x\approx s^2 \]

a Poisson may be plausible.


16. Overdispersion

If:

\[ s^2\gg\bar x \]

there may be overdispersion.

A negative binomial may be more suitable.


17. Exponential

For an exponential:

\[ E(X)=\sigma \]

Thus positive, right-skewed data with similar mean and standard deviation may suggest an exponential model.


18. Normal

For a normal one we expect:

  • symmetry;
  • mean and median close;
  • a single peak;
  • tails on both sides.

19. Rule 68–95–99.7

We can check how many values fall within:

\[ \bar x\pm s \]

\[ \bar x\pm2s \]

\[ \bar x\pm3s \]

and compare with:

\[ 68\%,95\%,99,7\% \]


20. Q-Q plot

The Q-Q plot compares:

  • observed quantiles;
  • theoretical quantiles.

If the points are nearly aligned, the model may be plausible.


21. Empirical function

\[ \boxed{ F_n(x) = \frac{ \#\{x_i\leq x\} }{ n } } \]

It can be compared with a theoretical function:

\[ F(x) \]


22. Adaptation test

We can use goodness-of-fit tests.

Example:

\[ \boxed{\chi^2} \]

with statistics:

\[ \boxed{ \chi^2 = \sum_i \frac{ (O_i-E_i)^2 }{ E_i } } \]


23. Hypothesis

\[ H_0: \text{il modello è compatibile con i dati} \]

against:

\[ H_1: \text{il modello non è adeguato} \]


24. Kolmogorov-Smirnov

Compare:

\[ F_n(x) \]

and:

\[ F(x) \]

through:

\[ \boxed{ D= \sup_x|F_n(x)-F(x)| } \]


25. Normality test

Among the most used:

  • Shapiro-Wilk;
  • Anderson-Darling;
  • Kolmogorov-Smirnov in appropriate versions.

26. One test is not enough

With small samples it may have little power.

With huge samples it can detect tiny differences.

Better to combine:

\[ \boxed{ \text{grafici} + \text{indicatori} + \text{test} + \text{conoscenza del fenomeno} } \]


27. Parameter estimation

Poisson:

\[ \boxed{ \hat\lambda=\bar x } \]

Exponential:

\[ \boxed{ \hat\lambda=\frac1{\bar x} } \]

Normal:

\[ \boxed{ \hat\mu=\bar x } \]

\[ \boxed{ \hat\sigma\approx s } \]

Binomial:

\[ \boxed{ \hat p=\frac{\bar x}{n} } \]

if \(n\) is known.


28. Method of moments

The idea is to equalize theoretical and sample moments.

If:

\[ E(X)=g(\theta) \]

let's say:

\[ \bar x=g(\hat\theta) \]

and we get:

\[ \hat\theta \]


29. Maximum likelihood

Maximum likelihood chooses parameters that make the observed data as plausible as possible.

\[ \boxed{ \hat\theta = \arg\max_\theta L(\theta) } \]


30. Bernoulli example

With:

\[ k \]

hits on:

\[ n \]

tests:

\[ \boxed{ \hat p=\frac kn } \]


31. Compare multiple models

We can compare:

  • histogram;
  • theoretical densities;
  • Q-Q plot;
  • adaptation test;
  • statistical criteria.

32. Thrift

If two models describe the data the same way, we often prefer the simpler one.


33. Example counts

Suppose:

\[ \bar x=3,1 \]

\[ s^2=3,3 \]

A Poisson may be plausible.

We estimate:

\[ \hat\lambda=3,1 \]


34. Example of waiting times

Suppose:

\[ \bar x=5,2 \]

\[ s=5,0 \]

and strong asymmetry to the right.

An exponential may be plausible.

\[ \hat\lambda = \frac1{5,2} \approx0,192 \]


35. Normal example

Suppose:

\[ \bar x=170,3 \]

\[ s=7,4 \]

with symmetric histogram.

Candidate model:

\[ \boxed{ N(170,3,7,4^2) } \]


36. Auto simulator

A program could:

  1. read the data;
  2. classify discrete/continuous;
  3. calculate statistics;
  4. build graphs;
  5. propose models;
  6. estimate parameters;
  7. compare models;
  8. return a compatibility opinion.

37. Correct output

Better to say:

the normal model appears compatible with the data

rather than:

the data certainly follows a normal.


38. If no model works

You can answer:

\[ \boxed{ \text{nessuna delle distribuzioni considerate descrive adeguatamente i dati} } \]

There is no need to force the model.


39. Simulate from the estimated model

After estimating a model we can generate new data and compare it with real data.

This allows for a very intuitive check of suitability.


40. Final outline

\[ \boxed{ \text{Dati} \rightarrow \text{Grafici} \rightarrow \text{Indicatori} \rightarrow \text{Modello} \rightarrow \text{Stima} \rightarrow \text{Controllo} } \]


41. Final idea

The probability starts from the model:

\[ \boxed{ \text{modello} \rightarrow \text{dati} } \]

Statistics often takes the opposite path:

\[ \boxed{ \text{dati} \rightarrow \text{modello} } \]

The complete cycle becomes:

\[ \boxed{ \text{osservare} \rightarrow \text{ipotizzare} \rightarrow \text{verificare} \rightarrow \text{simulare} } \]

This is one of the most comprehensive ways to understand probability and statistics.