Machine Learning is built on mathematical principles that allow models to:
represent data
learn patterns
optimise performance
flowchart LR
DATA[Data]
MATH[Math Models]
OPT[Optimisation]
MODEL[Trained Model]
DATA --> MATH
MATH --> OPT
OPT --> MODEL
ML requires core mathematical tools to understand how ML algorithms work internally. Algebra deals with relationships between variables and quantities, while Calculus focuses on change and optimization.
A/B testing, evaluating model improvements, significance vs noise, parameter estimation (MLE)
5. Prediction & Forecasting
Correlation, regression, time series (AR/MA/ARIMA/SARIMA etc.)
Linear regression, forecasting, sequential data modelling, baseline predictive modelling
6. GMM & EM
Mixtures of Gaussians; iterative estimation with EM
Unsupervised learning (soft clustering), density estimation, latent-variable models
flowchart TD
A["Statistical Methods<br/>AIML ZC418"] --> B["1. Basic Probability and Statistics"]
A --> C["2. Conditional Probability and Bayes"]
A --> D["3. Probability Distributions"]
A --> E["4. Hypothesis Testing"]
A --> F["5. Prediction and Forecasting"]
A --> G["6. Gaussian Mixture Model and EM"]
B --> B1["Central Tendency<br/>Mean - Median - Mode"]
B --> B2["Variability<br/>Range - Variance - SD - Quartiles"]
B --> B3["Basic Probability Concepts"]
B3 --> B31["Axioms of Probability"]
B3 --> B32["Definition of Probability"]
B3 --> B33["Mutually Exclusive vs Independent"]
C --> C1["Conditional Probability"]
C --> C2["Independence (conditional)"]
C --> C3["Bayes Theorem"]
C --> C4["Naive Bayes (intro)"]
D --> D1["Random Variables<br/>Discrete and Continuous"]
D --> D2["Expectation - Variance - Covariance"]
D --> D3["Transformations of RVs"]
D --> D4["Key Distributions"]
D4 --> D41["Bernoulli"]
D4 --> D42["Binomial"]
D4 --> D43["Poisson"]
D4 --> D44["Normal (Gaussian)"]
D4 --> D45["t - Chi-square - F (intro)"]
E --> E1["Sampling<br/>Random and Stratified"]
E --> E2["Sampling Distributions<br/>CLT"]
E --> E3["Estimation<br/>Confidence Intervals"]
E --> E4["Hypothesis Tests<br/>Means and Proportions"]
E --> E5["ANOVA<br/>Single and Dual factor"]
E --> E6["Maximum Likelihood"]
F --> F1["Correlation"]
F --> F2["Regression"]
F --> F3["Time Series Basics<br/>Components"]
F --> F4["Moving Averages<br/>Simple and Weighted"]
F --> F5["Time Series Models"]
F5 --> F51["AR"]
F5 --> F52["ARMA / ARIMA"]
F5 --> F53["SARIMA / SARIMAX"]
F5 --> F54["VAR / VARMAX"]
F --> F6["Exponential Smoothing"]
G --> G1["GMM<br/>Mixture of Gaussians"]
G --> G2["EM Algorithm<br/>E-step - M-step"]
B -.-> C
C -.-> D
D -.-> E
E -.-> F
F -.-> G
Probability often changes when we learn new information.
Conditional probability and Bayes’ theorem give a structured way to update beliefs using evidence.
Conditional probability updates probabilities after observing an event.
Bayes’ theorem lets you estimate a hidden cause from observed evidence.
Naïve Bayes turns Bayes’ theorem into a practical classifier by assuming conditional independence of features given the class.
flowchart TD
A[Conditional<br/>probability] -->|foundation| B[Bayes<br/>theorem]
D[Independent<br/>events] -->|implies| C[Independence]
C -->|simplifies| A
E[Prior] -->|with likelihood| B
F[Likelihood] -->|updates| H[Posterior]
G[Evidence] -->|normalises| B
B -->|yields| H
I[Naïve<br/>Bayes] -->|uses| B
J[Naïve<br/>assumption] -->|assumes| C
K[Features] -->|given class| J
L[Class] -->|conditions| J
I -->|predicts| M[Classification]
M -->|selects| L
style A fill:#90CAF9,stroke:#1E88E5,color:#000
style B fill:#90CAF9,stroke:#1E88E5,color:#000
style C fill:#90CAF9,stroke:#1E88E5,color:#000
style D fill:#CE93D8,stroke:#8E24AA,color:#000
style E fill:#CE93D8,stroke:#8E24AA,color:#000
style F fill:#CE93D8,stroke:#8E24AA,color:#000
style G fill:#CE93D8,stroke:#8E24AA,color:#000
style J fill:#CE93D8,stroke:#8E24AA,color:#000
style K fill:#CE93D8,stroke:#8E24AA,color:#000
style L fill:#CE93D8,stroke:#8E24AA,color:#000
style H fill:#C8E6C9,stroke:#2E7D32,color:#000
style I fill:#C8E6C9,stroke:#2E7D32,color:#000
style M fill:#C8E6C9,stroke:#2E7D32,color:#000
Probability distributions are the bridge between:
real-world randomness and mathematical modelling.
A random experiment produces outcomes.
A random variable turns those outcomes into numbers.
A probability distribution tells you how likely each number (or range of numbers) is.
Key takeaway:
A distribution is a complete “story” about uncertainty:
what values are possible, how likely they are, and how we summarise them (mean, variance).
Many ML models are probabilistic:
they assume data (or errors) follow a distribution.
Loss functions often come from distribution assumptions:
squared loss aligns with Gaussian noise.
Naïve Bayes (from the previous module) becomes practical once you can model:
\( P(X\mid Y) \)
using suitable distributions.
In practice:
choosing a distribution is a modelling decision.
It affects:
prediction, uncertainty estimates, and what “rare” or “typical” means in your data.
A random variable is a way to attach numbers to outcomes of a random experiment.
It lets us move from:
“what happened?”
to:
“what number should we analyse?”
Key takeaway:
A random variable is a function from the sample space to real numbers.
Once you define the random variable clearly, the rest (pmf/pdf/cdf, mean, variance) becomes systematic.
flowchart TD
PD["Probability<br/>distributions"] --> RV["Random<br/>variables"]
RV --> T["Types"]
T --> RV1["Discrete<br/>RVs"]
T --> RV2["Continuous<br/>RVs"]
RV --> F["PMF / PDF / CDF"]
RV --> S["Mean / Variance<br/>Covariance"]
RV --> J["Joint & Marginal<br/>distributions"]
RV --> X["Transformations"]
style PD fill:#90CAF9,stroke:#1E88E5,color:#000
style RV fill:#90CAF9,stroke:#1E88E5,color:#000
style T fill:#CE93D8,stroke:#8E24AA,color:#000
style F fill:#CE93D8,stroke:#8E24AA,color:#000
style S fill:#CE93D8,stroke:#8E24AA,color:#000
style J fill:#CE93D8,stroke:#8E24AA,color:#000
style X fill:#CE93D8,stroke:#8E24AA,color:#000
style RV1 fill:#CE93D8,stroke:#8E24AA,color:#000
style RV2 fill:#CE93D8,stroke:#8E24AA,color:#000