Gaussian Mixture Model & Expectation Maximization
#
A Gaussian Mixture Model represents data as a weighted combination of multiple Gaussian distributions.
It is commonly used for soft clustering and density estimation.
Key takeaway:
K-means gives hard cluster membership.
GMM gives probabilities of belonging to each cluster.
- Gaussian Mixture Model
- soft clustering
- mixing coefficients
- latent variables
- likelihood and log-likelihood
- Expectation-Maximization algorithm
- E-step and M-step
- responsibilities
- convergence
Motivation ☆
#
Many real datasets are not described well by one Gaussian distribution.
Instance-based Learning
#
Instance-based learning is a family of methods that do not build one explicit global model during training. Instead, they store training examples and delay most of the work until a new query arrives.
When a new point must be classified or predicted, the algorithm compares it with previously seen examples, finds the most relevant neighbours, and uses them to produce the answer.
Instance-based Learning covers three linked ideas:
May 8, 2026Support Vector Machine (SVM)
#
Support Vector Machine (SVM) is a supervised machine learning algorithm used for:
- Classification (most common)
- Regression (SVR – Support Vector Regression)
It connects many earlier ideas:
- classification and decision boundaries
- linear classifiers
- margins
- optimisation
- constrained optimisation
- kernels for non-linear data
SVM is a discriminative classifier.
That means it does not try to model how each class is generated.
Instead, it tries to find the best separating boundary between classes.
Attention Mechanism
#
Attention is a deep learning mechanism that allows a model to focus on the most relevant parts of an input sequence when producing an output.
Instead of compressing the whole input into one fixed vector, attention computes a weighted combination of useful information.
Key takeaway:
Attention answers a simple question:
For the current prediction, which input tokens should the model focus on most?
- Queries, Keys, and Values
- Attention Pooling by Similarity
- Attention Pooling via Nadaraya–Watson Regression
- Attention Scoring Functions
- Dot Product Attention
- Convenience Functions
- Scaled Dot Product Attention
- Additive Attention
- Bahdanau Attention Mechanism
- Multi-Head Attention
- Self-Attention
- Positional Encoding
Why Attention Is Needed ☆
#
Traditional encoder-decoder RNN models compress the full input sequence into one context vector.
Bayesian Learning
#
Bayesian Learning is a probabilistic approach to machine learning.
Instead of only asking, “Which output should the model predict?”, Bayesian Learning asks:
Given the data we have observed, how likely is each hypothesis, class, or parameter value?
This makes Bayesian Learning useful when uncertainty matters.
It is especially important in classification, probabilistic modelling, generative models, and situations where we want to combine prior knowledge with observed data.
Ensemble Learning
#
Ensemble Learning is a machine learning approach where we combine multiple models to produce a stronger final prediction.
Instead of depending on one model, an ensemble uses a group of models and combines their outputs.
The main idea is simple:
Many weak or moderately good models can work together to produce a better and more stable model.
Key takeaway:
Ensemble Learning improves prediction by combining several models.
ML System Optimisation
#
ML System Optimisation studies how to make machine learning workloads faster, more scalable, more memory-efficient, and suitable for different hardware platforms.
The subject connects machine learning algorithms with the systems that train and deploy them: multi-core CPUs, GPUs, distributed clusters, cloud platforms, edge devices, and embedded systems.
ML system optimisation = model quality + computational efficiency + hardware awareness + scalability
The learning path begins with performance measurement and parallel computing, progresses through distributed machine learning and scale-out platforms, and concludes with model compression and resource-constrained deployment.
A transformer is a neural network architecture that uses attention as its main mechanism for processing sequences.
Unlike RNNs, transformers do not process tokens one by one.
They process many tokens in parallel and use self-attention to learn relationships between tokens.
is an architecture of neural networks
based on the multi-head attention mechanism
text is converted to numerical representations called tokens, and each token is converted into a vector via lookup from a word embedding table
Optimisation of Deep models
#
Optimizers are algorithms that update neural network parameters to reduce the loss function.
Deep networks usually have millions or billions of parameters, so there is usually no closed-form solution.
Instead, training uses iterative optimisation.
Key takeaway:
An optimiser decides how the model moves through the loss landscape towards lower loss.
- Goal of Optimization
- Optimization Challenges in Deep Learning
- Gradient Descent
- Stochastic Gradient Descent
- Minibatch Stochastic Gradient Descent
- Momentum
- Adagrad and Algorithm
- RMSProp and Algorithm
- Adadelta and Algorithm
- Adam and Algorithm
- Code Implementation and comparison of algorithms (webinar)
flowchart TD
A["Optimisers in DNN"] --> B["Gradient Descent Variants"]
A --> C["Momentum-based Optimiser"]
A --> D["Adaptive Methods"]
A --> E["Learning Rate Schedules"]
D --> D1["Parameter-specific learning rates"]
E --> E1["Learning rate changes during training"]
style A fill:#E1F5FE,stroke:#4A90E2,stroke-width:2px
style B fill:#EDE7F6,stroke:#7E57C2
style C fill:#C8E6C9,stroke:#43A047
style D fill:#FFF9C4,stroke:#FBC02D
style E fill:#F8BBD0,stroke:#D81B60
Goal of Optimisation ☆
#
The goal is to find parameters
\( \theta \)
that minimise the loss.
Unsupervised Learning
#
Unsupervised Learning is used when we have input data but no target labels.
The model is not told the correct answer. Instead, it tries to discover hidden structure in the data.
- K-means Clustering and variants
- Review of EM algorithm
- GMM based Soft Clustering
- Applications
Supervised vs Unsupervised Learning
#
| Aspect | Supervised Learning | Unsupervised Learning |
|---|
| Data contains target label? | Yes | No |
| Learns from | Input-output pairs | Input features only |
| Main goal | Predict output | Discover structure |
| Example task | Classification, regression | Clustering |
| Example algorithm | Logistic regression, decision tree | K-means, GMM |
- Works on unlabelled raw data.
- The algorithm discovers hidden patterns without prior knowledge of outcomes.
- Requires no human intervention during training.
- Does not make direct predictions — it groups or organises data instead.
- Carries a higher risk because there’s no ground truth to verify results.
- Common techniques include Clustering, Association, and Dimensionality Reduction.
The most common example is clustering, where similar records are grouped together.