flowchart TD
A[Words and Documents] --> B[Observe Their Context]
B --> C[Represent Them as Vectors]
C --> D[Compare Vector Directions]
D --> E[Estimate Semantic Similarity]
E --> F[Search, Classify, Retrieve or Generate]
style A fill:#E1F5FE
style B fill:#C8E6C9
style C fill:#FFF9C4
style D fill:#EDE7F6
style E fill:#E1F5FE
style F fill:#C8E6C9
Probability distributions are the bridge between:
real-world randomness and mathematical modelling.
A random experiment produces outcomes.
A random variable turns those outcomes into numbers.
A probability distribution tells you how likely each number (or range of numbers) is.
Key takeaway:
A distribution is a complete “story” about uncertainty:
what values are possible, how likely they are, and how we summarise them (mean, variance).
Many ML models are probabilistic:
they assume data (or errors) follow a distribution.
Loss functions often come from distribution assumptions:
squared loss aligns with Gaussian noise.
Naïve Bayes (from the previous module) becomes practical once you can model:
\( P(X\mid Y) \)
using suitable distributions.
In practice:
choosing a distribution is a modelling decision.
It affects:
prediction, uncertainty estimates, and what “rare” or “typical” means in your data.
A linear neural network for regression is a model that predicts a continuous target by taking a weighted sum of input features and applying the identity activation (so the output can be any real number).
Single neuron for regression (predicting how much / how many)
Data + linear model (single neuron, no hidden layers) + squared loss
Training using batch gradient descent algorithm
Prediction (inference)
Eg: Auto MPG (UCI) style prediction with a single neuron (from-scratch code)
flowchart LR
D["Data<br/>X, y"] --> M["Linear model<br/>w, b<br/>Single neuron"]
M --> A["Activation<br/>Identity"]
A --> L["Loss<br/>MSE (Squared error)"]
L --> O["Optimiser<br/>Batch Gradient DescentBatch GD / Mini-batch GD"]
O --> P["Parameters<br/>w, b"]
P --> I["Inference<br/>Predict ŷ (number) for new x"]
%% Pastel colour scheme
style D fill:#E3F2FD,stroke:#1E88E5,stroke-width:1px
style M fill:#E8F5E9,stroke:#43A047,stroke-width:1px
style A fill:#FFF3E0,stroke:#FB8C00,stroke-width:1px
style L fill:#FCE4EC,stroke:#D81B60,stroke-width:1px
style O fill:#F3E5F5,stroke:#8E24AA,stroke-width:1px
style P fill:#E0F7FA,stroke:#00838F,stroke-width:1px
style I fill:#F1F8E9,stroke:#558B2F,stroke-width:1px
The AI pipeline is a continuous process where data is collected, prepared, used to train models, evaluated for performance, and continuously improved after deployment.
timeline
title AI Pipeline
Collect Data : Data Ingestion
: Data Understanding
Prepare Data : Cleaning
: Feature Engineering
: Sampling
Train Model : Model Training
: Validation & Metrics
Deploy Model : Deployment
: Monitoring & Retraining
A Markov Decision Process (MDP) is a mathematical framework for modelling sequential decisions. It describes the situations an agent may encounter, the actions it may take, how the environment may change, and the rewards produced by those changes.
Bandit problems ask which action is best in a single recurring situation. An MDP adds changing states: an action affects not only the immediate reward but also the situation faced next.
flowchart TD
A["Training Corpus"] --> B["Count N-grams"]
B --> C["Estimate Probabilities"]
C --> D["Handle Unseen Events"]
D --> E["Score Word Sequences"]
E --> F["Evaluate Model"]
style A fill:#E1F5FE
style B fill:#C8E6C9
style C fill:#FFF9C4
style D fill:#EDE7F6
style E fill:#E1F5FE
style F fill:#C8E6C9
Parallel algorithms need hardware that can execute independent work efficiently. Modern systems therefore combine multiple CPU cores, memory hierarchies, threads, instruction pipelines, GPUs, clusters, and specialised matrix processors.
This page covers:
multi-core CPU organisation
cache and memory hierarchy
processes, threads, scheduling, and synchronisation