N-gram Language Modelling
N-gram Language Modelling #
A language model assigns probabilities to sequences of words. It can compare complete sentences or predict which word is likely to come next.
Key ideas include:
- word prediction and sequence probability
- the chain rule and Markov assumption
- unigram, bigram and trigram models
- Maximum Likelihood Estimation
- unseen sequences and smoothing
- interpolation and backoff
- intrinsic and extrinsic evaluation
- perplexity
Learning Objectives #
- Explain what a language model represents.
- Calculate simple unigram and bigram probabilities.
- Explain why unseen N-grams create zero probabilities.
- Distinguish smoothing, interpolation and backoff.
- Interpret perplexity correctly.
Big Picture #
flowchart TD
A["Training Corpus"] --> B["Count N-grams"]
B --> C["Estimate Probabilities"]
C --> D["Handle Unseen Events"]
D --> E["Score Word Sequences"]
E --> F["Evaluate Model"]
style A fill:#E1F5FE
style B fill:#C8E6C9
style C fill:#FFF9C4
style D fill:#EDE7F6
style E fill:#E1F5FE
style F fill:#C8E6C9
1. What Is a Language Model? ☆ #
A language model estimates how probable a sequence of words is.