<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Statistics on Arshad Siddiqui</title><link>https://arshadhs.github.io/docs/ai/020-statistics/</link><description>Recent content in Statistics on Arshad Siddiqui</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 12 Mar 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://arshadhs.github.io/docs/ai/020-statistics/index.xml" rel="self" type="application/rss+xml"/><item><title>Formula Sheet</title><link>https://arshadhs.github.io/docs/ai/020-statistics/00_formulas/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/00_formulas/</guid><description>&lt;h1 id="formula-sheet">
 Formula Sheet
 
 &lt;a class="anchor" href="#formula-sheet">#&lt;/a>
 
&lt;/h1>
&lt;p>This page is a quick reference of &lt;strong>definitions + formulas&lt;/strong>, grouped by the modules.&lt;/p>
&lt;hr>
&lt;h2 id="notation">
 Notation
 
 &lt;a class="anchor" href="#notation">#&lt;/a>
 
&lt;/h2>
&lt;ul>
&lt;li>Sample size: 
&lt;link rel="stylesheet" href="https://arshadhs.github.io/katex/katex.min.css" />
&lt;script defer src="https://arshadhs.github.io/katex/katex.min.js">&lt;/script>

 &lt;script defer src="https://arshadhs.github.io/katex/auto-render.min.js" onload="renderMathInElement(document.body, {
 &amp;#34;delimiters&amp;#34;: [
 {&amp;#34;left&amp;#34;: &amp;#34;$$&amp;#34;, &amp;#34;right&amp;#34;: &amp;#34;$$&amp;#34;, &amp;#34;display&amp;#34;: true},
 {&amp;#34;left&amp;#34;: &amp;#34;$&amp;#34;, &amp;#34;right&amp;#34;: &amp;#34;$&amp;#34;, &amp;#34;display&amp;#34;: false},
 {&amp;#34;left&amp;#34;: &amp;#34;\\(&amp;#34;, &amp;#34;right&amp;#34;: &amp;#34;\\)&amp;#34;, &amp;#34;display&amp;#34;: false},
 {&amp;#34;left&amp;#34;: &amp;#34;\\[&amp;#34;, &amp;#34;right&amp;#34;: &amp;#34;\\]&amp;#34;, &amp;#34;display&amp;#34;: true}
 ]
});">&lt;/script>

&lt;span>
 \( n \)
 &lt;/span>

 (sample), 
&lt;span>
 \( N \)
 &lt;/span>

 (population)&lt;/li>
&lt;li>Sample mean: 
&lt;span>
 \( \bar{x} \)
 &lt;/span>

, population mean: 
&lt;span>
 \( \mu \)
 &lt;/span>

&lt;/li>
&lt;li>Sample variance: 
&lt;span>
 \( s^2 \)
 &lt;/span>

, population variance: 
&lt;span>
 \( \sigma^2 \)
 &lt;/span>

&lt;/li>
&lt;li>Sample SD: 
&lt;span>
 \( s \)
 &lt;/span>

, population SD: 
&lt;span>
 \( \sigma \)
 &lt;/span>

&lt;/li>
&lt;li>Complement: 
&lt;span>
 \( A^c \)
 &lt;/span>

&lt;/li>
&lt;li>Intersection (“and”): 
&lt;span>
 \( A\cap B \)
 &lt;/span>

, union (“or”): 
&lt;span>
 \( A\cup B \)
 &lt;/span>

&lt;/li>
&lt;li>Conditional probability: 
&lt;span>
 \( P(A\mid B) \)
 &lt;/span>

&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h1 id="1-basic-probability--statistics">
 1. Basic Probability &amp;amp; Statistics
 
 &lt;a class="anchor" href="#1-basic-probability--statistics">#&lt;/a>
 
&lt;/h1>
&lt;h2 id="11-measures-of-central-tendency">
 1.1 Measures of Central Tendency
 
 &lt;a class="anchor" href="#11-measures-of-central-tendency">#&lt;/a>
 
&lt;/h2>
&lt;h3 id="arithmetic-mean">
 Arithmetic mean
 
 &lt;a class="anchor" href="#arithmetic-mean">#&lt;/a>
 
&lt;/h3>
&lt;p>Sample mean (ungrouped):&lt;/p></description></item><item><title>Stats Formula Sheet</title><link>https://arshadhs.github.io/docs/ai/020-statistics/ism-formula-sheet/</link><pubDate>Wed, 25 Feb 2026 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/ism-formula-sheet/</guid><description>&lt;h1 id="stats-formula-sheet">
 Stats Formula Sheet
 
 &lt;a class="anchor" href="#stats-formula-sheet">#&lt;/a>
 
&lt;/h1>
&lt;p>Keep this page as a quick reference of &lt;strong>definitions + formulas&lt;/strong>.&lt;/p>
&lt;hr>
&lt;h2 id="notation">
 Notation
 
 &lt;a class="anchor" href="#notation">#&lt;/a>
 
&lt;/h2>
&lt;ul>
&lt;li>Sample size: 
&lt;span>
 \( n \)
 &lt;/span>

 (sample), 
&lt;span>
 \( N \)
 &lt;/span>

 (population)&lt;/li>
&lt;li>Mean: 
&lt;span>
 \( \bar{x} \)
 &lt;/span>

 (sample), 
&lt;span>
 \( \mu \)
 &lt;/span>

 (population)&lt;/li>
&lt;li>Variance: 
&lt;span>
 \( s^2 \)
 &lt;/span>

 (sample), 
&lt;span>
 \( \sigma^2 \)
 &lt;/span>

 (population)&lt;/li>
&lt;li>Standard deviation: 
&lt;span>
 \( s \)
 &lt;/span>

 (sample), 
&lt;span>
 \( \sigma \)
 &lt;/span>

 (population)&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="module-1-basic-statistics">
 Module 1: Basic Statistics
 
 &lt;a class="anchor" href="#module-1-basic-statistics">#&lt;/a>
 
&lt;/h2>
&lt;h3 id="measures-of-central-tendency">
 Measures of Central Tendency
 
 &lt;a class="anchor" href="#measures-of-central-tendency">#&lt;/a>
 
&lt;/h3>
&lt;p>&lt;strong>Sample mean (ungrouped):&lt;/strong>&lt;/p></description></item><item><title>Basic Statistics</title><link>https://arshadhs.github.io/docs/ai/020-statistics/01_basic_statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/01_basic_statistics/</guid><description>&lt;h1 id="basic-statistics">
 Basic Statistics
 
 &lt;a class="anchor" href="#basic-statistics">#&lt;/a>
 
&lt;/h1>
&lt;p>&lt;strong>Statistics&lt;/strong>: describes data (what you &lt;em>see&lt;/em>).&lt;br>
&lt;strong>Probability&lt;/strong>: models uncertainty (what you &lt;em>don’t know&lt;/em> yet).&lt;/p>
&lt;ul>
&lt;li>Summarise a dataset using central tendency and variability&lt;/li>
&lt;li>Explain core probability ideas using simple examples&lt;/li>
&lt;li>Apply the axioms of probability&lt;/li>
&lt;li>Distinguish mutually exclusive vs independent events&lt;/li>
&lt;/ul>
&lt;hr>


&lt;script src="https://arshadhs.github.io/mermaid.min.js">&lt;/script>

 &lt;script>mermaid.initialize({
 "flowchart": {
 "useMaxWidth":true
 },
 "theme": "default"
}
)&lt;/script>




&lt;pre class="mermaid">
flowchart TD
 A[Dataset] --&amp;gt; B[Central Tendency]
 A --&amp;gt; C[Variability]
 B --&amp;gt; B1[Mean]
 B --&amp;gt; B2[Median]
 B --&amp;gt; B3[Mode]
 C --&amp;gt; C1[Range]
 C --&amp;gt; C2[Variance]
 C --&amp;gt; C3[Standard Deviation]
 C --&amp;gt; C4[IQR]
&lt;/pre>

&lt;hr>
&lt;h2 id="measures-of-central-tendency">
 Measures of Central Tendency
 
 &lt;a class="anchor" href="#measures-of-central-tendency">#&lt;/a>
 
&lt;/h2>
&lt;p>Central tendency tells you where the “middle” of the data is.
Describes a set of scores with a &lt;strong>single number&lt;/strong> that describes the &lt;strong>PERFORMANCE&lt;/strong> of the group.&lt;/p></description></item><item><title>Basic Probability</title><link>https://arshadhs.github.io/docs/ai/020-statistics/01_basic_probability/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/01_basic_probability/</guid><description>&lt;h1 id="basic-probability">
 Basic Probability
 
 &lt;a class="anchor" href="#basic-probability">#&lt;/a>
 
&lt;/h1>
&lt;p>Probability models uncertainty:
what you &lt;em>don’t know&lt;/em> yet, but want to reason about.&lt;/p>
&lt;blockquote class="book-hint info">
&lt;p>Key takeaway:
Probability is a number between &lt;strong>0 and 1&lt;/strong> that measures how likely an event is.
The whole topic is about defining &lt;strong>events&lt;/strong> clearly and applying a few core rules consistently.&lt;/p>
&lt;/blockquote>
&lt;p>Probability quantifies uncertainty: a number between 0 and 1.&lt;/p>
&lt;ul>
&lt;li>0 means: impossible&lt;/li>
&lt;li>1 means: certain&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="terminology">
 Terminology
 
 &lt;a class="anchor" href="#terminology">#&lt;/a>
 
&lt;/h2>
&lt;h3 id="random-experiment">
 Random experiment
 
 &lt;a class="anchor" href="#random-experiment">#&lt;/a>
 
&lt;/h3>
&lt;p>A random experiment is an action whose outcome is not known in advance.&lt;/p></description></item><item><title>Hypothesis Testing</title><link>https://arshadhs.github.io/docs/ai/020-statistics/040-hypothesis-testing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/040-hypothesis-testing/</guid><description>&lt;h1 id="hypothesis-testing">
 Hypothesis Testing
 
 &lt;a class="anchor" href="#hypothesis-testing">#&lt;/a>
 
&lt;/h1>
&lt;p>Hypothesis testing is a statistical decision-making method used to decide whether sample evidence is strong enough to reject an initial assumption about a population.&lt;/p>
&lt;p>It connects probability, sampling distributions, confidence intervals, significance levels, and decision rules.&lt;/p>
&lt;blockquote class="book-hint info">
&lt;p>&lt;strong>Key takeaway:&lt;/strong>&lt;br>
Hypothesis testing is not about proving something with certainty.&lt;/p>
&lt;p>It is about asking:&lt;/p>

&lt;blockquote class='book-hint '>
 &lt;p>If the null hypothesis were true, how surprising would this sample result be?&lt;/p></description></item><item><title>Prediction &amp; Forecasting</title><link>https://arshadhs.github.io/docs/ai/020-statistics/050-prediction-forecasting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/050-prediction-forecasting/</guid><description>&lt;h1 id="prediction--forecasting">
 Prediction &amp;amp; Forecasting
 
 &lt;a class="anchor" href="#prediction--forecasting">#&lt;/a>
 
&lt;/h1>
&lt;p>Prediction and forecasting use statistical models to estimate unknown or future values.&lt;/p>
&lt;p>In this module, the focus is on correlation, regression, and time series forecasting.&lt;/p>
&lt;blockquote class="book-hint info">
&lt;p>&lt;strong>Key takeaway:&lt;/strong>&lt;br>
Prediction estimates a value using a model.&lt;/p>
&lt;p>Forecasting is prediction where the order of time matters.&lt;/p>
&lt;/blockquote>
&lt;ul>
&lt;li>Correlation&lt;/li>
&lt;li>Regression&lt;/li>
&lt;li>Time series analysis&lt;/li>
&lt;li>Components of time series data&lt;/li>
&lt;li>Moving average and weighted moving average&lt;/li>
&lt;li>AR model&lt;/li>
&lt;li>ARMA model&lt;/li>
&lt;li>ARIMA model&lt;/li>
&lt;li>SARIMA and SARIMAX&lt;/li>
&lt;li>VAR and VARMAX&lt;/li>
&lt;li>Simple exponential smoothing&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="prediction-vs-forecasting-">
 Prediction vs Forecasting ☆
 
 &lt;a class="anchor" href="#prediction-vs-forecasting-">#&lt;/a>
 
&lt;/h2>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Concept&lt;/th>
 &lt;th>Meaning&lt;/th>
 &lt;th>Example&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>Prediction&lt;/td>
 &lt;td>Estimate an unknown output&lt;/td>
 &lt;td>Predict house price from area and rooms&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Forecasting&lt;/td>
 &lt;td>Predict future values using time order&lt;/td>
 &lt;td>Forecast sales for next month&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;span style="color: red;">
 All forecasting is prediction, but not all prediction is forecasting.
&lt;/span>
&lt;hr>
&lt;h2 id="overall-workflow">
 Overall Workflow
 
 &lt;a class="anchor" href="#overall-workflow">#&lt;/a>
 
&lt;/h2>


&lt;pre class="mermaid">
flowchart LR
 A[Data] --&amp;gt; B[Explore Pattern]
 B --&amp;gt; C[Choose Model]
 C --&amp;gt; D[Train or Fit]
 D --&amp;gt; E[Validate]
 E --&amp;gt; F[Predict or Forecast]
 F --&amp;gt; G[Interpret Error]

 style A fill:#E1F5FE
 style B fill:#C8E6C9
 style C fill:#FFF9C4
 style D fill:#EDE7F6
 style E fill:#C8E6C9
 style F fill:#E1F5FE
 style G fill:#FFF9C4
&lt;/pre>

&lt;hr>
&lt;h2 id="correlation-">
 Correlation ☆
 
 &lt;a class="anchor" href="#correlation-">#&lt;/a>
 
&lt;/h2>
&lt;p>Correlation measures the direction and strength of linear relationship between two variables.&lt;/p></description></item><item><title>Gaussian Mixture Model &amp; Expectation Maximization</title><link>https://arshadhs.github.io/docs/ai/020-statistics/060-gaussian-mixture-model-em/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/020-statistics/060-gaussian-mixture-model-em/</guid><description>&lt;h1 id="gaussian-mixture-model--expectation-maximization">
 Gaussian Mixture Model &amp;amp; Expectation Maximization
 
 &lt;a class="anchor" href="#gaussian-mixture-model--expectation-maximization">#&lt;/a>
 
&lt;/h1>
&lt;p>A Gaussian Mixture Model represents data as a weighted combination of multiple Gaussian distributions.&lt;/p>
&lt;p>It is commonly used for soft clustering and density estimation.&lt;/p>
&lt;blockquote class="book-hint info">
&lt;p>&lt;strong>Key takeaway:&lt;/strong>&lt;br>
K-means gives hard cluster membership.&lt;/p>
&lt;p>GMM gives probabilities of belonging to each cluster.&lt;/p>
&lt;/blockquote>
&lt;ul>
&lt;li>Gaussian Mixture Model&lt;/li>
&lt;li>soft clustering&lt;/li>
&lt;li>mixing coefficients&lt;/li>
&lt;li>latent variables&lt;/li>
&lt;li>likelihood and log-likelihood&lt;/li>
&lt;li>Expectation-Maximization algorithm&lt;/li>
&lt;li>E-step and M-step&lt;/li>
&lt;li>responsibilities&lt;/li>
&lt;li>convergence&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="motivation-">
 Motivation ☆
 
 &lt;a class="anchor" href="#motivation-">#&lt;/a>
 
&lt;/h2>
&lt;p>Many real datasets are not described well by one Gaussian distribution.&lt;/p></description></item></channel></rss>