<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>M2: Distributed ML Algorithms on Arshad Siddiqui</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/</link><description>Recent content in M2: Distributed ML Algorithms on Arshad Siddiqui</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/index.xml" rel="self" type="application/rss+xml"/><item><title>Parallelisation of ML Algorithms</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/040-parallelisation-of-ml-algorithms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/040-parallelisation-of-ml-algorithms/</guid><description>&lt;h1 id="parallelisation-of-ml-algorithms">
 Parallelisation of ML Algorithms
 
 &lt;a class="anchor" href="#parallelisation-of-ml-algorithms">#&lt;/a>
 
&lt;/h1>
&lt;p>Machine-learning algorithms expose different kinds of parallel work. The correct decomposition depends on whether independent work occurs across records, features, trees, clusters, kernel entries, or optimisation steps.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Problem decomposition&lt;/strong> for parallel machine learning&lt;/li>
&lt;li>&lt;strong>Ensemble methods and XGBoost-style tree ensembles&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Parallel k-means&lt;/strong> assignment and centroid reduction&lt;/li>
&lt;li>&lt;strong>Distributed decision trees and random forests&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Support vector machine parallelisation&lt;/strong> using block and MapReduce-style computation&lt;/li>
&lt;/ol>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p></description></item><item><title>Communication-Aware Distributed ML</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/050-communication-aware-distributed-ml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/050-communication-aware-distributed-ml/</guid><description>&lt;h1 id="communication-aware-distributed-ml">
 Communication-Aware Distributed ML
 
 &lt;a class="anchor" href="#communication-aware-distributed-ml">#&lt;/a>
 
&lt;/h1>
&lt;p>Distributed machine learning is effective only when saved computation exceeds the cost of moving data and synchronising workers. Algorithm design must therefore account for message size, message frequency, barriers, and parameter placement.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Communication overhead&lt;/strong>, illustrated through distributed k-means&lt;/li>
&lt;li>&lt;strong>Model parallelism&lt;/strong> when parameters do not fit or compute must be divided&lt;/li>
&lt;li>&lt;strong>Distributed k-nearest neighbours:&lt;/strong> partitioning, indexing, and approximate search&lt;/li>
&lt;li>&lt;strong>Gradient descent, SGD, and mini-batch optimisation&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Synchronous versus asynchronous updates&lt;/strong>&lt;/li>
&lt;/ol>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p></description></item><item><title>Clusters, Hadoop, and Spark</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/060-clusters-hadoop-and-spark/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/060-clusters-hadoop-and-spark/</guid><description>&lt;h1 id="clusters-hadoop-and-spark">
 Clusters, Hadoop, and Spark
 
 &lt;a class="anchor" href="#clusters-hadoop-and-spark">#&lt;/a>
 
&lt;/h1>
&lt;p>Scale-out frameworks distribute data and computation across networked machines. Their value comes from aggregate compute and storage, while their main costs are communication, scheduling, serialisation, and repeated synchronisation.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Cluster computing with Hadoop and Spark&lt;/strong>&lt;/li>
&lt;li>&lt;strong>MapReduce k-means&lt;/strong>, mini-batch and streaming variants&lt;/li>
&lt;li>&lt;strong>Convergence and communication reduction&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Distributed convolutional neural networks:&lt;/strong> data and model parallelism&lt;/li>
&lt;li>&lt;strong>Layer or tensor slicing, mixed precision, pipelining, and distributed execution frameworks&lt;/strong>&lt;/li>
&lt;/ol>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p></description></item><item><title>Distributed Training Strategies</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/070-distributed-training-strategies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m2-distributed-ml-algorithms/070-distributed-training-strategies/</guid><description>&lt;h1 id="distributed-training-strategies">
 Distributed Training Strategies
 
 &lt;a class="anchor" href="#distributed-training-strategies">#&lt;/a>
 
&lt;/h1>
&lt;p>Distributed training combines devices to increase throughput, fit larger models, or shorten time to solution. The central design question is what to partition: examples, model parameters, layers, tensors, or the training schedule.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Data parallelism&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Model parallelism&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Pipeline parallelism and micro-batches&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Gradient checkpointing&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Mixed-precision training and loss scaling&lt;/strong>&lt;/li>
&lt;/ol>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p>
&lt;ul>
&lt;li>compare data, model, and pipeline parallelism&lt;/li>
&lt;li>calculate global batch size, speedup, efficiency, and effective throughput&lt;/li>
&lt;li>explain gradient aggregation through all-reduce&lt;/li>
&lt;li>calculate simple pipeline time and utilisation&lt;/li>
&lt;li>quantify the memory–computation trade-off of checkpointing&lt;/li>
&lt;li>explain mixed precision and loss scaling&lt;/li>
&lt;/ul>
&lt;h2 id="big-picture">
 Big Picture
 
 &lt;a class="anchor" href="#big-picture">#&lt;/a>
 
&lt;/h2>


&lt;pre class="mermaid">
flowchart TD
 A[&amp;#34;Training Constraint&amp;#34;] --&amp;gt; B[&amp;#34;More Data Throughput&amp;#34;]
 A --&amp;gt; C[&amp;#34;Model Too Large&amp;#34;]
 A --&amp;gt; D[&amp;#34;Activation Memory Too High&amp;#34;]
 B --&amp;gt; E[&amp;#34;Data Parallelism&amp;#34;]
 C --&amp;gt; F[&amp;#34;Model or Pipeline Parallelism&amp;#34;]
 D --&amp;gt; G[&amp;#34;Checkpointing or Mixed Precision&amp;#34;]

 style A fill:#E1F5FE
 style B fill:#C8E6C9
 style C fill:#FFF9C4
 style D fill:#EDE7F6
 style E fill:#E1F5FE
 style F fill:#C8E6C9
 style G fill:#FFF9C4
&lt;/pre>

&lt;h2 id="1-data-parallelism-">
 1. Data Parallelism ☆
 
 &lt;a class="anchor" href="#1-data-parallelism-">#&lt;/a>
 
&lt;/h2>
&lt;p>Each worker holds a complete model replica and processes a different shard of the mini-batch.&lt;/p></description></item></channel></rss>