<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>M3: Scale-Out Systems on Arshad Siddiqui</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/</link><description>Recent content in M3: Scale-Out Systems on Arshad Siddiqui</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/index.xml" rel="self" type="application/rss+xml"/><item><title>ML Platforms and Frameworks</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/080-ml-platforms-and-frameworks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/080-ml-platforms-and-frameworks/</guid><description>&lt;h1 id="ml-platforms-and-frameworks">
 ML Platforms and Frameworks
 
 &lt;a class="anchor" href="#ml-platforms-and-frameworks">#&lt;/a>
 
&lt;/h1>
&lt;p>Training at scale requires more than dividing data between GPUs. A platform must decide &lt;strong>where model state lives, who computes gradients, how updates are combined, and how information moves between devices&lt;/strong>.&lt;/p>
&lt;p>This page focuses on the &lt;strong>Parameter Server model&lt;/strong> and the practical constraints that determine whether adding resources makes training faster.&lt;/p>
&lt;p>Topics covered:&lt;/p>
&lt;ul>
&lt;li>Parameter servers, workers, and independent scaling of storage and computation&lt;/li>
&lt;li>Training-state memory and parameter sharding&lt;/li>
&lt;li>Distributed loss, gradient aggregation, and the pull–compute–push cycle&lt;/li>
&lt;li>Communication costs, RDMA, MPI, PCIe, and NVLink&lt;/li>
&lt;li>Global batch size, learning-rate scaling, and warm-up&lt;/li>
&lt;li>Monitoring, fault recovery, and sparse versus dense workloads&lt;/li>
&lt;/ul>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p></description></item><item><title>Distributed Deep Learning: Synchronisation and Communication</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/090-distributed-deep-learning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/090-distributed-deep-learning/</guid><description>&lt;h1 id="distributed-deep-learning-synchronisation-and-communication">
 Distributed Deep Learning: Synchronisation and Communication
 
 &lt;a class="anchor" href="#distributed-deep-learning-synchronisation-and-communication">#&lt;/a>
 
&lt;/h1>
&lt;p>Distributed training divides computation between workers, but those workers rarely finish together. A training system must decide &lt;strong>when to update the model, how much outdated information to accept, and how to reduce communication&lt;/strong>.&lt;/p>
&lt;p>Relevant themes from the content outline:&lt;/p>
&lt;ul>
&lt;li>Asynchronous parallelism and parameter staleness&lt;/li>
&lt;li>Effects on convergence and variance&lt;/li>
&lt;li>System and program optimisation for neural networks&lt;/li>
&lt;li>Communication–computation trade-offs&lt;/li>
&lt;/ul>
&lt;!-- Outline alignment: _COURSE_MLSysOps_ZG516.pdf, page 6. This page covers the parameter-server material supplied for this topic. Collective communication, matrix multiplication, and Local SGD remain separate coverage. -->
&lt;p>The &lt;a href="https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m3-scale-out-systems/080-ml-platforms-and-frameworks/">Parameter Server model&lt;/a> provides the architectural foundation. Here, the focus is on the policies that determine its performance.&lt;/p></description></item></channel></rss>