<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>M1: Foundations on Arshad Siddiqui</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/</link><description>Recent content in M1: Foundations on Arshad Siddiqui</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/index.xml" rel="self" type="application/rss+xml"/><item><title>ML and DL System Performance</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/010-ml-and-dl-system-performance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/010-ml-and-dl-system-performance/</guid><description>&lt;h1 id="ml-and-dl-system-performance">
 ML and DL System Performance
 
 &lt;a class="anchor" href="#ml-and-dl-system-performance">#&lt;/a>
 
&lt;/h1>
&lt;p>Machine learning system optimisation begins with measurement. Before changing an algorithm, adding processors, or moving work to a GPU, we need to understand &lt;strong>what is slow&lt;/strong>, &lt;strong>which resource is limiting performance&lt;/strong>, and &lt;strong>how performance changes as the workload grows&lt;/strong>.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Metrics:&lt;/strong> time complexity, running time, space, throughput, latency, and response time&lt;/li>
&lt;li>&lt;strong>Performance scaling and tuning:&lt;/strong> measuring bottlenecks and predicting how a workload grows&lt;/li>
&lt;li>&lt;strong>Training versus deployment:&lt;/strong> different objectives, workloads, and constraints&lt;/li>
&lt;li>&lt;strong>Deployment environments:&lt;/strong> distributed and cloud systems, embedded devices, and mobile systems&lt;/li>
&lt;/ol>
&lt;h2 id="learning-objectives">
 Learning Objectives
 
 &lt;a class="anchor" href="#learning-objectives">#&lt;/a>
 
&lt;/h2>
&lt;p>By the end of this page, you should be able to:&lt;/p></description></item><item><title>Parallel and Distributed Algorithms</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/020-parallel-and-distributed-algorithms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/020-parallel-and-distributed-algorithms/</guid><description>&lt;h1 id="parallel-and-distributed-algorithms">
 Parallel and Distributed Algorithms
 
 &lt;a class="anchor" href="#parallel-and-distributed-algorithms">#&lt;/a>
 
&lt;/h1>
&lt;p>Parallelisation divides computational work into parts that can execute concurrently. The purpose is to reduce completion time or increase throughput, but the gain depends on how much work is genuinely independent and how much overhead is introduced.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Systems and performance:&lt;/strong> execution time, throughput, efficiency, and bottlenecks&lt;/li>
&lt;li>&lt;strong>Speedup—approaches and issues:&lt;/strong> Amdahl&amp;rsquo;s Law, serial work, overhead, and load imbalance&lt;/li>
&lt;li>&lt;strong>Data parallelism versus task parallelism versus request parallelism&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Scale-out clusters:&lt;/strong> communication cost and its impact on speedup&lt;/li>
&lt;/ol>
&lt;p>The worked algorithms also show how divide-and-conquer and matrix operations expose parallel work.&lt;/p></description></item><item><title>Parallel Programming Models</title><link>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/030-parallel-programming-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arshadhs.github.io/docs/ai/038-ml-system-optimisation/m1-foundations/030-parallel-programming-models/</guid><description>&lt;h1 id="parallel-programming-models">
 Parallel Programming Models
 
 &lt;a class="anchor" href="#parallel-programming-models">#&lt;/a>
 
&lt;/h1>
&lt;p>Parallel algorithms need hardware that can execute independent work efficiently. Modern systems therefore combine multiple CPU cores, memory hierarchies, threads, instruction pipelines, GPUs, clusters, and specialised matrix processors.&lt;/p>
&lt;p>Course coverage:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Parallel programming models&lt;/strong> for expressing independent work&lt;/li>
&lt;li>&lt;strong>MapReduce&lt;/strong>, task-parallel, and request-parallel patterns&lt;/li>
&lt;li>&lt;strong>Multi-core CPUs and GPGPUs:&lt;/strong> SIMD, MIMD, SIMT, memory hierarchy, and bandwidth&lt;/li>
&lt;li>&lt;strong>Distributed clusters:&lt;/strong> CPU-only and GPU-accelerated systems, data sharding, parameter servers, and all-reduce&lt;/li>
&lt;/ol>
&lt;p>The hardware sections also cover specialised matrix processors and the performance limits introduced by memory and communication.&lt;/p></description></item></channel></rss>