M3: Scale-Out Systems

M3: Scale-Out Systems #

This module examines the platforms, communication patterns, and hardware organisation used to train across larger systems. It extends distributed algorithms into coordinated training services and federated settings.

Module coverage:

  • ML Platforms and Frameworks: the Parameter Server model, Spark MLlib, and TensorFlow/PyTorch distributed training.
  • Distributed Deep Learning: decentralised SGD, all-reduce, asynchronous parallelism, and large-scale neural networks.
  • Federated Learning: training that preserves data privacy, communication efficiency, and hardware utilisation.

Learning Objectives #

  • Explain how training platforms organise model state, workers, and updates.
  • Compare centralised parameter services with collective communication.
  • Assess the effects of bandwidth, latency, locality, and stale updates.
  • Relate federated training to data distribution, privacy goals, and communication limits.

Pages and Planned Topics #

ChapterTopicAvailability
8ML Platforms and FrameworksPage available; currently focuses on the Parameter Server model
9Distributed Deep Learning: Synchronisation and CommunicationPage available; parameter-server synchronisation, compression, overlap, and placement
10Locality and Large-Scale ParallelismPlanned
11GPU Architecture and Federated LearningPlanned

The outline below describes the full module scope. Planned topics have no chapter page yet. Chapter 9 currently covers parameter-server execution policies and communication optimisation; collective communication, matrix multiplication, and Local SGD remain to be added from their source material.

Topic Outline #

8. ML Platforms and Frameworks #

  • Implementation issues and the Parameter Server model ☆
  • Stochastic gradient descent
  • TensorFlow architecture: master, worker, and parameter-server roles
  • Graph mode, eager execution, and distributed training strategies
  • Batch and micro-batch sizing
  • Model compression, quantisation, I/O, and shuffling overhead

9. Distributed Deep Learning #

  • Decentralised SGD, all-reduce, and asynchronous parallelism
  • System and program optimisation for neural networks
  • Matrix multiplication
  • Ring-based and tree-based reduction, bandwidth, and latency
  • Hogwild!, parameter staleness, convergence, and variance
  • Local SGD, periodic model averaging, and communication–computation trade-offs

10. Locality and Large-Scale Parallelism #

  • Locality-aware programs, massive multithreading, and GPGPUs
  • Model parallelism: layer slicing, tensor splitting, and cross-device communication
  • Pipeline parallelism: micro-batches, scheduling, and bubble overhead
  • Large-scale neural-network training on GPU clusters
  • Memory-efficient checkpointing

11. GPU Architecture and Federated Learning #

  • GPU architecture and threading models
  • Matrix addition and multithreaded matrix multiplication
  • Federated learning: non-IID data and privacy goals
  • Secure aggregation
  • Decentralised federated learning, gossip, and blockchain-based architectures

How This Module Connects #

Build on M2: Distributed ML Algorithms. For model optimisation and deployment under tighter resource limits, continue to M4: Constrained Systems.

ML System Optimisation overview


Home | ML System Optimisation