Setting Up a YAML Configuration¶
AFL-Sim supports launching simulations over the CLI or in a Python script using YAML configuration files. This guide details the structure of a configuration file and its parameters. It also explains how to recover standard algorithmic benchmarks by tuning the configuration parameters.
Overview¶
A configuration file defines all aspects of a fresh AFL-Sim run, including hardware/reproducibility settings, federated learning communication modes, dataset partitioning, simulated system latency, and checkpoint management.
A configuration file can be passed to the simulator using the CLI or inside a Python script:
Use the CLI command afl-sim run to launch a new simulation from a YAML configuration:
Info
Check the run command section in the CLI Reference for the full list of arguments.
To launch a new simulation from a YAML configuration in a Python script, import the run_simulation function from the afl_sim library:
| run_new_from_yaml.py | |
|---|---|
Info
For the full list of run_simulation arguments and defaults, check its documentation in the API Reference.
Configuration Template¶
A complete YAML configuration template is provided below. You can copy this block directly to create your own configuration file.
Parameter Overrides And Dry Runs
AFL-Sim allows overriding the YAML learning_rate parameter for easy sweeps. Moreover, you can quickly check the correctness of a YAML configuration without running a full simulation by performing a dry run.
Use the --lr and --dry-run flags to override the YAML learning_rate and perform a dry run, respectively:
Parameter Reference¶
Dataset¶
| Parameter | Type | Default | Description |
|---|---|---|---|
dataset |
String | mnist |
The target dataset for the simulation. Supported options: (mnist, fashion_mnist, cifar10, cifar100). |
dirichlet_alpha |
Float | 0.1 |
Dirichlet distribution parameter. |
split_seed |
Integer | 42 |
The random seed ensuring reproducibility during the dataset partitioning process. |
Tip
Higher values of dirichlet_alpha produce more homogeneous data distributions across clients, moving closer to an Independent and Identically Distributed (IID) setting.
Warning
AFL-Sim requires each client to have a minimum number of samples equal to the training batch size. Because the training dataset is finite, this constraint becomes harder to satisfy if you increase the number of clients or decrease the Dirichlet parameter. If the simulator exceeds its maximum attempts to generate a valid data split, it will abort and prompt you to increase dirichlet_alpha.
Simulation Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
device |
String | auto |
The hardware accelerator used for the simulation. Options: (auto, cpu, cuda, mps). |
num_clients |
Integer | 10 |
Total number of clients. |
timeout_seconds |
Float | 300.0 |
Simulation duration in wall-clock seconds. |
client_rate_std |
Float | 1.0 |
Standard deviation of client latency. |
rate_seed |
Integer | 42 |
The random seed used to generate client arrival times and latency distributions. |
torch_seed |
Integer | 42 |
The random seed for all PyTorch operations. |
Clock Tips
- Decreasing
client_rate_stdreduces the variance in client latency, resulting in more consistent response times across the network (i.e., reducing the "straggler effect"). - The median client latency is hardcoded to
1simulation time unit.
Info
Setting the device parameter to auto selects the fastest available hardware accelerator. For example, AFL-Sim will prioritize cuda (NVIDIA GPUs) or mps (Apple Silicon) over standard cpu execution if those accelerators are detected.
Model¶
| Parameter | Type | Default | Description |
|---|---|---|---|
model_name |
String | cnn |
Model architecture to use. Options: (logreg, cnn, resnet18). |
Warning
AFL-Sim enforces strict dataset-model compatibility rules. While logreg and cnn can be paired with any supported dataset, resnet18 can only be used with cifar10 and cifar100. Attempting to pair resnet18 with mnist or fashion_mnist will cause the configuration validation to fail.
Client Memory Augmentation¶
| Parameter | Type | Default | Description |
|---|---|---|---|
type |
String | disabled |
Type of client memory augmentation (disabled, models, gradients). |
The type parameter defines how clients compute the updates they send to the server:
disabled: Standard Federated Learning. Clients send the local delta (i.e., the difference between their newly trained local model and the initial global model received before training).models: Clients maintain a local history, updating the server with the difference between their most recent trained local model and their second most recent trained local model.gradients: Clients maintain a gradient (delta) history, updating the server with the difference between their most recent local deltas and their second most recent local deltas.
Note
AFL-Sim implements memory augmentation strictly on the client side. However, this approach is mathematically equivalent to a server-side memory architecture where the server maintains a dedicated memory buffer for every individual client. In that alternative setup, clients would simply send their freshest updates (models or gradients), and the server would track the stale versions to compute the differences itself.
Communication Strategy¶
Failure
You must define exactly one communication strategy block (async OR sync) per configuration file. Attempting to define both will cause the configuration validation to fail and abort the simulation.
| Parameter | Type | Default | Description |
|---|---|---|---|
type |
String | async |
Asynchronous FL mode. |
buffer_size |
Integer | 3 |
Number of client updates that triggers a global model update. |
| Parameter | Type | Default | Description |
|---|---|---|---|
type |
String | sync |
Synchronous FL mode. |
sample_size |
Integer | 3 |
Number of clients sampled by the server at each round. |
In synchronous mode the server selects the sample_size subset of clients uniformly at random (i.e., with equal probability). In asynchronous mode, there is no central sampling; the server processes updates as they arrive in its buffer, and updates the global model after receiving a number of updates equal to buffer_size.
Client Optimizer Settings¶
| Parameter | Type | Default | Description |
|---|---|---|---|
learning_rate |
Float | 0.1 |
The step size applied during local client training. |
weight_decay |
Float | 0.0 |
The L2 penalty (weight decay) applied by the PyTorch optimizer to prevent overfitting. |
num_local_steps |
Integer | 100 |
The exact number of local SGD steps (batches) a client performs before communicating with the server. |
batch_size |
Integer | 32 |
The number of samples processed per local training step. |
Note
The local optimization algorithm is pre-set to Stochastic Gradient Descent (SGD). Adaptive algorithms such as Adam and its variants are not currently supported for local client training. In non-IID federated settings, naively applying adaptive optimizers locally causes objective inconsistency, arbitrarily distorting the geometry of the global objective function.
Evaluation Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
batch_size |
Integer | 32 |
The number of test dataset samples processed per batch during global model evaluation (for metric generation). |
num_workers |
Integer | 0 |
The number of subprocesses used for data loading, corresponding to the PyTorch DataLoader parameter. |
Tip
Leaving num_workers at 0 means data loading occurs sequentially on the main process. If you notice that evaluating the global model is creating a bottleneck and slowing down your simulation speed, try increasing this value (e.g., to 2 or 4) to enable multi-process data loading.
Checkpoints¶
| Parameter | Type | Default | Description |
|---|---|---|---|
keep_best |
Boolean | False |
If set to True, the simulator continuously saves a separate copy of the global model that achieved the highest accuracy on the test set. |
interval_seconds |
Float | 400.0 |
The interval (in wall-clock seconds) at which the simulator saves a resumable checkpoint. |
Tip
To effectively disable periodic checkpoints, set interval_seconds to a value greater than timeout_seconds (defined in the simulation parameters). Regardless of this interval, AFL-Sim will always attempt to save a resumable checkpoint once the simulation timeout is reached, or in the event of a user/system interruption (e.g., Ctrl+C).
Visualization¶
| Parameter | Type | Default | Description |
|---|---|---|---|
visualize_data_split |
Boolean | False |
Generates and saves a chart in .png format illustrating the distribution of dataset samples across the clients. |
visualize_client_arrivals |
Boolean | False |
Generates and saves a timeline plot in .png format depicting the simulated arrival times and latencies of the clients. |
Note
Visualizations are created with Matplotlib.
Implemented Algorithms¶
AFL-Sim provides a unified framework capable of recovering equivalent versions of several standard Federated Learning and distributed/parallel SGD algorithmic benchmarks. This is achieved by tuning three key configuration parameters: communication strategy, memory strategy, and the number of local steps.
Before mapping configurations to specific algorithms, it is important to understand how AFL-Sim's architecture manages communication and memory:
| Functionality | Implementation In AFL-Sim |
|---|---|
| Asynchrony | AFL-Sim's asynchronous mode is strictly client-driven. At no point does the server actively sample clients to send them the global model. Instead, the server broadcasts the global model to all clients once at initialization, and then accepts incoming client updates as they arrive. Whenever a client returns an update to the server, it receives the latest global model in return. |
| Memory Augmentation | With the exception of AREA, the algorithms listed in the table below natively implement memory augmentation on the server side, allocating dedicated buffers for each client. AFL-Sim, conversely, implements memory augmentation strictly on the client side. As noted in the Client Memory Augmentation section, the two architectures are mathematically equivalent. However, in real-world systems, client-side augmentation could drastically reduce memory requirements at the server and enhance scalability as the number of clients grows, provided clients have enough disk space to store a copy of the model. |
Info
For a deeper dive into the specific order of operations in asynchronous mode, refer to the Simulated Federated Learning Workflows section in the implementation notes of AFL-Sim.
Algorithm Recovery Matrix¶
Although the underlying memory-augmentation (or lack thereof) remains equivalent up to constants, due to the architectural deviations in the communication protocol and in the location of memory-augmentation described above, AFL-Sim can be said to return methods that are "like" the algorithms below (e.g., FedBuff-like), rather than exact replicas. The exception is AREA, which utilizes the communication and memory architecture native to AFL-Sim.
comm_strategy.type |
mem_strategy.type |
optimization.num_local_steps |
Recovered Algorithm |
|---|---|---|---|
sync |
disabled |
1 |
Synchronous SGD |
async |
disabled |
1 |
Asynchronous SGD |
sync |
disabled |
> 1 |
FedAvg |
async |
disabled |
> 1 |
FedBuff-like, AFA-CD-like |
sync |
gradients |
> 1 |
MIFA-like, FedVARP-like |
async |
gradients |
> 1 |
AFA-CS-like, CA2FL-like |
async |
models |
> 1 |
AREA |
Note on MIFA-like Recovery
MIFA was proposed to tackle client unavailability in the synchronous setting, hence this recovery assumes that the sampled clients are the available clients at each round. Because AFL-Sim hardcodes uniform client sampling under synchronous communication, this yields a particularly benign version of client unavailability where clients are equally likely to be unavailable.