Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
67
80%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./skills/nemo-mbridge-perf-tp-dp-comm-overlap/SKILL.mdFor stable background and recommendation level, see:
Minimal Bridge override:
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(
tp_comm_overlap=True,
)
cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = TrueOptional TP preset:
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048Precision knobs belong to mixed precision:
cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = FalseBridge overlap gating:
if self.user_comm_overlap_cfg.tp_comm_overlap is True:
if model_cfg.tensor_model_parallel_size < 2:
...
elif not model_cfg.sequence_parallel:
...
elif not HAVE_TE:
...PP overlap selection:
if model_cfg.pipeline_model_parallel_size > 1:
if vp_size > 1:
comm_overlap_cfg.overlap_p2p_comm = True
comm_overlap_cfg.batch_p2p_comm = False
else:
comm_overlap_cfg.overlap_p2p_comm = False
comm_overlap_cfg.batch_p2p_comm = TrueDP overlap defaults:
if self.data_parallel_size > 1:
comm_overlap_cfg.bucket_size = 128 * 1024 * 1024
comm_overlap_cfg.overlap_grad_reduce = True
comm_overlap_cfg.overlap_param_gather = TrueLaunch-time env tuning:
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)sequence_parallel=False or Transformer Engine is unavailable.overlap_p2p_comm=True when PP > 1 and VPP > 1.bucket_size is a parameter-count knob, not a byte-size knob.grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.Use the checked-in overlap unit coverage first:
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -qOptional second check if nemo_run is available:
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -qSuccess criteria:
26 passede3796e4
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.