Validate multi-rail routing before enabling RDMA on any non-primary subnet.
Don't assume a secondary subnet inherits the primary's RDMA configuration. It silently drops traffic.
Every lesson Portent applies is retrievable, actionable, and traced to a real event. It's categorized on four axes so knowledge learned in one place never leaks into a place it doesn't belong, and it grows only through a human-gated, ratcheted pipeline.
The tagging is what keeps an EFA lesson from being applied on InfiniBand, or a 70B-FSDP tuning from bleeding into an inference job. The axes and categories below are the real taxonomy; the bar magnitudes are illustrative until we publish the breakdown.
Each canonical lesson pairs an actionable rule with the mistake it prevents. Every customer, account, and host specific has been scrubbed.
Validate multi-rail routing before enabling RDMA on any non-primary subnet.
Don't assume a secondary subnet inherits the primary's RDMA configuration. It silently drops traffic.
Enumerate NICs on the node and verify the count before the workload starts.
Don't launch against 16 NICs because the instance type 'should' have 16. This one has 8.
Re-tune the NCCL algorithm and buffer sizes for the actual cluster size.
Don't reuse an 8-node ring configuration at 64 nodes; the topology assumptions no longer hold.
Verify the transport plugin's ABI matches the running NCCL build.
Don't accept a silent fallback to the socket transport as a 'slow' run; surface it.
Nothing enters the canonical corpus automatically. A candidate is mined, categorized, scrubbed, and verified; then a human approves the promote. The coverage ratchet guarantees every promoted lesson stays retrievable, actionable, and attributed, and the bar never lowers.
promote → recall: a promoted lesson is loaded on the next relevant run, closing the loop.