[SYNC] 2.22.3-1 #1426

BertanDogancay · 2024-11-18T16:13:58Z

Details

Do not mention proprietary info or link to internal work items in this PR.

Work item: "Internal", or link to GitHub issue (if applicable).

What were the changes?
One sentence describing the work done.

Why were the changes made?
Explain the motivation behind the work. Provide any publicly-available historical context.

How was the outcome achieved?
Technical details behind the work. Explain any publicly-available hardware peculiarities.

Additional Documentation:
What else should the reviewer know?

Approval Checklist

Do not approve until these items are satisfied.

Verify the CHANGELOG has been updated, if
- there are any NCCL API version changes,
- any changes impact library users, and/or
- any changes impact any other ROCm library.

Rework core for NVIDIA Trusted Computing * Compress work structs so that they are shared between channels * Utilize the full amount of kernel argument space permitted (4k) before resorting to work fifo. * Rework the task preprocessing phase. * Use a separate abortDevFlag which is kept in sync with abortFlag using cudaMemcpy operations. * Rename src/include/align.h to src/include/bitops.h Add lazy connection establishment for collective operations * Move buffer allocation and connection establishment to the first collective operation using that algorithm. * Accelerate init time and reduce memory usage. * Avoid allocating NVLS buffers if all calls are registered. * Compute algo/proto in ncclLaunchCollTasksInfo early on. * Connect peers in ncclCollPreconnectFunc if not connected already. * Also move shared buffer creation to the first send/recv call. Accelerate intra-node NVLink detection * Make each rank only detect NVLinks attached to its GPU. * Fuse XMLs to reconstruct the full NVLink topology Add init profiling to report time spend in different init phases. * Report timings of bootstrap, allgather, search, connect, etc. * Add new "PROFILE" category for NCCL_DEBUG_SUBSYS. Add support for PCI p2p on split PCI switches * Detect split PCI switches through a kernel module exposing switch information. * Update the topology XML and graph to add those inter-switch connections. Add cost estimation API * Add a new ncclGroupEndSimulate primitive to return the estimated time a group would take. Net/IB: Add separate traffic class for fifo messages * Add NCCL_IB_FIFO_TC to control the traffic class of fifo messages independently from NCCL_IB_TC. Merges PR ROCm#1194 Net/IB: Add support for IB router * Use flid instead of lid if subnets do not match * Warn if flid is 0 Optimizations and fixes for device network offload (unpack) * Double the default number of channels * Cache netDeviceType * Fix save/increment head logic to enable Tree support. Support ncclGroupStart/End for ncclCommAbort/Destroy * Allow Abort/Destroy to be called within a group when managing multiple GPUs with a single process. Improve Tuner API * Provide to the plugin the original cost table so that the plugin can leave unknown or disabled algo/proto combinations untouched. * Remove nvlsSupport and collnetSupport. Do not print version to stdout when using a debug file * Also print version from all processes with INFO debug level. Fixes issue ROCm#1271 Fix clang warnings in NVTX headers * Update NVTX headers to the latest version Fixes issue ROCm#1270 Disable port fusion in heterogeneous systems * Do not fuse ports if a mix of multi-port and single port are detected. Fix NVLS graphs search for dual NICs. * Fix NVLS graph search when we have more than one NIC per GPU. Fix crash with collnetDirect * Add separate graph search for collnetDirect, testing alltoall paths and working similarly to the NVLS search. Fix hang when nodes have different CPU types * Add the CPU type to the rank peer info. * Align all ranks on the CPU type after the first allgather. * Only use the aligned CPU type for all tuning operations. Fixes issue ROCm#1136 Fixes issue ROCm#1184 Fix performance of registered send/recv operations * Allow for single full size operations * Add INFO to confirm the registration of send/recv buffers. Move all sync ops to finalize stage * Ensure ncclCommDestroy is non-blocking if ncclCommFinalize has been called. Improve error reporting during SHM segment creation Improve support of various compilers Merges PR ROCm#1177 Merges PR ROCm#1228 Allow net and tuner plugins to be statically linked * Search for ncclNet or ncclTuner symbols in the main binary. Merges PR ROCm#979 Plugin examples includes cleanup * Harmonize err.h and common.h usage. * Add mixed plugin with both net and tuner.

BertanDogancay · 2024-12-18T23:45:38Z

I will clean up the history once the CI and other tests pass.

amd-jnovotny · 2024-12-19T13:34:43Z

@BertanDogancay It looks like I should probably update the develop version of https://rocm.docs.amd.com/projects/rccl/en/develop/how-to/using-nccl.html page to add the new feature? I guess you'll be adding something to the changelog during cleanup?

corey-derochie-amd · 2024-12-19T14:49:40Z

@BertanDogancay It looks like I should probably update the develop version of https://rocm.docs.amd.com/projects/rccl/en/develop/how-to/using-nccl.html page to add the new feature? I guess you'll be adding something to the changelog during cleanup?

That's correct, CHANGELOG, README, and docs should all be updated in this commit.

amd-jnovotny · 2024-12-19T15:20:47Z

@corey-derochie-amd If you'd like, I can create a new PR to update the docs (you're also welcome to add the changes and I can review as part of the PR) but 100% agree about updating CHANGELOG.md as part of this PR so the log and changes don't get separated.

sjeaugey and others added 18 commits June 14, 2024 01:57

Add decription for regIsGlobal in the NET API documentation

529ee69

Merge remote-tracking branch 'nccl/master' into develop

1b972cd

modify loadWorkBatchToShmem for WARP_SIZE of 64

cef4621

fix send/recv merge

aed4ced

fix multistream kernel launch

568eae6

Merge remote-tracking branch 'rccl/develop' into 2.22-sync

e842e08

msccl needs teardown before freeing comm

26128c8

missing chunksize optmizations

874f7fc

npkit fix

6be8ad4

Maintain extra version info

50433c1

fix colltrace

5ec2c59

use local tid when loading work batch to shmem

8a2f939

Merge remote-tracking branch 'rccl/develop' into 2.22-sync

315124b

update notices.txt

98578c1

Merge remote-tracking branch 'rccl/develop' into 2.22-sync

7695ff9

support up to 128 channels

3b082ca

Merge remote-tracking branch 'rccl/develop' into nccl-2.22-sync

2cbbee4

BertanDogancay added ci:extended gfx942 labels Nov 18, 2024

Allow max 2 work/batch

c6d742c

BertanDogancay force-pushed the nccl-2.22-sync branch from 31ed980 to c6d742c Compare December 18, 2024 06:22

BertanDogancay added 3 commits December 18, 2024 00:32

Map channels in sequential order

f042066

Fix partidx based on channel in device side

875a2b4

Merge remote-tracking branch 'rccl/develop' into nccl-2.22-sync

164d91a

BertanDogancay marked this pull request as ready for review December 18, 2024 23:41

BertanDogancay requested review from a team, wenkaidu, gilbertlee-amd and akolliasAMD as code owners December 18, 2024 23:41

BertanDogancay requested review from edgargabriel, PedramAlizadeh, nusislam, nileshnegi, KawtharShafie, AtlantaPepsi, mberenjk, corey-derochie-amd, mustafabar, thananon and haripriya-amd as code owners December 18, 2024 23:41

BertanDogancay added the gfx942-multinode label Dec 19, 2024

corey-derochie-amd mentioned this pull request Dec 19, 2024

Documentation for new NCCL Net plugin API change #1472

Open

1 task

BertanDogancay added 3 commits December 19, 2024 18:15

Covert to int8 in case of AlltoAllPivot

7a65e8f

Make adjustments for warp size 32

1d91c67

Temporarily disable alltoall pivot kernel

6e4ea12

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SYNC] 2.22.3-1 #1426

[SYNC] 2.22.3-1 #1426

BertanDogancay commented Nov 18, 2024

BertanDogancay commented Dec 18, 2024

amd-jnovotny commented Dec 19, 2024

corey-derochie-amd commented Dec 19, 2024

amd-jnovotny commented Dec 19, 2024

[SYNC] 2.22.3-1 #1426

Are you sure you want to change the base?

[SYNC] 2.22.3-1 #1426

Conversation

BertanDogancay commented Nov 18, 2024

Details

Approval Checklist

BertanDogancay commented Dec 18, 2024

amd-jnovotny commented Dec 19, 2024

corey-derochie-amd commented Dec 19, 2024

amd-jnovotny commented Dec 19, 2024