[CI][CD] Fix `install_nvshem` function #159907

malfet · 2025-08-05T22:07:20Z

When one builds CD docker, all CUDA dependencies must be installed into /usr/local/cuda/ folder

Test plan: Looks at the binary build logs, for example here:

2025-08-06T05:58:00.7347471Z -- NVSHMEM_HOME set to:  ''
2025-08-06T05:58:00.7348378Z -- NVSHMEM wheel installed at:  ''
2025-08-06T05:58:00.7392528Z -- NVSHMEM_HOST_LIB:  '/usr/local/cuda/lib64/libnvshmem_host.so'
2025-08-06T05:58:00.7393251Z -- NVSHMEM_DEVICE_LIB:  '/usr/local/cuda/lib64/libnvshmem_device.a'
2025-08-06T05:58:00.7393792Z -- NVSHMEM_INCLUDE_DIR:  '/usr/local/cuda/include'
2025-08-06T05:58:00.7394252Z -- NVSHMEM found, building with NVSHMEM support

When one builds CD docker, all CUDA dependencies must be installed into `/usr/local/cuda/` folder

pytorch-bot · 2025-08-05T22:07:23Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159907

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki or our office hours

Note: Links to docs will display an error until the docs builds have been completed.

❗ 1 Active SEVs

There are 1 currently active SEVs. If your PR is affected, please view them below:

ghstack-mergeability-check and Check labels failing with 'Resource not accessible by integration'

✅ You can merge normally! (1 Unrelated Failure)

As of commit 4076d96 with merge base 64cc6f0 ():

UNSTABLE - The following job is marked as unstable, possibly due to flakiness on trunk:

pull / linux-jammy-py3_9-clang9-xla / test (xla, 1, 1, lf.linux.12xlarge, unstable) (gh) (#158876)
/var/lib/jenkins/workspace/xla/torch_xla/csrc/runtime/BUILD:476:14: Compiling torch_xla/csrc/runtime/xla_util_test.cpp failed: (Exit 1): gcc failed: error executing CppCompile command (from target //torch_xla/csrc/runtime:xla_util_test) /usr/bin/gcc -U_FORTIFY_SOURCE -fstack-protector -Wall -Wunused-but-set-parameter -Wno-free-nonheap-object -fno-omit-frame-pointer -g0 -O2 '-D_FORTIFY_SOURCE=1' -DNDEBUG -ffunction-sections ... (remaining 229 arguments skipped)

This comment was automatically generated by Dr. CI and updates every 15 minutes.

.ci/docker/common/install_cuda.sh

malfet · 2025-08-06T14:42:41Z

@pytorchbot merge -f "Lint + binary builds are green"

pytorchmergebot · 2025-08-06T14:44:21Z

Merge started

Your change will be merged immediately since you used the force (-f) flag, bypassing any CI checks (ETA: 1-5 minutes). Please use -f as last resort and instead consider -i/--ignore-current to continue the merge ignoring current failures. This will allow currently pending tests to finish and report signal before the merge.

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

kwen2501 · 2025-08-12T20:34:06Z

Thanks for the fix! Indeed I wouldn't know CD has a different Docker build way than CI.

Just a minor thing -- I am actually a bit surprised that CMake can still find the header file after this move.
The following is the CMake scripts for the header search:

pytorch/caffe2/CMakeLists.txt

Lines 1014 to 1017 in 8e6a313

    
               find_path(NVSHMEM_INCLUDE_DIR 
        
                 NAMES nvshmem.h 
        
                 HINTS $ENV{NVSHMEM_HOME}/include ${NVSHMEM_PY_DIR}/include 
        
                 DOC "The location of NVSHMEM headers.")

Both NVSHMEM_HOME and NVSHMEM_PY_DIR are empty in your test above.
Yet CMake is still able to find the header under /usr/local/cuda. That seems to indicate that /usr/local/cuda is now one of the default search path of find_path?

When one builds CD docker, all CUDA dependencies must be installed into `/usr/local/cuda/` folder Test plan: Looks at the binary build logs, for example [here](https://github.com/pytorch/pytorch/actions/runs/16768141521/job/47477380147?pr=159907): ``` 2025-08-06T05:58:00.7347471Z -- NVSHMEM_HOME set to: '' 2025-08-06T05:58:00.7348378Z -- NVSHMEM wheel installed at: '' 2025-08-06T05:58:00.7392528Z -- NVSHMEM_HOST_LIB: '/usr/local/cuda/lib64/libnvshmem_host.so' 2025-08-06T05:58:00.7393251Z -- NVSHMEM_DEVICE_LIB: '/usr/local/cuda/lib64/libnvshmem_device.a' 2025-08-06T05:58:00.7393792Z -- NVSHMEM_INCLUDE_DIR: '/usr/local/cuda/include' 2025-08-06T05:58:00.7394252Z -- NVSHMEM found, building with NVSHMEM support ``` Pull Request resolved: pytorch#159907 Approved by: https://github.com/Skylion007, https://github.com/ngimel

[CI][CD] Fix install_nvshem function

ea7c28c

When one builds CD docker, all CUDA dependencies must be installed into `/usr/local/cuda/` folder

malfet requested a review from jeffdaily as a code owner August 5, 2025 22:07

malfet requested review from Skylion007 and removed request for jeffdaily August 5, 2025 22:07

pytorch-bot bot added the topic: not user facing topic category label Aug 5, 2025

malfet added the ciflow/trunk Trigger trunk jobs on your pull request label Aug 5, 2025

Skylion007 approved these changes Aug 5, 2025

View reviewed changes

ngimel approved these changes Aug 5, 2025

View reviewed changes

malfet commented Aug 6, 2025

View reviewed changes

.ci/docker/common/install_cuda.sh Outdated Show resolved Hide resolved

Update .ci/docker/common/install_cuda.sh

4076d96

pytorchmergebot added the merging label Aug 6, 2025

pytorchmergebot closed this in 2231c3c Aug 6, 2025

pytorchmergebot added Merged and removed merging labels Aug 6, 2025

ngimel added the ciflow/h100-symm-mem label Aug 6, 2025

ngimel mentioned this pull request Aug 6, 2025

[SymmMem] Add Triton 3.4 support to NVSHMEM Triton and fix CI tests (make device library discoverable + fix peer calculation bug) #159701

Closed

This was referenced Aug 12, 2025

[CD] nvshem-3.3.9 wheels for aarch64 is not manylinux2_28 compliant #160425

Closed

Use different include and libs path for nvshmem install on x86 and aarch64 builds #160433

Closed

github-actions bot deleted the malfet-patch-6 branch September 12, 2025 02:07

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[CI][CD] Fix `install_nvshem` function #159907

[CI][CD] Fix `install_nvshem` function #159907

Uh oh!

malfet commented Aug 5, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Aug 5, 2025 •

edited

Loading

Uh oh!

Uh oh!

malfet commented Aug 6, 2025

Uh oh!

pytorchmergebot commented Aug 6, 2025

Uh oh!

kwen2501 commented Aug 12, 2025 •

edited

Loading

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

[CI][CD] Fix install_nvshem function #159907

[CI][CD] Fix install_nvshem function #159907

Uh oh!

Conversation

malfet commented Aug 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Aug 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159907

❗ 1 Active SEVs

✅ You can merge normally! (1 Unrelated Failure)

Uh oh!

Uh oh!

malfet commented Aug 6, 2025

Uh oh!

pytorchmergebot commented Aug 6, 2025

Merge started

Uh oh!

kwen2501 commented Aug 12, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

[CI][CD] Fix `install_nvshem` function #159907

[CI][CD] Fix `install_nvshem` function #159907

malfet commented Aug 5, 2025 •

edited

Loading

pytorch-bot bot commented Aug 5, 2025 •

edited

Loading

kwen2501 commented Aug 12, 2025 •

edited

Loading