KEMBAR78
move thread-local capture mode guard to include work.isStarted by ngimel · Pull Request #160398 · pytorch/pytorch · GitHub
Skip to content

Conversation

@ngimel
Copy link
Collaborator

@ngimel ngimel commented Aug 12, 2025

Per title, should fix capture errors that happen because nccl watchdog races with capture start.

cc @H-Huang @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @pragupta

@pytorch-bot
Copy link

pytorch-bot bot commented Aug 12, 2025

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/160398

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d44850a with merge base ca7315c (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@pytorch-bot pytorch-bot bot added oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (c10d) release notes category labels Aug 12, 2025
@facebook-github-bot
Copy link
Contributor

@ngimel has imported this pull request. If you are a Meta employee, you can view this in D80063590.

@pytorch-bot pytorch-bot bot added the ciflow/trunk Trigger trunk jobs on your pull request label Aug 12, 2025
@ngimel
Copy link
Collaborator Author

ngimel commented Aug 12, 2025

@pytorchbot merge

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

@eqy
Copy link
Collaborator

eqy commented Aug 12, 2025

Couldn't repro the original condition in

def test_nccl_watchdog_cudagraph(self):
after playing around but the change looks fine

chuanhaozhuge pushed a commit that referenced this pull request Aug 14, 2025
Per title, should fix capture errors that happen because nccl watchdog races with capture start.

Pull Request resolved: #160398
Approved by: https://github.com/aorenste
chuanhaozhuge pushed a commit that referenced this pull request Aug 18, 2025
Per title, should fix capture errors that happen because nccl watchdog races with capture start.

Pull Request resolved: #160398
Approved by: https://github.com/aorenste
can-gaa-hou pushed a commit to can-gaa-hou/pytorch that referenced this pull request Aug 22, 2025
…ch#160398)

Per title, should fix capture errors that happen because nccl watchdog races with capture start.

Pull Request resolved: pytorch#160398
Approved by: https://github.com/aorenste
@github-actions github-actions bot deleted the ngimel/modeguard branch September 12, 2025 02:07
markc-614 pushed a commit to markc-614/pytorch that referenced this pull request Sep 17, 2025
…ch#160398)

Per title, should fix capture errors that happen because nccl watchdog races with capture start.

Pull Request resolved: pytorch#160398
Approved by: https://github.com/aorenste
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk Trigger trunk jobs on your pull request Merged oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (c10d) release notes category

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants