KEMBAR78
[DeviceMesh] Fixed `from_group` when passing tensor `mesh` by awgu · Pull Request #137713 · pytorch/pytorch · GitHub
Skip to content

Conversation

@awgu
Copy link
Collaborator

@awgu awgu commented Oct 10, 2024

@pytorch-bot pytorch-bot bot added the oncall: distributed Add this issue/PR to distributed oncall triage queue label Oct 10, 2024
@pytorch-bot
Copy link

pytorch-bot bot commented Oct 10, 2024

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/137713

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 58e7c71 with merge base d1b87e2 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

awgu pushed a commit that referenced this pull request Oct 10, 2024
ghstack-source-id: 3c95471
Pull Request resolved: #137713
@awgu awgu marked this pull request as ready for review October 10, 2024 18:24
) or (mesh is not None and mesh != group_ranks):
) or (
mesh is not None
and not isinstance(mesh, torch.Tensor)
Copy link
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

add this check

@awgu awgu requested a review from wz337 October 10, 2024 18:24
@awgu awgu added ciflow/trunk Trigger trunk jobs on your pull request ciflow/periodic Trigger jobs ran periodically on master (periodic.yml) on the PR labels Oct 10, 2024
Copy link
Contributor

@wz337 wz337 left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fix!

@awgu
Copy link
Collaborator Author

awgu commented Oct 11, 2024

@pytorchbot merge

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

@pytorchmergebot
Copy link
Collaborator

Merge failed

Reason: 1 jobs have failed, first few of them are: periodic / parallelnative-linux-jammy-py3.9-gcc11 / test (default, 1, 3, linux.2xlarge)

Details for Dev Infra team Raised by workflow job

@awgu
Copy link
Collaborator Author

awgu commented Oct 11, 2024

@pytorchbot merge -f "unrelated default periodic test failing"

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged immediately since you used the force (-f) flag, bypassing any CI checks (ETA: 1-5 minutes). Please use -f as last resort and instead consider -i/--ignore-current to continue the merge ignoring current failures. This will allow currently pending tests to finish and report signal before the merge.

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

@github-actions github-actions bot deleted the gh/awgu/651/head branch November 11, 2024 02:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/periodic Trigger jobs ran periodically on master (periodic.yml) on the PR ciflow/trunk Trigger trunk jobs on your pull request Merged oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: DeviceMesh topic: bug fixes topic category

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants