KEMBAR78
[FSDP2] Added check for contiguous parameters by awgu · Pull Request #137000 · pytorch/pytorch · GitHub
Skip to content

Conversation

@awgu
Copy link
Collaborator

@awgu awgu commented Sep 30, 2024

Stack from ghstack (oldest at bottom):

Since our implementation currently assumes contiguous strides, let us add an explicit check and raise an error at construction time if the parameter is not contiguous.

We can try to support this in the future. Mainly, I want to first learn more about how DTensor support for non-contiguous memory formats works.

cc @XilunWu @H-Huang @kwen2501 @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @c-p-i-o

@pytorch-bot
Copy link

pytorch-bot bot commented Sep 30, 2024

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/137000

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit b3b3882 with merge base 7ff8e66 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@pytorch-bot pytorch-bot bot added oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (fsdp) release notes category labels Sep 30, 2024
awgu pushed a commit that referenced this pull request Sep 30, 2024
ghstack-source-id: c4a2448
Pull Request resolved: #137000
@awgu awgu added release notes: distributed (fsdp2) release notes category and removed release notes: distributed (fsdp) release notes category labels Sep 30, 2024
@awgu awgu added the ciflow/trunk Trigger trunk jobs on your pull request label Sep 30, 2024
@awgu awgu marked this pull request as ready for review September 30, 2024 16:00
@awgu awgu requested a review from weifengpy September 30, 2024 16:00
@awgu
Copy link
Collaborator Author

awgu commented Sep 30, 2024

I am going to assume that we are not missing signal from this test Lint / Test run_test.py is usable without boto3/rockset:

ERROR: Could not find a version that satisfies the requirement torch (from versions: none)
ERROR: No matching distribution found for torch

@awgu
Copy link
Collaborator Author

awgu commented Sep 30, 2024

@pytorchbot merge -f "unrelated test failure"

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged immediately since you used the force (-f) flag, bypassing any CI checks (ETA: 1-5 minutes). Please use -f as last resort and instead consider -i/--ignore-current to continue the merge ignoring current failures. This will allow currently pending tests to finish and report signal before the merge.

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

AnantGulati pushed a commit to AnantGulati/pytorch that referenced this pull request Oct 2, 2024
Since our implementation currently assumes contiguous strides, let us add an explicit check and raise an error at construction time if the parameter is not contiguous.

We can try to support this in the future. Mainly, I want to first learn more about how DTensor support for non-contiguous memory formats works.

Pull Request resolved: pytorch#137000
Approved by: https://github.com/weifengpy
@github-actions github-actions bot deleted the gh/awgu/643/head branch November 3, 2024 02:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk Trigger trunk jobs on your pull request Merged oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (fsdp2) release notes category

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants