[DTensor] Cached hash for `DTensorSpec` #113915

awgu · 2023-11-17T03:06:00Z

Stack from ghstack (oldest at bottom):

Overview
Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing DTensorSpec, which turns out to be relatively expensive.

Details
It looks like tensor_meta is only updated in _wrap_output_spec_tensor_meta, which only runs if the propagation was not cached:

pytorch/torch/distributed/_tensor/sharding_prop.py

Line 137 in ae94c7e

output_spec.tensor_meta = output_tensor_meta

pytorch/torch/distributed/_tensor/sharding_prop.py

Line 153 in ae94c7e

spec.tensor_meta = output_tensor_meta_i

In that case, I think we can cache the hash for the DTensorSpec and only update it when one of the hashed attributes changes, which we only really expect to happen for tensor_meta.

To ensure correctness, we need that all hashed attributes are immutable.

DeviceMesh caches its hash:

pytorch/torch/distributed/_device_mesh.py

Line 181 in a9134fa

self._hash = hash((self._flatten_mesh_list, self.mesh.shape))
This PR makes each Placement a frozen dataclass, making them immutable (relying on the fact that they do not have references to any mutable objects).

TensorMeta is a NamedTuple of torch.Size, Tuple[int, ...], and torch.dtype, so it is immutable:

pytorch/torch/distributed/_tensor/placement_types.py

Lines 369 to 375 in 9916d8a

    
           class TensorMeta(NamedTuple): 
        
               # simple named tuple to represent tensor metadata 
        
               # intentionally to stay simple only for sharding 
        
               # propagation purposes. 
        
               shape: torch.Size 
        
               stride: Tuple[int, ...] 
        
               dtype: torch.dtype

Example
For some simple small GPT model:
Before: 0.125 ms

After: 0.048 ms

The overall Adam CPU step time decreases from 7.647 ms to 6.451 ms.

[ghstack-poisoned]

pytorch-bot · 2023-11-17T03:06:03Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/113915

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki or our office hours

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 9fd0091 with merge base 140c54e ():
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

[ghstack-poisoned]

Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing `DTensorSpec`, which turns out to be relatively expensive. It looks like `tensor_meta` is only updated in `_wrap_output_spec_tensor_meta`, which only runs if the propagation was not cached: https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L137 https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L153 In that case, I think we can cache the hash for the `DTensorSpec` and only update it when one of the hashed attributes changes, which we only really expect to happen for `tensor_meta`. --- For some simple small GPT model: Before: 0.125 ms <img width="509" alt="Screenshot 2023-11-16 at 10 08 05 PM" src="https://github.com/pytorch/pytorch/assets/31054793/10e59401-f635-431f-80b5-1b48df3a706e"> After: 0.048 ms <img width="294" alt="Screenshot 2023-11-16 at 10 08 47 PM" src="https://github.com/pytorch/pytorch/assets/31054793/09a3b0b9-f68c-4afc-bca1-c29a4b01c2fb"> [ghstack-poisoned]

**Overview** Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing `DTensorSpec`, which turns out to be relatively expensive. **Details** It looks like `tensor_meta` is only updated in `_wrap_output_spec_tensor_meta`, which only runs if the propagation was not cached: https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L137 https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L153 In that case, I think we can cache the hash for the `DTensorSpec` and only update it when one of the hashed attributes changes, which we only really expect to happen for `tensor_meta`. To ensure correctness, we need that all hashed attributes are immutable. - `DeviceMesh` caches its hash: https://github.com/pytorch/pytorch/blob/a9134fa99a8986adf478a12db2ea5729d24554db/torch/distributed/_device_mesh.py#L181 - This PR makes each `Placement` a frozen `dataclass`, making them immutable (relying on the fact that they do not have references to any mutable objects). - `TensorMeta` is a `NamedTuple` of `torch.Size`, `Tuple[int, ...]`, and `torch.dtype`, so it is immutable: https://github.com/pytorch/pytorch/blob/9916d8a9eaaf2c05c131f2a2dbe9eabeeaa9dffc/torch/distributed/_tensor/placement_types.py#L369-L375 **Example** For some simple small GPT model: Before: 0.125 ms <img width="509" alt="Screenshot 2023-11-16 at 10 08 05 PM" src="https://github.com/pytorch/pytorch/assets/31054793/10e59401-f635-431f-80b5-1b48df3a706e"> After: 0.048 ms <img width="294" alt="Screenshot 2023-11-16 at 10 08 47 PM" src="https://github.com/pytorch/pytorch/assets/31054793/09a3b0b9-f68c-4afc-bca1-c29a4b01c2fb"> The overall Adam CPU step time decreases from 7.647 ms to 6.451 ms. [ghstack-poisoned]

awgu · 2023-11-17T05:22:19Z

CUDA_VISIBLE_DEVICES=2,3 python -m pytest test/distributed/_tensor/test_tensor_ops.py -k test_index

RuntimeError: Process 0 exited with error code 10 and exception:
Traceback (most recent call last):
  File "/data/users/andgu/pytorch/torch/testing/_internal/common_distributed.py", line 658, in run_test
    getattr(self, test_name)()
  File "/data/users/andgu/pytorch/torch/testing/_internal/common_distributed.py", line 544, in wrapper
    fn()
  File "/data/users/andgu/pytorch/torch/testing/_internal/common_utils.py", line 2536, in wrapper
    method(*args, **kwargs)
  File "/data/users/andgu/pytorch/torch/testing/_internal/distributed/_tensor/common_dtensor.py", line 193, in wrapper
    func(self, *args, **kwargs)  # type: ignore[misc]
  File "/data/users/andgu/pytorch/test/distributed/_tensor/test_tensor_ops.py", line 307, in test_index
    self._test_op(
  File "/data/users/andgu/pytorch/test/distributed/_tensor/test_tensor_ops.py", line 279, in _test_op
    self.assertEqual(d_out.full_tensor(), out)
  File "/data/users/andgu/pytorch/torch/testing/_internal/common_utils.py", line 3439, in assertEqual
    raise error_metas.pop()[0].to_error(
AssertionError: The values for attribute 'shape' do not match: torch.Size([12, 16, 32]) != torch.Size([12, 32, 16]).

Update: I think I fixed it. I did not account for the case when a DTensorSpec could be constructed with tensor_meta directly specified (I incorrectly thought it would only be set after).

**Overview** Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing `DTensorSpec`, which turns out to be relatively expensive. **Details** It looks like `tensor_meta` is only updated in `_wrap_output_spec_tensor_meta`, which only runs if the propagation was not cached: https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L137 https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L153 In that case, I think we can cache the hash for the `DTensorSpec` and only update it when one of the hashed attributes changes, which we only really expect to happen for `tensor_meta`. To ensure correctness, we need that all hashed attributes are immutable. - `DeviceMesh` caches its hash: https://github.com/pytorch/pytorch/blob/a9134fa99a8986adf478a12db2ea5729d24554db/torch/distributed/_device_mesh.py#L181 - This PR makes each `Placement` a frozen `dataclass`, making them immutable (relying on the fact that they do not have references to any mutable objects). - `TensorMeta` is a `NamedTuple` of `torch.Size`, `Tuple[int, ...]`, and `torch.dtype`, so it is immutable: https://github.com/pytorch/pytorch/blob/9916d8a9eaaf2c05c131f2a2dbe9eabeeaa9dffc/torch/distributed/_tensor/placement_types.py#L369-L375 **Example** For some simple small GPT model: Before: 0.125 ms <img width="509" alt="Screenshot 2023-11-16 at 10 08 05 PM" src="https://github.com/pytorch/pytorch/assets/31054793/10e59401-f635-431f-80b5-1b48df3a706e"> After: 0.048 ms <img width="294" alt="Screenshot 2023-11-16 at 10 08 47 PM" src="https://github.com/pytorch/pytorch/assets/31054793/09a3b0b9-f68c-4afc-bca1-c29a4b01c2fb"> The overall Adam CPU step time decreases from 7.647 ms to 6.451 ms. [ghstack-poisoned]

ghstack-source-id: 633bac0 Pull Request resolved: #113915

wanchaol

Awesome! great to see we can precompute DTensorSpec hash and save a such big CPU overhead!

**Overview** Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing `DTensorSpec`, which turns out to be relatively expensive. **Details** It looks like `tensor_meta` is only updated in `_wrap_output_spec_tensor_meta`, which only runs if the propagation was not cached: https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L137 https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L153 In that case, I think we can cache the hash for the `DTensorSpec` and only update it when one of the hashed attributes changes, which we only really expect to happen for `tensor_meta`. To ensure correctness, we need that all hashed attributes are immutable. - `DeviceMesh` caches its hash: https://github.com/pytorch/pytorch/blob/a9134fa99a8986adf478a12db2ea5729d24554db/torch/distributed/_device_mesh.py#L181 - This PR makes each `Placement` a frozen `dataclass`, making them immutable (relying on the fact that they do not have references to any mutable objects). - `TensorMeta` is a `NamedTuple` of `torch.Size`, `Tuple[int, ...]`, and `torch.dtype`, so it is immutable: https://github.com/pytorch/pytorch/blob/9916d8a9eaaf2c05c131f2a2dbe9eabeeaa9dffc/torch/distributed/_tensor/placement_types.py#L369-L375 **Example** For some simple small GPT model: Before: 0.125 ms <img width="509" alt="Screenshot 2023-11-16 at 10 08 05 PM" src="https://github.com/pytorch/pytorch/assets/31054793/10e59401-f635-431f-80b5-1b48df3a706e"> After: 0.048 ms <img width="294" alt="Screenshot 2023-11-16 at 10 08 47 PM" src="https://github.com/pytorch/pytorch/assets/31054793/09a3b0b9-f68c-4afc-bca1-c29a4b01c2fb"> The overall Adam CPU step time decreases from 7.647 ms to 6.451 ms. [ghstack-poisoned]

awgu · 2023-11-21T01:20:50Z

@pytorchbot merge

pytorchmergebot · 2023-11-21T01:23:31Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

This is a nit change to save one `isinstance` call for when `dim` is not `None` but the placement is not `Shard`. Pull Request resolved: #114140 Approved by: https://github.com/Skylion007, https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930, #114141, #113915

This is a forward fix for #113781. We lazily compute the hash so that we do not try to compute the hash on `SymInt`s (for the stride) during Dynamo tracing. Tested via: ``` python test/distributed/_tensor/test_dtensor_compile.py -k test_2d_fsdp_tp_ac_compile ``` Pull Request resolved: #114322 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930, #114141, #113915, #114140

[DTensor] Cached hash for DTensorSpec

594cc8c

[ghstack-poisoned]

Update on "[DTensor] Cached hash for DTensorSpec"

90b1538

[ghstack-poisoned]

awgu requested a review from wanchaol November 17, 2023 03:20

awgu mentioned this pull request Nov 17, 2023

[DTensor] Renamed shard_spec -> placements in test file #113917

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Made _Partial, Replicate frozen dataclasses #113919

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Returned new placements for neg dim in global info #113922

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Used new placements for neg dim in redistribute #113924

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Ensured grad_placements was tuple #113925

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Used new placements for neg dim in distribute_tensor #113930

Closed

awgu pushed a commit that referenced this pull request Nov 17, 2023

[DTensor] Cached hash for DTensorSpec

b357350

ghstack-source-id: 633bac0 Pull Request resolved: #113915

wanchaol approved these changes Nov 19, 2023

View reviewed changes

awgu mentioned this pull request Nov 20, 2023

[DTensor] Used new placements for neg dim in from_local #114134

Closed

This was referenced Nov 20, 2023

[DTensor] Reduced to one isinstance call in is_shard #114140

Closed

[DTensor] Replaced neg dim normalization with assert in helper #114141

Closed

awgu marked this pull request as ready for review November 20, 2023 17:32

awgu added ciflow/trunk Trigger trunk jobs on your pull request release notes: distributed (dtensor) release notes category labels Nov 20, 2023

pytorchmergebot added the merging label Nov 21, 2023

pytorchmergebot added Merged and removed merging labels Nov 21, 2023

pytorchmergebot closed this in 3e49621 Nov 21, 2023

awgu mentioned this pull request Nov 22, 2023

[DTensor] Computed DTensorSpec hash lazily #114322

Closed

awgu mentioned this pull request Nov 22, 2023

[FSDP] Added DDP parity test for CPU training #114372

Closed

facebook-github-bot deleted the gh/awgu/455/head branch November 24, 2023 15:27

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[DTensor] Cached hash for `DTensorSpec` #113915

[DTensor] Cached hash for `DTensorSpec` #113915

Uh oh!

awgu commented Nov 17, 2023 •

edited

Loading

Uh oh!

pytorch-bot bot commented Nov 17, 2023 •

edited

Loading

Uh oh!

awgu commented Nov 17, 2023 •

edited

Loading

Uh oh!

wanchaol left a comment

Uh oh!

awgu commented Nov 21, 2023

Uh oh!

pytorchmergebot commented Nov 21, 2023

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

	class TensorMeta(NamedTuple):
	# simple named tuple to represent tensor metadata
	# intentionally to stay simple only for sharding
	# propagation purposes.
	shape: torch.Size
	stride: Tuple[int, ...]
	dtype: torch.dtype

[DTensor] Cached hash for DTensorSpec #113915

[DTensor] Cached hash for DTensorSpec #113915

Uh oh!

Conversation

awgu commented Nov 17, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Nov 17, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/113915

✅ No Failures

Uh oh!

awgu commented Nov 17, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

wanchaol left a comment

Choose a reason for hiding this comment

Uh oh!

awgu commented Nov 21, 2023

Uh oh!

pytorchmergebot commented Nov 21, 2023

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

[DTensor] Cached hash for `DTensorSpec` #113915

[DTensor] Cached hash for `DTensorSpec` #113915

awgu commented Nov 17, 2023 •

edited

Loading

pytorch-bot bot commented Nov 17, 2023 •

edited

Loading

awgu commented Nov 17, 2023 •

edited

Loading