🚀 Motivation and context
It would be great to extend the compatibility of HTA to AMD Devices. Some basic features of HTA (Temporal/Kernel Breakdown) work with Kineto traces from AMD GPUs. However, more advanced modules such as the critical path analyzer do not work.
Moreover, other tools that use HTA/parts of HTA ( ex. Chakra - Open source Execution Trace Schema) run into errors on AMD Kineto Traces, caused by HTA's lack of support (see this PR on the Chakra Repo).
Description
-
In hta/common/trace_symbol_table.py, the symbol table recognizes certain trace events as runtime launch events from their names. A list of known cuda launch op names is checked and if the event name matches, it gets assigned as a launch event. However, AMD Kineto traces do not have these launch event names. The respective known runtime launch names must be added for AMD (Recently, a similar change was made for MTIA).
-
In hta/common/trace_filter.py, GPU ops are filtered using the "stream" value from the Kineto trace. This differentiates them from kernel launch ops, etc. Currently, HTA checks for a stream value of greater than 0. This is done because cuda assigns a default non-zero stream value (in most cases, a value of 7). However, ROCm assigns a default stream value of 0. A simple change can be made that changes the > sign to a >= sign so that AMD GPU events are not incorrectly filtered out.
I have already made the changes locally and run all the unit tests, which pass successfully. Let me know if you have any suggestions; if not, I could open a PR. Thank you!
Alternatives
No response
Additional context
No response
🚀 Motivation and context
It would be great to extend the compatibility of HTA to AMD Devices. Some basic features of HTA (Temporal/Kernel Breakdown) work with Kineto traces from AMD GPUs. However, more advanced modules such as the critical path analyzer do not work.
Moreover, other tools that use HTA/parts of HTA ( ex. Chakra - Open source Execution Trace Schema) run into errors on AMD Kineto Traces, caused by HTA's lack of support (see this PR on the Chakra Repo).
Description
In hta/common/trace_symbol_table.py, the symbol table recognizes certain trace events as runtime launch events from their names. A list of known cuda launch op names is checked and if the event name matches, it gets assigned as a launch event. However, AMD Kineto traces do not have these launch event names. The respective known runtime launch names must be added for AMD (Recently, a similar change was made for MTIA).
In hta/common/trace_filter.py, GPU ops are filtered using the "stream" value from the Kineto trace. This differentiates them from kernel launch ops, etc. Currently, HTA checks for a stream value of greater than 0. This is done because cuda assigns a default non-zero stream value (in most cases, a value of 7). However, ROCm assigns a default stream value of 0. A simple change can be made that changes the > sign to a >= sign so that AMD GPU events are not incorrectly filtered out.
I have already made the changes locally and run all the unit tests, which pass successfully. Let me know if you have any suggestions; if not, I could open a PR. Thank you!
Alternatives
No response
Additional context
No response