Replace FloatUtils Conversion to Integral types through unions with casting - #7676
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #7676 +/- ##
==========================================
- Coverage 69.88% 69.88% -0.01%
==========================================
Files 1487 1487
Lines 276240 276236 -4
Branches 28287 28287
==========================================
- Hits 193046 193043 -3
- Misses 75703 75705 +2
+ Partials 7491 7488 -3
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Thank you for the PR, the changes looks good to me. There might be other areas where we can do similar improvements (e.g.
machinelearning/src/Microsoft.ML.Core/Utilities/FloatUtils.cs
Lines 429 to 441 in 677a0b8
Thanks for approving the PR! Regarding the changes to IsFinite() and other functions that have a similar structure to them, is the idea that we can get rid of the default value write and just overwrite whatever is present in the struct? |
My idea was something along these lines But it would have to be measured to see if it actually brings improvements over the current implementation. Let me know what you think. |
|
I see. Then I think I'd like to those as part of a separate PR because it will take me some time to go through all the "similar" functions and verify how the underlying assembly changes. |
FloatUtils.cs contains an overloaded function GetBits that takes Single/Double precision floating point values and converts them into the corresponding unsigned integral types with the same width. The disassembly for this function will have a store-load dependency which can be reduced to a single mov instruction.
The function utilizes the internal union that is stored in the class, by writing into the floating point field, and then reading from the integral field. This can be observed in the Microsoft.ML.PerformanceTests.HashBench.HashScalarDouble testcase, when inspecting the HashRound function in Hashing.cs:
When value is converted from a double to a ulong, the corresponding assembly code will look like this because of union semantics:
This can be collapsed into a single instruction like so:
When benchmarked on both Microsoft.ML.PerformanceTests.HashBench.HashScalarDouble and Microsoft.ML.PerformanceTests.HashBench.HashScalarFloat, it shows a performance improvement.
Before:
After: