Add SSE multiplyMatrix4x4 hook (0x7bc6a0) with A/B benchmark — 4.3x speedup
Standalone SSE 4x4 matrix multiply replacing 542 bytes of x87 FPU. Refactored rotateMatrixByAxisAngle to call the new function (with temp buffer to avoid aliasing). All 5 SSE replacements now confirmed winners: clip (4x), triplane (10x), rotmat (4.3x), raytri (1.3x), matmul (4.3x).
This commit is contained in:
+1027
-17
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user