Add SSE multiplyMatrix4x4 hook (0x7bc6a0) with A/B benchmark — 4.3x speedup

Standalone SSE 4x4 matrix multiply replacing 542 bytes of x87 FPU.
Refactored rotateMatrixByAxisAngle to call the new function (with temp
buffer to avoid aliasing). All 5 SSE replacements now confirmed winners:
clip (4x), triplane (10x), rotmat (4.3x), raytri (1.3x), matmul (4.3x).
This commit is contained in:
MarcelineVQ
2026-03-12 18:44:02 -07:00
parent 53fb100368
commit fc7f6d637b
2 changed files with 1403 additions and 17 deletions
File diff suppressed because it is too large Load Diff