perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44

- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's
  search result reused for scale/translation when tracks share temporal
  structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction.
- silicon: SSE replacement for processLinkedListCollision (0x6ABC40,
  1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW.
  Benched at 3.2x speedup (378→115 cyc/call, 8 nodes).
- transform44: remove A/B toggle, always use SSE path (teardown guard
  kept). A/B infrastructure remains for other hooks.
- bench: add processLinkedListCollision benchmark with fake linked list
  test fixture and stubbed addGeometryToBuffer.
This commit is contained in:
MarcelineVQ
2026-03-23 20:27:26 -07:00
parent b0b70a744e
commit 7819d6d914
5 changed files with 283 additions and 8 deletions
+1 -3
View File
@@ -298,10 +298,8 @@ fn transformDetour(this: u32, mat1: u32, mat2: u32, mat3: u32, mat4: u32) callco
if (teardown_active) {
transform_hook.callOriginal(.{ this, mat1, mat2, mat3, mat4 });
} else if (ab_use_custom) {
transformMatrix4x4_SSE(this, mat1, mat2, mat3, mat4);
} else {
transformMatrix4x4_REF(this, mat1, mat2, mat3, mat4);
transformMatrix4x4_SSE(this, mat1, mat2, mat3, mat4);
}
t44_depth -|= 1;