perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44
- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's search result reused for scale/translation when tracks share temporal structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction. - silicon: SSE replacement for processLinkedListCollision (0x6ABC40, 1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW. Benched at 3.2x speedup (378→115 cyc/call, 8 nodes). - transform44: remove A/B toggle, always use SSE path (teardown guard kept). A/B infrastructure remains for other hooks. - bench: add processLinkedListCollision benchmark with fake linked list test fixture and stubbed addGeometryToBuffer.
This commit is contained in:
@@ -298,10 +298,8 @@ fn transformDetour(this: u32, mat1: u32, mat2: u32, mat3: u32, mat4: u32) callco
|
||||
|
||||
if (teardown_active) {
|
||||
transform_hook.callOriginal(.{ this, mat1, mat2, mat3, mat4 });
|
||||
} else if (ab_use_custom) {
|
||||
transformMatrix4x4_SSE(this, mat1, mat2, mat3, mat4);
|
||||
} else {
|
||||
transformMatrix4x4_REF(this, mat1, mat2, mat3, mat4);
|
||||
transformMatrix4x4_SSE(this, mat1, mat2, mat3, mat4);
|
||||
}
|
||||
|
||||
t44_depth -|= 1;
|
||||
|
||||
Reference in New Issue
Block a user