Commit Graph
213 Commits
Author SHA1 Message Date
MarcelineVQ 10bd922cc3 silicon: fix CC mismatches (normalizeVec3InPlace TC, packParticleColor TC, addVec3ToAccumulator remove phantom scale param, revert classifyPointFrustum/testOBBFrustum to TC); runtime JMP address resolution 2026-03-17 13:33:51 -07:00
MarcelineVQ dc37042e6a silicon: naked asm for transpose/setAlpha/addColor/normalize, remove vec3Dot/distPlane/insideBounds from patches, direct patch transpose (7cy, 3.7x) 2026-03-17 03:52:50 -07:00
MarcelineVQ 3dfa559eac silicon: direct byte patch for ftol, JMP patches for all others, removed compare mode 2026-03-17 03:20:59 -07:00
MarcelineVQ 4152d1dc78 silicon: JMP patch infrastructure + native CC for all functions, no naked asm except ftol 2026-03-17 03:12:28 -07:00
MarcelineVQ 6140a3a333 silicon_sse: native calling conventions (thiscall/fastcall/stdcall) for all functions, ready for JMP patching 2026-03-17 02:51:50 -07:00
MarcelineVQ 48e542dcf4 silicon_sse: packParticleColor 26cy->16cy (6.3x) via packed V4 round+convert; checkBoxLineIntersect branchless @min/@max cleanup; all tasks complete 2026-03-17 02:31:01 -07:00
MarcelineVQ 741d44c973 silicon_sse: FSINCOS tested (slower), normalizeVec3/translateBoundingVol/createRotMat3x4 documented at optimum 2026-03-17 02:24:56 -07:00
MarcelineVQ 350cd93162 silicon_sse: normalizeVec3InPlace + testSphereFrustum at optimum, documented attempts 2026-03-17 02:06:08 -07:00
MarcelineVQ d1c3c4654f silicon_sse: batched testOBBFrustum 79cy->43cy (2.3x), 4-corner V4 dots + branchless mask 2026-03-17 01:49:37 -07:00
MarcelineVQ 9715e32ed7 silicon_sse: mulMat3x4InPlace eliminate tmp copy + preload b (1.7x -> 1.8x) 2026-03-17 01:47:13 -07:00
MarcelineVQ 9f2fa5213a bench: best-of-5 for all silicon functions; silicon_sse: V4 column mulMat3x4/InPlace (1.8x/1.7x) 2026-03-17 01:43:16 -07:00
MarcelineVQ 662c854376 silicon_sse: patch-in-place isPointInsideBounds 6cy->5cy (1.2x), naked vucomiss 2026-03-17 01:26:55 -07:00
MarcelineVQ 16f1a073ae silicon_sse: branchless classifyPointFrustum 35cy->30cy (1.8x), revert addToColorAccum 2026-03-17 01:17:23 -07:00
MarcelineVQ cb0888e0fa silicon_sse: naked FMA asm for vec3Dot (0.4x->0.8x) and distanceToPlane (0.7x->1.0x) 2026-03-17 01:01:07 -07:00
MarcelineVQ 9488fb568c silicon_sse: V4/@mulAdd/@shuffle rewrites for all functions
Major improvements from SSE4.1+FMA+AVX target + explicit SIMD:
- transposeMat4x4: 0.9x -> 2.4x (V4 shuffle)
- rotateMatByQuat: 4.0x -> 5.0x (V4 matmul)
- testOBBFrustum: 0.9x -> 1.2x (V4 corner transform + dot4)
- classifyPointFrustum: 1.4x -> 2.0x (V4 dot4 with {x,y,z,1} trick)
- translateBoundingVol: 1.3x -> 1.7x (@mulAdd plane distances)
- testSphereFrustum: 1.1x -> 1.3x (V4 dot)
- createRotMat3x4: @mulAdd for all 9 matrix entries
- mulMat3x4/InPlace: @mulAdd chains
- quatSlerp: V4 blend + @mulAdd dot
- addVec3ToAccumulator: @mulAdd for scale multiply

All parity tests pass.
2026-03-17 00:34:12 -07:00
MarcelineVQ bad3126973 build: enable SSE4.1+FMA+AVX for silicon_sse compilation unit
Was compiling with baseline SSE2 only. Now matches bone_sse target.
Free wins: packParticleColor 1.2x->4.6x, rotateMatByQuat 3.4x->4.0x,
normalizeVec3 1.1x->1.8x, mulMat3x4InPlace 1.4x->1.7x.
2026-03-17 00:27:48 -07:00
MarcelineVQ b27a987355 silicon: FTOL_ONLY debug flag, ftolSSE2 compare mode, h67 disabled
Temporary debug state for isolating visual issues:
- FTOL_ONLY gates all hooks except __ftol and world update reporter
- ftolSSE2 compare mode: calls original + SSE2, counts mismatches
- h67 (ConvertPixelsToScreenAlt 0x5C7010) disabled: crashes with ECX=0
2026-03-17 00:22:21 -07:00
MarcelineVQ 8537587df3 silicon: si_ftol SSE3 FISTTP replacement, bench patch-in-place framework
si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.

Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.

Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
2026-03-17 00:22:07 -07:00
MarcelineVQ 61ee4f48e1 bone_sse: f32 callFtol, fastMod conditional subtract for looping anims
callFtol: use f32 multiply instead of f64 intermediate. Parity holds --
the delta*scale product is well within f32 precision range.

fastMod: replace integer modulo (idiv, ~25 cycles) with conditional
subtract (~2 cycles) for looping animation frame computation. Falls
back to real modulo for large time skips (alt-tab, etc).

3609 cycles (-14% vs 4176 baseline), parity PASS.
2026-03-16 17:06:20 -07:00
MarcelineVQ dd43f6aca3 bone_sse: return interp values in registers, hoist runtime constants, value-based local matrix
- interpAnimKF returns [4]f32, interpVec3Track returns [3]f32, interpFloatTrack returns f32
- Internal crossfade blends stay in registers instead of writing then re-reading from memory
- Bone loop uses returned values directly for buildRotationMatrix/scaleMatrix3x3/translation
- Hoist getShortToFloat() reads to function entry in section functions (Proposal D)
- Remove dead blendVec3, getHermite5, HERMITE_3
- buildRotationMatrixVal returns [16]f32; bone loop uses array ops instead of u32 pointer casts
- matMul4x4Local/matMul4x4InPlace variants for local array operands
- shortInterpToFloat takes pre-read stf parameter

3574 cycles (-14% vs 4176 baseline), parity PASS (9120 bytes).
2026-03-16 16:25:26 -07:00
MarcelineVQ d9d2e41eb4 bone_sse: return InterpResult in registers, eliminate store-forward latency
findInterpIdx now returns {idx0, idx1, t} as a struct instead of writing
all three to the output buffer. Only output[0] is written for next-frame
cache persistence. All 29 call sites updated to use returned values.

3574 cycles (-14% vs 4176 baseline), was 3841 (-8%). Parity PASS.
2026-03-16 15:37:28 -07:00
MarcelineVQ 0419f39833 bench: 2M iterations, baseline 4176 cycles, SSE 3841 (-8%), parity PASS 2026-03-16 15:25:00 -07:00
MarcelineVQ 99b4c8b2d0 bone_sse: f32 findInterpIdx t division, non-inline sections with quota — 3697 cycles (-6%) 2026-03-16 14:12:53 -07:00
MarcelineVQ 502fd9ba04 bone_sse: remove align(1) for naturally-aligned game data — 3641 cycles (was 3744) 2026-03-16 12:43:51 -07:00
MarcelineVQ 88fec305be bone_sse: aggressive inlining + skip identity init + fused rotateByQuaternion — 3744 cycles (was 4040) 2026-03-16 12:41:54 -07:00
MarcelineVQ 22f9e712f4 bench: 100% code path coverage — multi-track ranges, FloatTrack12 mode=0, 9120 bytes parity PASS 2026-03-16 12:34:07 -07:00
MarcelineVQ 4776a979c9 bench: 94% code path coverage, 8988 bytes parity check — all interp modes, billboards, particles, crossfade, attachments 2026-03-16 12:29:31 -07:00
MarcelineVQ 5baa49d97a bench: fix frame_ctr=0 so section functions execute — 3249 cycles, parity PASS 2026-03-16 12:20:22 -07:00
MarcelineVQ 9c215a0bb1 bench: full parity check across all 13 output buffers (8136 bytes), dual BASELINE/SSE 2026-03-16 12:12:57 -07:00
MarcelineVQ 4ed5495219 bench: dual BASELINE/SSE with parity check — 2177 vs 2112 cycles, PASS 2026-03-16 12:10:51 -07:00
MarcelineVQ 4198a00c26 bench: full path coverage + determinism validation — 2065 cycles/call, PASS 2026-03-16 11:57:37 -07:00
MarcelineVQ 75bb38f5af bench: full coverage fixture — 1012 cycles/call (ribbon, particle, attach, billboard, crossfade, clamped, GS, time delta) 2026-03-16 11:50:38 -07:00
MarcelineVQ b6ab4ac59f bench: comprehensive fixture exercising all code paths — 636 cycles/call baseline 2026-03-16 11:46:16 -07:00
MarcelineVQ f119c15b37 bench: comprehensive transform44 fixture — 578 cycles/call baseline (12 bones + texAnim + colorAnim + wordAnim + boneKF) 2026-03-16 11:42:31 -07:00
MarcelineVQ 93100a1a7e bench: add transform44 SSE benchmark — 288 cycles/call baseline (8 bones, 4 animated) 2026-03-16 11:30:06 -07:00
MarcelineVQ 91444f8104 bone_sse: incremental bone pointer advancement, V4 copyMat4 2026-03-16 11:18:50 -07:00
MarcelineVQ 3a522fd1e4 bone_sse: cache frame_ctr, pass to all section functions — eliminates ~29 redundant reads 2026-03-16 11:14:56 -07:00
MarcelineVQ c80e63522b bone_sse: @mulAdd (FMA) for all lerp/blend paths across interp and section functions 2026-03-16 11:06:59 -07:00
MarcelineVQ c36352f4ff bone_sse: @mulAdd (FMA) for lerpVec3, applyTranslation, vec3SqMag 2026-03-16 11:00:01 -07:00
MarcelineVQ a7e2f978e2 bone_sse: simplify rotateByQuaternion to use matMul4x4 V4+FMA path 2026-03-16 10:56:48 -07:00
MarcelineVQ 5d41cd03a3 bone_sse: V4 FMA matmul, enable SSE4.1+FMA+AVX target 2026-03-16 10:54:25 -07:00
MarcelineVQ 08d1992eaf bone_sse: replace IsParticleBufferEmpty with pure Zig — 1 game call remains (atexit) 2026-03-16 10:45:06 -07:00
MarcelineVQ 1987dc5622 bone_sse: pure Zig — only 2 game calls remain (atexit init + particle buffer check)
Replaced all remaining game function calls:
- findInterpIdx (0x713D50): full temporal-coherence search reimplementation
- interpAnimKF in boneKeyframeLoop: reuses existing pure Zig version
- applyTranslation/rotateByQuaternion/scaleMatrix3x3 in boneKeyframeLoop
- getInterpolatedFloat (0x71AF20): replaced with interpFloatTrack (identical)
- extractByte (0x71AE90): findInterpIdx + direct byte read
- Child recursion: direct call to transformImpl_SSE instead of 0x714260 hook

Only 2 game calls remain (cannot be replaced):
- 0x409AEF: one-time atexit registration in boneKeyframeLoop
- 0x7B5F60: IsParticleBufferEmpty (reads game particle state)
2026-03-16 10:38:34 -07:00
MarcelineVQ 16c913b9b7 bone_sse: pure Zig findInterpIdx — last major game call in interpolation path 2026-03-16 10:33:40 -07:00
MarcelineVQ 800adbe185 bone_sse: pure Zig bone loop — interpAnimKF, buildRot, scaleMat, matMul all replaced 2026-03-16 10:27:33 -07:00
MarcelineVQ 0aab0d3678 bone_sse: replace getIndexOffset/setShortValue with direct ri16 reads 2026-03-16 10:22:06 -07:00
MarcelineVQ 8736e0b3da bone_sse: clean A/B testing, remove diagnostic comparison code
Remove the double-call REF/SSE bone output comparison diagnostic.
Clean detour: REF baseline, SSE custom, simple toggle.
2026-03-16 10:18:35 -07:00
MarcelineVQ 008db88403 bone_sse: replace matMul/ftol/vec3SqMag with pure Zig, fix visual artifacts
Root cause of billboard visual artifacts: callVec3SqMag used inline asm
to call game's x87 vec3SqMag (0x4549F0) with fstps to capture ST0.
With SSE2 codegen, the x87/SSE state interaction caused corrupted float
values in billboard bone matrices (mat[0][2] wildly wrong).

Replaced with pure Zig: x*x + y*y + z*z — no x87, no inline asm.

Also replaced:
- matMul (0x74A7C0): pure Zig f64 scalar matmul, no alignment needs
- callFtol (0x40A2B0): f64 intermediate + @intFromFloat (cvttsd2si)

Architecture: bone_sse.zig is REF code compiled with SSE2, called as
cdecl from thiscall wrapper in transform44.zig (cross-object to prevent
LLVM inlining AND ESP alignment into thiscall frame).

A/B: other hooks gated behind AB_OTHER_HOOKS=false for isolated testing.
2026-03-16 10:10:02 -07:00
MarcelineVQ 252853dd8f bone_sse: fix findInterpIdx from assembly, runtime constants, interpAnimKF stride
Assembly-verified fixes from t44_helpers_asm.txt:

findInterpIdx (0x713D50):
- Range format is [start, last] not [start, count] (DEC EDI pattern)
- Backward scan entry: delta >= 0xFFFFFE0C not > (JC = unsigned below)
- 500-tick threshold (0x1F4) for forward/backward vs binary search
- GS check is CMP AX,0xFFFF (word compare), not >= 0
- t computation: FILD qword (i64 numer) / FIDIV dword (i32 denom)

interpAnimKF (0x713EA0):
- Keyframe stride is 16 bytes (SHL EAX,0x4), NOT 8 (CompQuat)
- Values are raw floats, no short-to-float conversion needed

Runtime constants — all now read from game memory:
- 0x80297C (3.0) and 0x802990 (6.0) for Hermite/Bezier basis
- 0x80C5C8 for billboard squared magnitude threshold
- 0x811610 and 0x8029D4 already read at runtime

Wrapper pattern: thiscall export delegates to normal fn for AVX alignment.
2026-03-15 18:39:27 -07:00
MarcelineVQ 3a031803ef bone_sse: pure Zig SSE/FMA reimplementation, zero game function calls
Replace bone_sse.zig with a complete pure Zig implementation compiled
with SSE4.1 + FMA + AVX. All 18 game function calls replaced:

- findInterpIdx (0x713D50): temporal-coherence keyframe search
- interpAnimKF (0x713EA0): CompQuat lerp for rotation keyframes
- extractByte (0x71AE90): byte keyframe extraction
- getInterpolatedFloat (0x71AF20): float track with direct blend read
- callFtol (0x40A2B0): @intFromFloat replaces x87 __ftol
- callVec3SqMag (0x4549F0): inline FMA dot product
- callGetIndexOffset/callSetShortValue (0x71AFF0/0x71B010): direct ri16
- matMul (0x74A7C0): V4 FMA matmul (broadcast + 3 @mulAdd per row)
- buildRotFn (0x74B6B5): inline quat→matrix
- rotateQuat (0x7BDDB0): quat→matrix then FMA matmul
- scaleMat (0x7BDCA0): inline scale from vec3 ptr
- applyTrans (0x7BDC40): inline FMA dot product translation

Only 2 game calls remain:
- 0x409AEF: one-time atexit init (boneKeyframeLoop)
- 0x7B5F60: IsParticleBufferEmpty (reads game particle state)

Child recursion calls transformMatrix4x4_SSE directly instead of
going through the hook at 0x714260.

Detour cleaned up: REF is baseline, SSE activates via ab_use_custom
toggle. Diagnostic/bisect/FPU-comparison scaffolding removed.
build.zig: bone_sse gets dedicated target with sse4_1+fma+avx features.
2026-03-15 18:16:07 -07:00