si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.
Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.
Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
callFtol: use f32 multiply instead of f64 intermediate. Parity holds --
the delta*scale product is well within f32 precision range.
fastMod: replace integer modulo (idiv, ~25 cycles) with conditional
subtract (~2 cycles) for looping animation frame computation. Falls
back to real modulo for large time skips (alt-tab, etc).
3609 cycles (-14% vs 4176 baseline), parity PASS.
findInterpIdx now returns {idx0, idx1, t} as a struct instead of writing
all three to the output buffer. Only output[0] is written for next-frame
cache persistence. All 29 call sites updated to use returned values.
3574 cycles (-14% vs 4176 baseline), was 3841 (-8%). Parity PASS.
Replaced all remaining game function calls:
- findInterpIdx (0x713D50): full temporal-coherence search reimplementation
- interpAnimKF in boneKeyframeLoop: reuses existing pure Zig version
- applyTranslation/rotateByQuaternion/scaleMatrix3x3 in boneKeyframeLoop
- getInterpolatedFloat (0x71AF20): replaced with interpFloatTrack (identical)
- extractByte (0x71AE90): findInterpIdx + direct byte read
- Child recursion: direct call to transformImpl_SSE instead of 0x714260 hook
Only 2 game calls remain (cannot be replaced):
- 0x409AEF: one-time atexit registration in boneKeyframeLoop
- 0x7B5F60: IsParticleBufferEmpty (reads game particle state)
Root cause of billboard visual artifacts: callVec3SqMag used inline asm
to call game's x87 vec3SqMag (0x4549F0) with fstps to capture ST0.
With SSE2 codegen, the x87/SSE state interaction caused corrupted float
values in billboard bone matrices (mat[0][2] wildly wrong).
Replaced with pure Zig: x*x + y*y + z*z — no x87, no inline asm.
Also replaced:
- matMul (0x74A7C0): pure Zig f64 scalar matmul, no alignment needs
- callFtol (0x40A2B0): f64 intermediate + @intFromFloat (cvttsd2si)
Architecture: bone_sse.zig is REF code compiled with SSE2, called as
cdecl from thiscall wrapper in transform44.zig (cross-object to prevent
LLVM inlining AND ESP alignment into thiscall frame).
A/B: other hooks gated behind AB_OTHER_HOOKS=false for isolated testing.
Assembly-verified fixes from t44_helpers_asm.txt:
findInterpIdx (0x713D50):
- Range format is [start, last] not [start, count] (DEC EDI pattern)
- Backward scan entry: delta >= 0xFFFFFE0C not > (JC = unsigned below)
- 500-tick threshold (0x1F4) for forward/backward vs binary search
- GS check is CMP AX,0xFFFF (word compare), not >= 0
- t computation: FILD qword (i64 numer) / FIDIV dword (i32 denom)
interpAnimKF (0x713EA0):
- Keyframe stride is 16 bytes (SHL EAX,0x4), NOT 8 (CompQuat)
- Values are raw floats, no short-to-float conversion needed
Runtime constants — all now read from game memory:
- 0x80297C (3.0) and 0x802990 (6.0) for Hermite/Bezier basis
- 0x80C5C8 for billboard squared magnitude threshold
- 0x811610 and 0x8029D4 already read at runtime
Wrapper pattern: thiscall export delegates to normal fn for AVX alignment.
Full 5317-instruction walkthrough of t44_full_asm.txt vs bone_sse_reference.zig.
Crash fix (Issues 1-2): When anim_start >= anim_end in the looping animation
path, assembly always writes prim_time/sec_time = anim_start as fallback.
REF skipped the write, leaving garbage in bone_rt timing fields. On newly
loaded world SceneObjects this propagated through findInterpIdx → extractByte
→ ACCESS_VIOLATION at 0x71AEBC with ECX=0x7FFFFFFF (self-reinforcing bad
cached index).
Time clamp fix (Issues 3-4): Clamped-not-passed animation path now clamps
cur_time to sec_start when sec_start > cur_time, matching assembly at
0x7145EB/0x71474B.
Crossfade fix (Issues 5-6): interpVec3Track36 and interpFloatTrack12 had
'else return' for unknown interp modes. Assembly's JNZ skips primary interp
but falls through to crossfade check. Changed to 'else {}' fallthrough.
Also includes prior uncommitted fixes: particle crossfade blend_weight
restoration, bw>0→bw!=0, attach_count==0 early-return removal.
- Particle interpVec3Track/interpFloatTrack: pass 0.0 blend_weight to
disable crossfade, matching original which has no crossfade in particle
sections (only bone loop and boneKeyframeLoop have crossfade).
- interpFloatTrack: add explicit blend_weight parameter instead of
reading from bone_rt internally, allowing callers to control crossfade.
- Revert extractByte guard (was added then removed during investigation).
Known crash: extractByte (0x71AE90) crashes at 0x71AEBC with idx=0x7FFFFFFF
on world entry. findInterpIdx reads output[0] as cached search position;
if hierarchy buffer contains stale 0x7FFFFFFF, search overflows and
self-reinforces. Investigation ongoing — REF's attachment section matches
original assembly instruction-for-instruction.
- texAnimLoop alpha: add crossfade blend for mode != 0 (mode 0 skips
crossfade per original assembly JMP at 0x715B49). Fix alpha output
base to output+0x30 matching original ESI.
- colorAnimLoop: add crossfade blend, same mode 0 skip pattern.
Mode 0 uses direct short->float copy matching original.
- New wordAnimLoop: implements model_hdr+0x6C/0x70 word animation
section (assembly 0x715E46-0x715F25). Word copy with crossfade,
no float blending. Data stride 0x1C, output stride 0x20.
- Fix billboard cross product z-component for types 0x10/0x20:
was +cross.z, should be -cross.z (r0y*r1x - r0x*r1y).
- Extract shortInterpToFloat helper shared by alpha/color crossfade.
Known: particle emitter crash (pre-existing, idx=0x7FFFFFFF in
secondary findInterpIdx) — exposed by corrected colorAnimLoop count.
Assembly-level comparison of compiled REF against original 0x714260 revealed:
1. Billboard cross product sign error (types 0x10/0x20): computed +cross
instead of -cross for components 0/1, corrupting billboard bone matrices
2. colorAnimLoop wrong count field: read model_hdr+0x6C instead of +0x64
3. colorAnimLoop wrong gate offset: checked anim_data+0x04 instead of +0x0C
4. Timestamp delta guard inverted: REF guarded on cur_ts!=0 and always
wrote to this+0x4C; original guards on this+0x4C!=0 first and never
seeds the field (something else initializes it)
5. Section 5 emitter_ctx cached instead of re-read after matMul call
Also: build REF with x87-only target (subtract SSE/SSE2 features) to
match original's FLD/FMUL/FSTP codegen, and use callVec3SqMag for all
magnitude computations instead of inline SSE math.
Remaining known issues (not yet fixed):
- texAnimLoop alpha track missing crossfade blend
- colorAnimLoop missing crossfade blend
- Missing word animation section (model_hdr+0x6C/0x70)
- Bisect infrastructure and diagnostic code still present (test scaffolding)
- Replace all reimplemented game functions with actual game calls:
vec3_sqmag (0x4549F0), __ftol (0x40A2B0), getIndexOffset (0x71AFF0),
setShortValue (0x71B010) — matching assembly exactly
- Fix 3 wrong hardcoded constants that differ at runtime from Ghidra static values:
SHORT_TO_FLOAT: 0x38000000→0x38000100 (1/32767 not 1/32768)
BILLBOARD_EPSILON: 0x3727c5ac→0x34800000
HERMITE_5: 5.0→6.0
All now read from game memory at runtime
- Fix timestamp delta guard (this+0x4C): was guarding on stored value,
assembly guards on anim_ctx pointer — prevents first-frame initialization
- Change REF calling convention to thiscall matching original
- Add comprehensive memory comparison diagnostic (original vs REF)
- Disable interpKfDetour hook (was pure passthrough)
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.
packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.
42/42 benchmarks now pass correctness. 0 MISMATCHes.
Extracted via Ghidra from WoW.exe for the 22 silicon SSE functions that
don't overlap with ssemaths addresses. These cover frustum culling,
bounding volume transforms, quaternion slerp, matrix operations, and
various geometry functions. Ready for benchmarking.