si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.
Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.
Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.
packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.
42/42 benchmarks now pass correctness. 0 MISMATCHes.
Probe infrastructure using comptime probeDetour() that generates
per-function detours: atomic counter increment + callOriginal passthrough.
Hit counts reported on shutdown. Hooks installed at lateInit (engine init),
not DLL load time.
Key findings during decompilation:
- 0x40CF81 is GetFPUControlWord, not ftol (silicon mislabel)
- Real __ftol at 0x40A2B0 (51+ callers, hot path)
- 0x7B7A80/7B7B10 use normal stack floats, not FPU register params
- 0x686640/686820/6868E0 are bounding volume ops, not vector ops
- 0x7786A0 is a UI model constructor, not SetModelLighting
Complete catalog of libSiliconPatch240.dll hook targets with verified
calling conventions, parameter layouts, RET stack cleanup, and functional
descriptions from Ghidra analysis. Zero unknown stubs remain.