Commit Graph
27 Commits
Author SHA1 Message Date
MarcelineVQ 0576a06ef0 fix: Zig operator precedence bugs in silicon SSE functions
All bitwise & and | comparisons were missing parentheses — Zig's ==
and != bind tighter than & and |, so `flags & 0x8 == 0` parsed as
`flags & (0x8 == 0)` = `flags & 0` = always 0.

Affected: si_frustumCullBBox (behind-camera check never ran, occlusion
flag check always passed), si_processLinkedListCollision (early-out
never triggered, node skip broken, AABB hit test only checked X-axis).

Also adds fastRecip (vrcpss+NR) and cvtss2si helpers for frustumCull
perspective divide optimization (2.0x speedup, was 1.6x).
2026-03-23 20:57:19 -07:00
MarcelineVQ 035355b757 perf: FrustumCullBoundingBox SSE replacement, 1.6x speedup
- silicon_sse: si_frustumCullBBox (0x686000) — inline V4 mat*vec3
  transforms, SSE perspective divide, 4-wide horizon buffer scan.
  Benched 1.6x (117→72 cyc/call). Installed via JMP patch (544 bytes,
  won't fit in 380-byte original).
- bench: add frustumCullBBox benchmark with mapped globals, identity
  matrices, and horizon buffer test fixture.
- Remove patch table entry for frustumCullBBox, install via detour hook
  instead (allows future A/B testing if needed).
2026-03-23 20:46:28 -07:00
MarcelineVQ 7819d6d914 perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44
- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's
  search result reused for scale/translation when tracks share temporal
  structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction.
- silicon: SSE replacement for processLinkedListCollision (0x6ABC40,
  1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW.
  Benched at 3.2x speedup (378→115 cyc/call, 8 nodes).
- transform44: remove A/B toggle, always use SSE path (teardown guard
  kept). A/B infrastructure remains for other hooks.
- bench: add processLinkedListCollision benchmark with fake linked list
  test fixture and stubbed addGeometryToBuffer.
2026-03-23 20:27:26 -07:00
MarcelineVQ 10bd922cc3 silicon: fix CC mismatches (normalizeVec3InPlace TC, packParticleColor TC, addVec3ToAccumulator remove phantom scale param, revert classifyPointFrustum/testOBBFrustum to TC); runtime JMP address resolution 2026-03-17 13:33:51 -07:00
MarcelineVQ dc37042e6a silicon: naked asm for transpose/setAlpha/addColor/normalize, remove vec3Dot/distPlane/insideBounds from patches, direct patch transpose (7cy, 3.7x) 2026-03-17 03:52:50 -07:00
MarcelineVQ 3dfa559eac silicon: direct byte patch for ftol, JMP patches for all others, removed compare mode 2026-03-17 03:20:59 -07:00
MarcelineVQ 4152d1dc78 silicon: JMP patch infrastructure + native CC for all functions, no naked asm except ftol 2026-03-17 03:12:28 -07:00
MarcelineVQ 6140a3a333 silicon_sse: native calling conventions (thiscall/fastcall/stdcall) for all functions, ready for JMP patching 2026-03-17 02:51:50 -07:00
MarcelineVQ 48e542dcf4 silicon_sse: packParticleColor 26cy->16cy (6.3x) via packed V4 round+convert; checkBoxLineIntersect branchless @min/@max cleanup; all tasks complete 2026-03-17 02:31:01 -07:00
MarcelineVQ 741d44c973 silicon_sse: FSINCOS tested (slower), normalizeVec3/translateBoundingVol/createRotMat3x4 documented at optimum 2026-03-17 02:24:56 -07:00
MarcelineVQ 350cd93162 silicon_sse: normalizeVec3InPlace + testSphereFrustum at optimum, documented attempts 2026-03-17 02:06:08 -07:00
MarcelineVQ d1c3c4654f silicon_sse: batched testOBBFrustum 79cy->43cy (2.3x), 4-corner V4 dots + branchless mask 2026-03-17 01:49:37 -07:00
MarcelineVQ 9715e32ed7 silicon_sse: mulMat3x4InPlace eliminate tmp copy + preload b (1.7x -> 1.8x) 2026-03-17 01:47:13 -07:00
MarcelineVQ 9f2fa5213a bench: best-of-5 for all silicon functions; silicon_sse: V4 column mulMat3x4/InPlace (1.8x/1.7x) 2026-03-17 01:43:16 -07:00
MarcelineVQ 662c854376 silicon_sse: patch-in-place isPointInsideBounds 6cy->5cy (1.2x), naked vucomiss 2026-03-17 01:26:55 -07:00
MarcelineVQ 16f1a073ae silicon_sse: branchless classifyPointFrustum 35cy->30cy (1.8x), revert addToColorAccum 2026-03-17 01:17:23 -07:00
MarcelineVQ cb0888e0fa silicon_sse: naked FMA asm for vec3Dot (0.4x->0.8x) and distanceToPlane (0.7x->1.0x) 2026-03-17 01:01:07 -07:00
MarcelineVQ 9488fb568c silicon_sse: V4/@mulAdd/@shuffle rewrites for all functions
Major improvements from SSE4.1+FMA+AVX target + explicit SIMD:
- transposeMat4x4: 0.9x -> 2.4x (V4 shuffle)
- rotateMatByQuat: 4.0x -> 5.0x (V4 matmul)
- testOBBFrustum: 0.9x -> 1.2x (V4 corner transform + dot4)
- classifyPointFrustum: 1.4x -> 2.0x (V4 dot4 with {x,y,z,1} trick)
- translateBoundingVol: 1.3x -> 1.7x (@mulAdd plane distances)
- testSphereFrustum: 1.1x -> 1.3x (V4 dot)
- createRotMat3x4: @mulAdd for all 9 matrix entries
- mulMat3x4/InPlace: @mulAdd chains
- quatSlerp: V4 blend + @mulAdd dot
- addVec3ToAccumulator: @mulAdd for scale multiply

All parity tests pass.
2026-03-17 00:34:12 -07:00
MarcelineVQ b27a987355 silicon: FTOL_ONLY debug flag, ftolSSE2 compare mode, h67 disabled
Temporary debug state for isolating visual issues:
- FTOL_ONLY gates all hooks except __ftol and world update reporter
- ftolSSE2 compare mode: calls original + SSE2, counts mismatches
- h67 (ConvertPixelsToScreenAlt 0x5C7010) disabled: crashes with ECX=0
2026-03-17 00:22:21 -07:00
MarcelineVQ 8537587df3 silicon: si_ftol SSE3 FISTTP replacement, bench patch-in-place framework
si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.

Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.

Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
2026-03-17 00:22:07 -07:00
MarcelineVQ c9249c6c96 bench: fix 3 correctness bugs found by benchmarker
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.

packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.

42/42 benchmarks now pass correctness. 0 MISMATCHes.
2026-03-15 12:27:59 -07:00
MarcelineVQ 7c88423a41 bench: add remaining silicon functions, total 38 benchmarks
Added 8 more silicon SSE functions to benchmark: normalizeVec3,
testOBBFrustum, calculateSinCos, translateBoundingVol, addToColorAccum,
packParticleColor, setParticleAlpha, and addVec3ToAccumulator (export fn).

Notable results from new entries:
  packParticleColor: 6.8x (109->16) MISMATCH -- alpha byte diff
  setParticleAlpha:  3.4x (24->7)
  calculateSinCos:   3.0x (138->45)
  normalizeVec3:     1.8x (18->10)
  testOBBFrustum:    1.3x (107->81)

3 MISMATCHes total (mulMat3x4, mulMat3x4InPlace, packParticleColor)
need correctness investigation -- likely matrix layout and struct
byte-order differences.
2026-03-15 12:08:59 -07:00
MarcelineVQ 55b4931fcb bench: map WoW PE sections for full-fidelity benchmarking, add silicon SSE
Major bench harness upgrade:
- Maps WoW .text (4MB) and .rdata (160KB) at original virtual addresses
  instead of individual function byte arrays. All CALL targets and float
  constants resolve automatically -- no more manual mapGameConstants().
- Bench binary linked at 0x10000000 to avoid address conflict with WoW
  PE sections at 0x400000-0xD00000.

Added silicon_sse.zig: 18 pure math functions extracted from silicon.zig
as export fn (C ABI) for standalone compilation. Covers frustum culling,
bounding volume ops, quaternion slerp, matrix multiplies, trig, etc.

Silicon benchmark results (all 14 new entries pass correctness):
  checkBoxLineIntersect: 3.2x (72->22)  -- slab AABB intersection
  rotateMatByQuat:       3.1x (172->54) -- quat->mat + mat multiply
  quatSlerp:             2.5x (454->178)
  createZRotMat3x3:      2.5x (140->55)
  createRotMat3x4:       2.2x (163->73)
  normalizeVec3InPlace:  1.8x (34->18)
  mulMat3x4InPlace:      1.6x (106->66) -- MISMATCH (layout diff, needs investigation)
  classifyPointFrustum:  1.5x (60->40)
  mulMat3x4:             1.2x (64->52)  -- MISMATCH (same layout issue)

Two MISMATCH entries on mat3x4 multiply -- likely row/column order
difference between original and our implementation. Correctness needs
verification against game behavior.
2026-03-15 12:04:02 -07:00
MarcelineVQ 28dd1cf6fa silicon: implement 32 SSE math replacements for x87 FPU functions
Replace probe-only detours with direct SSE/scalar math implementations
for the highest-impact WoW.exe functions by call frequency:

Core transforms (10.8M+ calls/7.5s):
- transformVector3ByMatrix4x4: @Vector(4,f32) dot products
- transformVector4ByMatrix4x4: 4-row vector dots
- MultiplyMatrix3x4, multiplyMatrix3x3: affine matrix multiply
- MultiplyMatrix3x4InPlace: aliasing-safe with stack temp

Frustum/collision (8.2M+ combined):
- ClassifyPointAgainstFrustum: 6-plane dot + bitmask
- CheckBoxLineIntersection: slab-method AABB test
- TestSphereAgainstFrustum: 6-plane sphere test
- TestOBBAgainstFrustum: 8-corner OBB vs 6 planes
- IsPointInsideBounds: 3-component compare
- CalculateDistanceToPlane: ray-plane intersection

Matrix builders (300K+):
- createAxisAngleRotationMatrix 4x4/3x3/3x4: Rodrigues formula
- createZRotationMatrix3x3: sin/cos rotation
- rotateMatrixByQuaternion: quat→matrix + 4x4 multiply
- getTransposedMatrix4x4: inline transpose
- scaleMatrix3x3ByVector, ApplyTranslationMatrix

Vector/scalar:
- normalizeVector3, NormalizeVector3_InPlace
- Vector3_DotProduct, quaternion_slerp
- addVector3ToAccumulator, addToColorAccumulator
- calculateSinCos

Particle:
- packParticleColorToBytes, setParticleAlphaFromFloat

Bounding volume:
- TranslateBoundingVolume, TransformBoundingVolume

84 functions remain as probe-only (orchestrators, Lua internals,
complex renderers — optimization is in their callees, now replaced).
2026-03-15 11:18:35 -07:00
MarcelineVQ 05e016da46 silicon: fix 5 param count bugs, add periodic probe reporting
Fix calling convention mismatches found via systematic RET purge audit:
- interpolateKeyframes: FC4d→FC3d (3 params, not 4)
- ConvertPixelsToScreenAlt: TC2r→TC2d (returns x87 float, not u32)
- UpdateObjectTransform: FC2r→FC3r (3 params, not 2)
- packParticleColorToBytes: FC4v→FC5v (5 params, not 4)
- setParticleAlphaFromFloat: FC2v→FC3v (3 params, not 2)

Disable GetFPUControlWord probe (corrupts FLDCW state).
Use naked asm for __ftol probe (preserves implicit x87 ST(0)).
Add OnWorldUpdate hook for periodic hit count dumps every ~7.5s.
All 115 active probes verified with non-zero counts in-game.
2026-03-15 03:41:28 -07:00
MarcelineVQ e70015579f silicon: add 116 probe hooks for call-counting all explored functions
Probe infrastructure using comptime probeDetour() that generates
per-function detours: atomic counter increment + callOriginal passthrough.
Hit counts reported on shutdown. Hooks installed at lateInit (engine init),
not DLL load time.

Key findings during decompilation:
- 0x40CF81 is GetFPUControlWord, not ftol (silicon mislabel)
- Real __ftol at 0x40A2B0 (51+ callers, hot path)
- 0x7B7A80/7B7B10 use normal stack floats, not FPU register params
- 0x686640/686820/6868E0 are bounding volume ops, not vector ops
- 0x7786A0 is a UI model constructor, not SetModelLighting
2026-03-15 01:55:09 -07:00
MarcelineVQ 66696e993a silicon: decompile and annotate all ~100 hooked WoW.exe functions
Complete catalog of libSiliconPatch240.dll hook targets with verified
calling conventions, parameter layouts, RET stack cleanup, and functional
descriptions from Ghidra analysis. Zero unknown stubs remain.
2026-03-15 01:22:43 -07:00