MarcelineVQ
0576a06ef0
fix: Zig operator precedence bugs in silicon SSE functions
...
All bitwise & and | comparisons were missing parentheses — Zig's ==
and != bind tighter than & and |, so `flags & 0x8 == 0` parsed as
`flags & (0x8 == 0)` = `flags & 0` = always 0.
Affected: si_frustumCullBBox (behind-camera check never ran, occlusion
flag check always passed), si_processLinkedListCollision (early-out
never triggered, node skip broken, AABB hit test only checked X-axis).
Also adds fastRecip (vrcpss+NR) and cvtss2si helpers for frustumCull
perspective divide optimization (2.0x speedup, was 1.6x).
2026-03-23 20:57:19 -07:00
MarcelineVQ
035355b757
perf: FrustumCullBoundingBox SSE replacement, 1.6x speedup
...
- silicon_sse: si_frustumCullBBox (0x686000) — inline V4 mat*vec3
transforms, SSE perspective divide, 4-wide horizon buffer scan.
Benched 1.6x (117→72 cyc/call). Installed via JMP patch (544 bytes,
won't fit in 380-byte original).
- bench: add frustumCullBBox benchmark with mapped globals, identity
matrices, and horizon buffer test fixture.
- Remove patch table entry for frustumCullBBox, install via detour hook
instead (allows future A/B testing if needed).
2026-03-23 20:46:28 -07:00
MarcelineVQ
7819d6d914
perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44
...
- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's
search result reused for scale/translation when tracks share temporal
structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction.
- silicon: SSE replacement for processLinkedListCollision (0x6ABC40,
1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW.
Benched at 3.2x speedup (378→115 cyc/call, 8 nodes).
- transform44: remove A/B toggle, always use SSE path (teardown guard
kept). A/B infrastructure remains for other hooks.
- bench: add processLinkedListCollision benchmark with fake linked list
test fixture and stubbed addGeometryToBuffer.
2026-03-23 20:27:26 -07:00
MarcelineVQ
10bd922cc3
silicon: fix CC mismatches (normalizeVec3InPlace TC, packParticleColor TC, addVec3ToAccumulator remove phantom scale param, revert classifyPointFrustum/testOBBFrustum to TC); runtime JMP address resolution
2026-03-17 13:33:51 -07:00
MarcelineVQ
dc37042e6a
silicon: naked asm for transpose/setAlpha/addColor/normalize, remove vec3Dot/distPlane/insideBounds from patches, direct patch transpose (7cy, 3.7x)
2026-03-17 03:52:50 -07:00
MarcelineVQ
3dfa559eac
silicon: direct byte patch for ftol, JMP patches for all others, removed compare mode
2026-03-17 03:20:59 -07:00
MarcelineVQ
4152d1dc78
silicon: JMP patch infrastructure + native CC for all functions, no naked asm except ftol
2026-03-17 03:12:28 -07:00
MarcelineVQ
6140a3a333
silicon_sse: native calling conventions (thiscall/fastcall/stdcall) for all functions, ready for JMP patching
2026-03-17 02:51:50 -07:00
MarcelineVQ
48e542dcf4
silicon_sse: packParticleColor 26cy->16cy (6.3x) via packed V4 round+convert; checkBoxLineIntersect branchless @min/@max cleanup; all tasks complete
2026-03-17 02:31:01 -07:00
MarcelineVQ
741d44c973
silicon_sse: FSINCOS tested (slower), normalizeVec3/translateBoundingVol/createRotMat3x4 documented at optimum
2026-03-17 02:24:56 -07:00
MarcelineVQ
350cd93162
silicon_sse: normalizeVec3InPlace + testSphereFrustum at optimum, documented attempts
2026-03-17 02:06:08 -07:00
MarcelineVQ
d1c3c4654f
silicon_sse: batched testOBBFrustum 79cy->43cy (2.3x), 4-corner V4 dots + branchless mask
2026-03-17 01:49:37 -07:00
MarcelineVQ
9715e32ed7
silicon_sse: mulMat3x4InPlace eliminate tmp copy + preload b (1.7x -> 1.8x)
2026-03-17 01:47:13 -07:00
MarcelineVQ
9f2fa5213a
bench: best-of-5 for all silicon functions; silicon_sse: V4 column mulMat3x4/InPlace (1.8x/1.7x)
2026-03-17 01:43:16 -07:00
MarcelineVQ
662c854376
silicon_sse: patch-in-place isPointInsideBounds 6cy->5cy (1.2x), naked vucomiss
2026-03-17 01:26:55 -07:00
MarcelineVQ
16f1a073ae
silicon_sse: branchless classifyPointFrustum 35cy->30cy (1.8x), revert addToColorAccum
2026-03-17 01:17:23 -07:00
MarcelineVQ
cb0888e0fa
silicon_sse: naked FMA asm for vec3Dot (0.4x->0.8x) and distanceToPlane (0.7x->1.0x)
2026-03-17 01:01:07 -07:00
MarcelineVQ
9488fb568c
silicon_sse: V4/@mulAdd/@shuffle rewrites for all functions
...
Major improvements from SSE4.1+FMA+AVX target + explicit SIMD:
- transposeMat4x4: 0.9x -> 2.4x (V4 shuffle)
- rotateMatByQuat: 4.0x -> 5.0x (V4 matmul)
- testOBBFrustum: 0.9x -> 1.2x (V4 corner transform + dot4)
- classifyPointFrustum: 1.4x -> 2.0x (V4 dot4 with {x,y,z,1} trick)
- translateBoundingVol: 1.3x -> 1.7x (@mulAdd plane distances)
- testSphereFrustum: 1.1x -> 1.3x (V4 dot)
- createRotMat3x4: @mulAdd for all 9 matrix entries
- mulMat3x4/InPlace: @mulAdd chains
- quatSlerp: V4 blend + @mulAdd dot
- addVec3ToAccumulator: @mulAdd for scale multiply
All parity tests pass.
2026-03-17 00:34:12 -07:00
MarcelineVQ
b27a987355
silicon: FTOL_ONLY debug flag, ftolSSE2 compare mode, h67 disabled
...
Temporary debug state for isolating visual issues:
- FTOL_ONLY gates all hooks except __ftol and world update reporter
- ftolSSE2 compare mode: calls original + SSE2, counts mismatches
- h67 (ConvertPixelsToScreenAlt 0x5C7010) disabled: crashes with ECX=0
2026-03-17 00:22:21 -07:00
MarcelineVQ
8537587df3
silicon: si_ftol SSE3 FISTTP replacement, bench patch-in-place framework
...
si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.
Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.
Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
2026-03-17 00:22:07 -07:00
MarcelineVQ
c9249c6c96
bench: fix 3 correctness bugs found by benchmarker
...
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.
packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.
42/42 benchmarks now pass correctness. 0 MISMATCHes.
2026-03-15 12:27:59 -07:00
MarcelineVQ
7c88423a41
bench: add remaining silicon functions, total 38 benchmarks
...
Added 8 more silicon SSE functions to benchmark: normalizeVec3,
testOBBFrustum, calculateSinCos, translateBoundingVol, addToColorAccum,
packParticleColor, setParticleAlpha, and addVec3ToAccumulator (export fn).
Notable results from new entries:
packParticleColor: 6.8x (109->16) MISMATCH -- alpha byte diff
setParticleAlpha: 3.4x (24->7)
calculateSinCos: 3.0x (138->45)
normalizeVec3: 1.8x (18->10)
testOBBFrustum: 1.3x (107->81)
3 MISMATCHes total (mulMat3x4, mulMat3x4InPlace, packParticleColor)
need correctness investigation -- likely matrix layout and struct
byte-order differences.
2026-03-15 12:08:59 -07:00
MarcelineVQ
55b4931fcb
bench: map WoW PE sections for full-fidelity benchmarking, add silicon SSE
...
Major bench harness upgrade:
- Maps WoW .text (4MB) and .rdata (160KB) at original virtual addresses
instead of individual function byte arrays. All CALL targets and float
constants resolve automatically -- no more manual mapGameConstants().
- Bench binary linked at 0x10000000 to avoid address conflict with WoW
PE sections at 0x400000-0xD00000.
Added silicon_sse.zig: 18 pure math functions extracted from silicon.zig
as export fn (C ABI) for standalone compilation. Covers frustum culling,
bounding volume ops, quaternion slerp, matrix multiplies, trig, etc.
Silicon benchmark results (all 14 new entries pass correctness):
checkBoxLineIntersect: 3.2x (72->22) -- slab AABB intersection
rotateMatByQuat: 3.1x (172->54) -- quat->mat + mat multiply
quatSlerp: 2.5x (454->178)
createZRotMat3x3: 2.5x (140->55)
createRotMat3x4: 2.2x (163->73)
normalizeVec3InPlace: 1.8x (34->18)
mulMat3x4InPlace: 1.6x (106->66) -- MISMATCH (layout diff, needs investigation)
classifyPointFrustum: 1.5x (60->40)
mulMat3x4: 1.2x (64->52) -- MISMATCH (same layout issue)
Two MISMATCH entries on mat3x4 multiply -- likely row/column order
difference between original and our implementation. Correctness needs
verification against game behavior.
2026-03-15 12:04:02 -07:00
MarcelineVQ
28dd1cf6fa
silicon: implement 32 SSE math replacements for x87 FPU functions
...
Replace probe-only detours with direct SSE/scalar math implementations
for the highest-impact WoW.exe functions by call frequency:
Core transforms (10.8M+ calls/7.5s):
- transformVector3ByMatrix4x4: @Vector(4,f32) dot products
- transformVector4ByMatrix4x4: 4-row vector dots
- MultiplyMatrix3x4, multiplyMatrix3x3: affine matrix multiply
- MultiplyMatrix3x4InPlace: aliasing-safe with stack temp
Frustum/collision (8.2M+ combined):
- ClassifyPointAgainstFrustum: 6-plane dot + bitmask
- CheckBoxLineIntersection: slab-method AABB test
- TestSphereAgainstFrustum: 6-plane sphere test
- TestOBBAgainstFrustum: 8-corner OBB vs 6 planes
- IsPointInsideBounds: 3-component compare
- CalculateDistanceToPlane: ray-plane intersection
Matrix builders (300K+):
- createAxisAngleRotationMatrix 4x4/3x3/3x4: Rodrigues formula
- createZRotationMatrix3x3: sin/cos rotation
- rotateMatrixByQuaternion: quat→matrix + 4x4 multiply
- getTransposedMatrix4x4: inline transpose
- scaleMatrix3x3ByVector, ApplyTranslationMatrix
Vector/scalar:
- normalizeVector3, NormalizeVector3_InPlace
- Vector3_DotProduct, quaternion_slerp
- addVector3ToAccumulator, addToColorAccumulator
- calculateSinCos
Particle:
- packParticleColorToBytes, setParticleAlphaFromFloat
Bounding volume:
- TranslateBoundingVolume, TransformBoundingVolume
84 functions remain as probe-only (orchestrators, Lua internals,
complex renderers — optimization is in their callees, now replaced).
2026-03-15 11:18:35 -07:00
MarcelineVQ
05e016da46
silicon: fix 5 param count bugs, add periodic probe reporting
...
Fix calling convention mismatches found via systematic RET purge audit:
- interpolateKeyframes: FC4d→FC3d (3 params, not 4)
- ConvertPixelsToScreenAlt: TC2r→TC2d (returns x87 float, not u32)
- UpdateObjectTransform: FC2r→FC3r (3 params, not 2)
- packParticleColorToBytes: FC4v→FC5v (5 params, not 4)
- setParticleAlphaFromFloat: FC2v→FC3v (3 params, not 2)
Disable GetFPUControlWord probe (corrupts FLDCW state).
Use naked asm for __ftol probe (preserves implicit x87 ST(0)).
Add OnWorldUpdate hook for periodic hit count dumps every ~7.5s.
All 115 active probes verified with non-zero counts in-game.
2026-03-15 03:41:28 -07:00
MarcelineVQ
e70015579f
silicon: add 116 probe hooks for call-counting all explored functions
...
Probe infrastructure using comptime probeDetour() that generates
per-function detours: atomic counter increment + callOriginal passthrough.
Hit counts reported on shutdown. Hooks installed at lateInit (engine init),
not DLL load time.
Key findings during decompilation:
- 0x40CF81 is GetFPUControlWord, not ftol (silicon mislabel)
- Real __ftol at 0x40A2B0 (51+ callers, hot path)
- 0x7B7A80/7B7B10 use normal stack floats, not FPU register params
- 0x686640/686820/6868E0 are bounding volume ops, not vector ops
- 0x7786A0 is a UI model constructor, not SetModelLighting
2026-03-15 01:55:09 -07:00
MarcelineVQ
66696e993a
silicon: decompile and annotate all ~100 hooked WoW.exe functions
...
Complete catalog of libSiliconPatch240.dll hook targets with verified
calling conventions, parameter layouts, RET stack cleanup, and functional
descriptions from Ghidra analysis. Zero unknown stubs remain.
2026-03-15 01:22:43 -07:00