Commit Graph
30 Commits
Author SHA1 Message Date
MarcelineVQ d205693dbb particle: Ghidra decompilations, particle_sse.zig scaffold, bench
- Ghidra C decompilation of RenderParticleSprites (422 lines) and 5
  helper functions (calculateColorValues, matVec3Transform, etc.)
- particle_sse.zig with calcColorValues_SSE (10.7x bench but cache-miss
  bound in-game — needs inlining into full function replacement)
- Bench harness for calcColorValues with correctness check
- build.zig: particle_sse as separate ReleaseFast compilation unit
- colorDetour reverted to pass-through (SSE has no in-game effect due
  to L1 cache misses on scattered ColorCtx structs)
2026-03-23 21:56:57 -07:00
MarcelineVQ 035355b757 perf: FrustumCullBoundingBox SSE replacement, 1.6x speedup
- silicon_sse: si_frustumCullBBox (0x686000) — inline V4 mat*vec3
  transforms, SSE perspective divide, 4-wide horizon buffer scan.
  Benched 1.6x (117→72 cyc/call). Installed via JMP patch (544 bytes,
  won't fit in 380-byte original).
- bench: add frustumCullBBox benchmark with mapped globals, identity
  matrices, and horizon buffer test fixture.
- Remove patch table entry for frustumCullBBox, install via detour hook
  instead (allows future A/B testing if needed).
2026-03-23 20:46:28 -07:00
MarcelineVQ 7819d6d914 perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44
- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's
  search result reused for scale/translation when tracks share temporal
  structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction.
- silicon: SSE replacement for processLinkedListCollision (0x6ABC40,
  1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW.
  Benched at 3.2x speedup (378→115 cyc/call, 8 nodes).
- transform44: remove A/B toggle, always use SSE path (teardown guard
  kept). A/B infrastructure remains for other hooks.
- bench: add processLinkedListCollision benchmark with fake linked list
  test fixture and stubbed addGeometryToBuffer.
2026-03-23 20:27:26 -07:00
MarcelineVQ 10bd922cc3 silicon: fix CC mismatches (normalizeVec3InPlace TC, packParticleColor TC, addVec3ToAccumulator remove phantom scale param, revert classifyPointFrustum/testOBBFrustum to TC); runtime JMP address resolution 2026-03-17 13:33:51 -07:00
MarcelineVQ dc37042e6a silicon: naked asm for transpose/setAlpha/addColor/normalize, remove vec3Dot/distPlane/insideBounds from patches, direct patch transpose (7cy, 3.7x) 2026-03-17 03:52:50 -07:00
MarcelineVQ 4152d1dc78 silicon: JMP patch infrastructure + native CC for all functions, no naked asm except ftol 2026-03-17 03:12:28 -07:00
MarcelineVQ 6140a3a333 silicon_sse: native calling conventions (thiscall/fastcall/stdcall) for all functions, ready for JMP patching 2026-03-17 02:51:50 -07:00
MarcelineVQ 9f2fa5213a bench: best-of-5 for all silicon functions; silicon_sse: V4 column mulMat3x4/InPlace (1.8x/1.7x) 2026-03-17 01:43:16 -07:00
MarcelineVQ 662c854376 silicon_sse: patch-in-place isPointInsideBounds 6cy->5cy (1.2x), naked vucomiss 2026-03-17 01:26:55 -07:00
MarcelineVQ cb0888e0fa silicon_sse: naked FMA asm for vec3Dot (0.4x->0.8x) and distanceToPlane (0.7x->1.0x) 2026-03-17 01:01:07 -07:00
MarcelineVQ 8537587df3 silicon: si_ftol SSE3 FISTTP replacement, bench patch-in-place framework
si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.

Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.

Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
2026-03-17 00:22:07 -07:00
MarcelineVQ 0419f39833 bench: 2M iterations, baseline 4176 cycles, SSE 3841 (-8%), parity PASS 2026-03-16 15:25:00 -07:00
MarcelineVQ 99b4c8b2d0 bone_sse: f32 findInterpIdx t division, non-inline sections with quota — 3697 cycles (-6%) 2026-03-16 14:12:53 -07:00
MarcelineVQ 22f9e712f4 bench: 100% code path coverage — multi-track ranges, FloatTrack12 mode=0, 9120 bytes parity PASS 2026-03-16 12:34:07 -07:00
MarcelineVQ 4776a979c9 bench: 94% code path coverage, 8988 bytes parity check — all interp modes, billboards, particles, crossfade, attachments 2026-03-16 12:29:31 -07:00
MarcelineVQ 5baa49d97a bench: fix frame_ctr=0 so section functions execute — 3249 cycles, parity PASS 2026-03-16 12:20:22 -07:00
MarcelineVQ 9c215a0bb1 bench: full parity check across all 13 output buffers (8136 bytes), dual BASELINE/SSE 2026-03-16 12:12:57 -07:00
MarcelineVQ 4ed5495219 bench: dual BASELINE/SSE with parity check — 2177 vs 2112 cycles, PASS 2026-03-16 12:10:51 -07:00
MarcelineVQ 4198a00c26 bench: full path coverage + determinism validation — 2065 cycles/call, PASS 2026-03-16 11:57:37 -07:00
MarcelineVQ 75bb38f5af bench: full coverage fixture — 1012 cycles/call (ribbon, particle, attach, billboard, crossfade, clamped, GS, time delta) 2026-03-16 11:50:38 -07:00
MarcelineVQ b6ab4ac59f bench: comprehensive fixture exercising all code paths — 636 cycles/call baseline 2026-03-16 11:46:16 -07:00
MarcelineVQ f119c15b37 bench: comprehensive transform44 fixture — 578 cycles/call baseline (12 bones + texAnim + colorAnim + wordAnim + boneKF) 2026-03-16 11:42:31 -07:00
MarcelineVQ 93100a1a7e bench: add transform44 SSE benchmark — 288 cycles/call baseline (8 bones, 4 animated) 2026-03-16 11:30:06 -07:00
MarcelineVQ c9249c6c96 bench: fix 3 correctness bugs found by benchmarker
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.

packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.

42/42 benchmarks now pass correctness. 0 MISMATCHes.
2026-03-15 12:27:59 -07:00
MarcelineVQ 7c88423a41 bench: add remaining silicon functions, total 38 benchmarks
Added 8 more silicon SSE functions to benchmark: normalizeVec3,
testOBBFrustum, calculateSinCos, translateBoundingVol, addToColorAccum,
packParticleColor, setParticleAlpha, and addVec3ToAccumulator (export fn).

Notable results from new entries:
  packParticleColor: 6.8x (109->16) MISMATCH -- alpha byte diff
  setParticleAlpha:  3.4x (24->7)
  calculateSinCos:   3.0x (138->45)
  normalizeVec3:     1.8x (18->10)
  testOBBFrustum:    1.3x (107->81)

3 MISMATCHes total (mulMat3x4, mulMat3x4InPlace, packParticleColor)
need correctness investigation -- likely matrix layout and struct
byte-order differences.
2026-03-15 12:08:59 -07:00
MarcelineVQ 55b4931fcb bench: map WoW PE sections for full-fidelity benchmarking, add silicon SSE
Major bench harness upgrade:
- Maps WoW .text (4MB) and .rdata (160KB) at original virtual addresses
  instead of individual function byte arrays. All CALL targets and float
  constants resolve automatically -- no more manual mapGameConstants().
- Bench binary linked at 0x10000000 to avoid address conflict with WoW
  PE sections at 0x400000-0xD00000.

Added silicon_sse.zig: 18 pure math functions extracted from silicon.zig
as export fn (C ABI) for standalone compilation. Covers frustum culling,
bounding volume ops, quaternion slerp, matrix multiplies, trig, etc.

Silicon benchmark results (all 14 new entries pass correctness):
  checkBoxLineIntersect: 3.2x (72->22)  -- slab AABB intersection
  rotateMatByQuat:       3.1x (172->54) -- quat->mat + mat multiply
  quatSlerp:             2.5x (454->178)
  createZRotMat3x3:      2.5x (140->55)
  createRotMat3x4:       2.2x (163->73)
  normalizeVec3InPlace:  1.8x (34->18)
  mulMat3x4InPlace:      1.6x (106->66) -- MISMATCH (layout diff, needs investigation)
  classifyPointFrustum:  1.5x (60->40)
  mulMat3x4:             1.2x (64->52)  -- MISMATCH (same layout issue)

Two MISMATCH entries on mat3x4 multiply -- likely row/column order
difference between original and our implementation. Correctness needs
verification against game behavior.
2026-03-15 12:04:02 -07:00
MarcelineVQ 008d74ddb6 bench: add original x87 bytes for 22 silicon module functions
Extracted via Ghidra from WoW.exe for the 22 silicon SSE functions that
don't overlap with ssemaths addresses. These cover frustum culling,
bounding volume transforms, quaternion slerp, matrix operations, and
various geometry functions. Ready for benchmarking.
2026-03-15 11:26:38 -07:00
MarcelineVQ b9603c75f9 bench: add inlined x87 vs SSE comparison, update release notes
Inlined benchmarks use x87 inline asm vs direct Zig @Vector code with
no CALL/RET on either side. Shows the true instruction-level comparison
that in-place patching would achieve:

  dotProduct:  x87=3 SSE=1 -> 3.0x (was 0.3x when called)
  evalPoly:    x87=4 SSE=1 -> 4.0x (was 0.7x when called)
  vec3MulScal: x87=3 SSE=2 -> 1.5x (was 0.8x when called)

Every "loser" from the called benchmarks flips to a winner when inlined.
The entire performance gap was function call overhead (~5 cycles), not
instruction quality. Confirms in-place patching as the right strategy.

Also added RADV_TEX_ANISO env var exploration item to release notes and
updated math polyfill section with full benchmark breakdown.
2026-03-14 22:12:04 -07:00
MarcelineVQ 99503f883f bench: fresh data each iteration to fix overflow artifacts
In-place functions (vec3MulAssign, scaleByVec, etc.) were showing fake
speedups (121x, 10x) because repeated application on the same data
caused values to overflow to Inf/NaN. x87 FPU traps on denormals while
SSE handles them in hardware, making the comparison meaningless.

Now each iteration resets from a template copy. Real results show these
functions are ~1.0x (neutral), not 100x wins. The actual winners are
rotMat3x3/4x4 at 2.1x, and several functions where our scalar-through-
pointers approach is slower than the original x87 pipeline (dotProduct
0.3x, evaluatePolynomial 0.4x) — candidates for real SIMD optimization.
2026-03-14 21:41:56 -07:00
MarcelineVQ db28746182 bench: add x86 Linux micro-benchmark harness for math_sse functions
Extracts original x87 FPU bytes from WoW.exe via Ghidra, mmaps them
executable, and benchmarks against our SSE replacements. Covers all 17
UnitXP polyfill functions with correctness validation and cycle counts.

Maps a page at 0x7ff000 for the float 1.0 constant referenced by
rotMat3x3/rotMat4x4/planeNormal via absolute address 0x7ff9d8.

Build: zig build bench / zig build run-bench
2026-03-14 19:29:41 -07:00