Commit Graph
18 Commits
Author SHA1 Message Date
MarcelineVQ 99b4c8b2d0 bone_sse: f32 findInterpIdx t division, non-inline sections with quota — 3697 cycles (-6%) 2026-03-16 14:12:53 -07:00
MarcelineVQ 22f9e712f4 bench: 100% code path coverage — multi-track ranges, FloatTrack12 mode=0, 9120 bytes parity PASS 2026-03-16 12:34:07 -07:00
MarcelineVQ 4776a979c9 bench: 94% code path coverage, 8988 bytes parity check — all interp modes, billboards, particles, crossfade, attachments 2026-03-16 12:29:31 -07:00
MarcelineVQ 5baa49d97a bench: fix frame_ctr=0 so section functions execute — 3249 cycles, parity PASS 2026-03-16 12:20:22 -07:00
MarcelineVQ 9c215a0bb1 bench: full parity check across all 13 output buffers (8136 bytes), dual BASELINE/SSE 2026-03-16 12:12:57 -07:00
MarcelineVQ 4ed5495219 bench: dual BASELINE/SSE with parity check — 2177 vs 2112 cycles, PASS 2026-03-16 12:10:51 -07:00
MarcelineVQ 4198a00c26 bench: full path coverage + determinism validation — 2065 cycles/call, PASS 2026-03-16 11:57:37 -07:00
MarcelineVQ 75bb38f5af bench: full coverage fixture — 1012 cycles/call (ribbon, particle, attach, billboard, crossfade, clamped, GS, time delta) 2026-03-16 11:50:38 -07:00
MarcelineVQ b6ab4ac59f bench: comprehensive fixture exercising all code paths — 636 cycles/call baseline 2026-03-16 11:46:16 -07:00
MarcelineVQ f119c15b37 bench: comprehensive transform44 fixture — 578 cycles/call baseline (12 bones + texAnim + colorAnim + wordAnim + boneKF) 2026-03-16 11:42:31 -07:00
MarcelineVQ 93100a1a7e bench: add transform44 SSE benchmark — 288 cycles/call baseline (8 bones, 4 animated) 2026-03-16 11:30:06 -07:00
MarcelineVQ c9249c6c96 bench: fix 3 correctness bugs found by benchmarker
mulMat3x4 / mulMat3x4InPlace: translation row had A*B operands swapped.
Original computes A_translation * B_rotation + B_translation, but our
code was doing A_rotation * B_translation + A_translation. Verified by
tracing x87 disassembly: first element loads A[9]*B[col] pattern.
Fixed in both silicon_sse.zig and silicon.zig.

packParticleColor: x87 rounds 127.5 to 128 (round-to-nearest), but SSE
@intFromFloat truncates to 127. Added @round() before @intFromFloat.

42/42 benchmarks now pass correctness. 0 MISMATCHes.
2026-03-15 12:27:59 -07:00
MarcelineVQ 7c88423a41 bench: add remaining silicon functions, total 38 benchmarks
Added 8 more silicon SSE functions to benchmark: normalizeVec3,
testOBBFrustum, calculateSinCos, translateBoundingVol, addToColorAccum,
packParticleColor, setParticleAlpha, and addVec3ToAccumulator (export fn).

Notable results from new entries:
  packParticleColor: 6.8x (109->16) MISMATCH -- alpha byte diff
  setParticleAlpha:  3.4x (24->7)
  calculateSinCos:   3.0x (138->45)
  normalizeVec3:     1.8x (18->10)
  testOBBFrustum:    1.3x (107->81)

3 MISMATCHes total (mulMat3x4, mulMat3x4InPlace, packParticleColor)
need correctness investigation -- likely matrix layout and struct
byte-order differences.
2026-03-15 12:08:59 -07:00
MarcelineVQ 55b4931fcb bench: map WoW PE sections for full-fidelity benchmarking, add silicon SSE
Major bench harness upgrade:
- Maps WoW .text (4MB) and .rdata (160KB) at original virtual addresses
  instead of individual function byte arrays. All CALL targets and float
  constants resolve automatically -- no more manual mapGameConstants().
- Bench binary linked at 0x10000000 to avoid address conflict with WoW
  PE sections at 0x400000-0xD00000.

Added silicon_sse.zig: 18 pure math functions extracted from silicon.zig
as export fn (C ABI) for standalone compilation. Covers frustum culling,
bounding volume ops, quaternion slerp, matrix multiplies, trig, etc.

Silicon benchmark results (all 14 new entries pass correctness):
  checkBoxLineIntersect: 3.2x (72->22)  -- slab AABB intersection
  rotateMatByQuat:       3.1x (172->54) -- quat->mat + mat multiply
  quatSlerp:             2.5x (454->178)
  createZRotMat3x3:      2.5x (140->55)
  createRotMat3x4:       2.2x (163->73)
  normalizeVec3InPlace:  1.8x (34->18)
  mulMat3x4InPlace:      1.6x (106->66) -- MISMATCH (layout diff, needs investigation)
  classifyPointFrustum:  1.5x (60->40)
  mulMat3x4:             1.2x (64->52)  -- MISMATCH (same layout issue)

Two MISMATCH entries on mat3x4 multiply -- likely row/column order
difference between original and our implementation. Correctness needs
verification against game behavior.
2026-03-15 12:04:02 -07:00
MarcelineVQ 008d74ddb6 bench: add original x87 bytes for 22 silicon module functions
Extracted via Ghidra from WoW.exe for the 22 silicon SSE functions that
don't overlap with ssemaths addresses. These cover frustum culling,
bounding volume transforms, quaternion slerp, matrix operations, and
various geometry functions. Ready for benchmarking.
2026-03-15 11:26:38 -07:00
MarcelineVQ b9603c75f9 bench: add inlined x87 vs SSE comparison, update release notes
Inlined benchmarks use x87 inline asm vs direct Zig @Vector code with
no CALL/RET on either side. Shows the true instruction-level comparison
that in-place patching would achieve:

  dotProduct:  x87=3 SSE=1 -> 3.0x (was 0.3x when called)
  evalPoly:    x87=4 SSE=1 -> 4.0x (was 0.7x when called)
  vec3MulScal: x87=3 SSE=2 -> 1.5x (was 0.8x when called)

Every "loser" from the called benchmarks flips to a winner when inlined.
The entire performance gap was function call overhead (~5 cycles), not
instruction quality. Confirms in-place patching as the right strategy.

Also added RADV_TEX_ANISO env var exploration item to release notes and
updated math polyfill section with full benchmark breakdown.
2026-03-14 22:12:04 -07:00
MarcelineVQ 99503f883f bench: fresh data each iteration to fix overflow artifacts
In-place functions (vec3MulAssign, scaleByVec, etc.) were showing fake
speedups (121x, 10x) because repeated application on the same data
caused values to overflow to Inf/NaN. x87 FPU traps on denormals while
SSE handles them in hardware, making the comparison meaningless.

Now each iteration resets from a template copy. Real results show these
functions are ~1.0x (neutral), not 100x wins. The actual winners are
rotMat3x3/4x4 at 2.1x, and several functions where our scalar-through-
pointers approach is slower than the original x87 pipeline (dotProduct
0.3x, evaluatePolynomial 0.4x) — candidates for real SIMD optimization.
2026-03-14 21:41:56 -07:00
MarcelineVQ db28746182 bench: add x86 Linux micro-benchmark harness for math_sse functions
Extracts original x87 FPU bytes from WoW.exe via Ghidra, mmaps them
executable, and benchmarks against our SSE replacements. Covers all 17
UnitXP polyfill functions with correctness validation and cycle counts.

Maps a page at 0x7ff000 for the float 1.0 constant referenced by
rotMat3x3/rotMat4x4/planeNormal via absolute address 0x7ff9d8.

Build: zig build bench / zig build run-bench
2026-03-14 19:29:41 -07:00