bench: map WoW PE sections for full-fidelity benchmarking, add silicon SSE
Major bench harness upgrade:
- Maps WoW .text (4MB) and .rdata (160KB) at original virtual addresses
instead of individual function byte arrays. All CALL targets and float
constants resolve automatically -- no more manual mapGameConstants().
- Bench binary linked at 0x10000000 to avoid address conflict with WoW
PE sections at 0x400000-0xD00000.
Added silicon_sse.zig: 18 pure math functions extracted from silicon.zig
as export fn (C ABI) for standalone compilation. Covers frustum culling,
bounding volume ops, quaternion slerp, matrix multiplies, trig, etc.
Silicon benchmark results (all 14 new entries pass correctness):
checkBoxLineIntersect: 3.2x (72->22) -- slab AABB intersection
rotateMatByQuat: 3.1x (172->54) -- quat->mat + mat multiply
quatSlerp: 2.5x (454->178)
createZRotMat3x3: 2.5x (140->55)
createRotMat3x4: 2.2x (163->73)
normalizeVec3InPlace: 1.8x (34->18)
mulMat3x4InPlace: 1.6x (106->66) -- MISMATCH (layout diff, needs investigation)
classifyPointFrustum: 1.5x (60->40)
mulMat3x4: 1.2x (64->52) -- MISMATCH (same layout issue)
Two MISMATCH entries on mat3x4 multiply -- likely row/column order
difference between original and our implementation. Correctness needs
verification against game behavior.