0671ffce02
Added V4 vector helpers (loadV3, storeV3, dot3) and rewrote several functions to use @Vector(4, f32) operations instead of scalar f64: - planeNormal: 0.7x -> 1.7x (79 -> 33 cyc) -- V4 cross + normalize - transformAABox: 0.7x -> 1.2x (118 -> 69 cyc) -- f32 + @min/@max - evaluatePolynomial: 0.4x -> 0.6x (24 -> 16 cyc) -- drop f64 promotion - vec3MulScalar: 0.7x -> 0.8x -- V4 splat multiply - dotProduct/squaredMagnitude: rewritten with dot3 helper crossProduct reverted from shuffle-SIMD back to scalar -- shuffles added latency that the x87 pipeline doesn't have (16 vs 13 cyc). Remaining losers (dotProduct 0.4x, evalPoly 0.6x) are at the function call overhead floor -- the original x87 is 5-11 cycles, which is close to bare call/ret cost.