Commit Graph
222 Commits
Author SHA1 Message Date
MarcelineVQ d52d4bc661 docs: add DPSLog to DLL_README and release notes
- DLL_README.md: full feature description, GetSpellInfo API, module name in dev API list
- RELEASE_NOTES.md: updated DPSLog entry with WotLK parity details, removed from gaps list
- README.md: updated module table
2026-03-18 00:16:57 -07:00
MarcelineVQ 56f5c0405d dpslog: WotLK CLEU parity — overkill, overheal, spell names, unit names, GetSpellInfo
Full WotLK 3.3.5 COMBAT_LOG_EVENT_UNFILTERED layout parity:

Event structure:
- Base: subevent, sourceGUID, sourceName, destGUID, destName
- Spell prefix: spellId, spellName, spellSchool
- All suffix field orders match WotLK (blocked before absorbed, etc.)

New fields:
- overkill (-1 if alive) on all _DAMAGE events
- overheal on SPELL_HEAL and SPELL_PERIODIC_HEAL
- sourceName/destName via getNameByGUID (name cache 0x55F080)
- spellName via SpellRec+0x1E0+locale*4 (verified from SuperWoW)
- glancing/crushing as separate booleans (was packed flags)
- amountMissed field (0 — downstream hook lacks amount)

New hooks:
- ProcessStandardPowerGainMessage (0x62CA00) for SPELL_ENERGIZE with
  actual amount and powerType from SMSG_SPELLENERGIZELOG (opcode 0x151)
- Periodic energize: divide by power display factor (rage/10, happiness/1000)

Leech/drain reclassification:
- Drain Life routed through SPELL_PERIODIC_LEECH (aura 53 check in spell DB)
- SPELL_HEAL suppressed for leech spells (covered by leech event)
- Power leech (aura 64): gainMultiplier read as f32, >0 = LEECH, ==0 = DRAIN
- _LEECH/_DRAIN suffix: amount, powerType(-2=health), extraAmount

Lua API:
- GetSpellInfo(spellId) -> name, rank, icon, castTime, minRange, maxRange, spellId
  Matches WotLK returns. DBC addresses verified from nampower + Ghidra:
  SpellRange.dbc (0xC0D79C), SpellIcon.dbc (0xC0D7EC),
  SpellCastTimes.dbc (0xC0D878). SpellRec offsets from nampower struct.

Reference docs:
- WOTLK_CLEU_SPEC.md: full suffix spec from Blizzard_CombatLog.lua + Skada
- WARDEN_CONCERNS.md: hook risk analysis and remaining parity gaps
2026-03-18 00:08:42 -07:00
MarcelineVQ 76f0691540 dpslog: fix DAMAGE_SPLIT aura offset (0x178->0x16C for client Spell.dbc), fix SPELL_EXTRA_ATTACKS count, DEFLECT note, EXHAUSTED note 2026-03-17 16:02:39 -07:00
MarcelineVQ aadd467976 dpslog: SPELL_EXTRA_ATTACKS now includes count (EDX was the count all along, not unused) 2026-03-17 15:18:21 -07:00
MarcelineVQ 2ad6ed380e dpslog addon: remove SWING_MISSED DEFLECT (melee can't deflect), add SPELL_MISSED DEFLECT (dead in vanilla server, kept for protocol completeness) 2026-03-17 15:06:06 -07:00
MarcelineVQ 7089ffa399 dpslog addon: variant-expanded tracker checklist (67 entries) — miss types, crit/glancing/crushing, env types, aura BUFF/DEBUFF, power types 2026-03-17 15:01:51 -07:00
MarcelineVQ 8bd82aeea7 dpslog: SPELL_CAST_FAILED for other units via SMSG_SPELL_FAILED_OTHER (0x6E75F0); addon tracks DAMAGE_SPLIT, SPELL_DISPEL_FAILED, UNIT_DESTROYED (37 subevents) 2026-03-17 14:57:22 -07:00
MarcelineVQ 3eff6928b9 dpslog: UNIT_DESTROYED for totems (creature type 11 via name cache), DAMAGE_SPLIT, SPELL_DISPEL_FAILED 2026-03-17 14:39:42 -07:00
MarcelineVQ fec7387148 dpslog: add DAMAGE_SPLIT (Soul Link/BoSac via spell DB aura check) and SPELL_DISPEL_FAILED (hook ProcessMultipleSpellInterrupts 0x628C20) 2026-03-17 14:27:31 -07:00
MarcelineVQ 10bd922cc3 silicon: fix CC mismatches (normalizeVec3InPlace TC, packParticleColor TC, addVec3ToAccumulator remove phantom scale param, revert classifyPointFrustum/testOBBFrustum to TC); runtime JMP address resolution 2026-03-17 13:33:51 -07:00
MarcelineVQ dc37042e6a silicon: naked asm for transpose/setAlpha/addColor/normalize, remove vec3Dot/distPlane/insideBounds from patches, direct patch transpose (7cy, 3.7x) 2026-03-17 03:52:50 -07:00
MarcelineVQ 3dfa559eac silicon: direct byte patch for ftol, JMP patches for all others, removed compare mode 2026-03-17 03:20:59 -07:00
MarcelineVQ 4152d1dc78 silicon: JMP patch infrastructure + native CC for all functions, no naked asm except ftol 2026-03-17 03:12:28 -07:00
MarcelineVQ 6140a3a333 silicon_sse: native calling conventions (thiscall/fastcall/stdcall) for all functions, ready for JMP patching 2026-03-17 02:51:50 -07:00
MarcelineVQ 48e542dcf4 silicon_sse: packParticleColor 26cy->16cy (6.3x) via packed V4 round+convert; checkBoxLineIntersect branchless @min/@max cleanup; all tasks complete 2026-03-17 02:31:01 -07:00
MarcelineVQ 741d44c973 silicon_sse: FSINCOS tested (slower), normalizeVec3/translateBoundingVol/createRotMat3x4 documented at optimum 2026-03-17 02:24:56 -07:00
MarcelineVQ 350cd93162 silicon_sse: normalizeVec3InPlace + testSphereFrustum at optimum, documented attempts 2026-03-17 02:06:08 -07:00
MarcelineVQ d1c3c4654f silicon_sse: batched testOBBFrustum 79cy->43cy (2.3x), 4-corner V4 dots + branchless mask 2026-03-17 01:49:37 -07:00
MarcelineVQ 9715e32ed7 silicon_sse: mulMat3x4InPlace eliminate tmp copy + preload b (1.7x -> 1.8x) 2026-03-17 01:47:13 -07:00
MarcelineVQ 9f2fa5213a bench: best-of-5 for all silicon functions; silicon_sse: V4 column mulMat3x4/InPlace (1.8x/1.7x) 2026-03-17 01:43:16 -07:00
MarcelineVQ 662c854376 silicon_sse: patch-in-place isPointInsideBounds 6cy->5cy (1.2x), naked vucomiss 2026-03-17 01:26:55 -07:00
MarcelineVQ 16f1a073ae silicon_sse: branchless classifyPointFrustum 35cy->30cy (1.8x), revert addToColorAccum 2026-03-17 01:17:23 -07:00
MarcelineVQ cb0888e0fa silicon_sse: naked FMA asm for vec3Dot (0.4x->0.8x) and distanceToPlane (0.7x->1.0x) 2026-03-17 01:01:07 -07:00
MarcelineVQ 9488fb568c silicon_sse: V4/@mulAdd/@shuffle rewrites for all functions
Major improvements from SSE4.1+FMA+AVX target + explicit SIMD:
- transposeMat4x4: 0.9x -> 2.4x (V4 shuffle)
- rotateMatByQuat: 4.0x -> 5.0x (V4 matmul)
- testOBBFrustum: 0.9x -> 1.2x (V4 corner transform + dot4)
- classifyPointFrustum: 1.4x -> 2.0x (V4 dot4 with {x,y,z,1} trick)
- translateBoundingVol: 1.3x -> 1.7x (@mulAdd plane distances)
- testSphereFrustum: 1.1x -> 1.3x (V4 dot)
- createRotMat3x4: @mulAdd for all 9 matrix entries
- mulMat3x4/InPlace: @mulAdd chains
- quatSlerp: V4 blend + @mulAdd dot
- addVec3ToAccumulator: @mulAdd for scale multiply

All parity tests pass.
2026-03-17 00:34:12 -07:00
MarcelineVQ bad3126973 build: enable SSE4.1+FMA+AVX for silicon_sse compilation unit
Was compiling with baseline SSE2 only. Now matches bone_sse target.
Free wins: packParticleColor 1.2x->4.6x, rotateMatByQuat 3.4x->4.0x,
normalizeVec3 1.1x->1.8x, mulMat3x4InPlace 1.4x->1.7x.
2026-03-17 00:27:48 -07:00
MarcelineVQ b27a987355 silicon: FTOL_ONLY debug flag, ftolSSE2 compare mode, h67 disabled
Temporary debug state for isolating visual issues:
- FTOL_ONLY gates all hooks except __ftol and world update reporter
- ftolSSE2 compare mode: calls original + SSE2, counts mismatches
- h67 (ConvertPixelsToScreenAlt 0x5C7010) disabled: crashes with ECX=0
2026-03-17 00:22:21 -07:00
MarcelineVQ 8537587df3 silicon: si_ftol SSE3 FISTTP replacement, bench patch-in-place framework
si_ftol: 9-byte naked asm using FISTTP (SSE3 truncate-from-x87) replaces
the 39-byte FSTCW/FLDCW/FISTP rounding mode dance. 4 vs 7 cycles (1.7x).
13.2M calls/7.5s in-game -- ~13ms savings per period.

Benchmark: patch-in-place at mapped 0x40A2B0, test parity across 19 values,
best-of-5 timing with varying inputs. Framework for all silicon functions.

Also disabled h67 (ConvertPixelsToScreenAlt) probe -- game passes ECX=0
as valid input, thiscall probe crashes on null this.
2026-03-17 00:22:07 -07:00
MarcelineVQ 61ee4f48e1 bone_sse: f32 callFtol, fastMod conditional subtract for looping anims
callFtol: use f32 multiply instead of f64 intermediate. Parity holds --
the delta*scale product is well within f32 precision range.

fastMod: replace integer modulo (idiv, ~25 cycles) with conditional
subtract (~2 cycles) for looping animation frame computation. Falls
back to real modulo for large time skips (alt-tab, etc).

3609 cycles (-14% vs 4176 baseline), parity PASS.
2026-03-16 17:06:20 -07:00
MarcelineVQ dd43f6aca3 bone_sse: return interp values in registers, hoist runtime constants, value-based local matrix
- interpAnimKF returns [4]f32, interpVec3Track returns [3]f32, interpFloatTrack returns f32
- Internal crossfade blends stay in registers instead of writing then re-reading from memory
- Bone loop uses returned values directly for buildRotationMatrix/scaleMatrix3x3/translation
- Hoist getShortToFloat() reads to function entry in section functions (Proposal D)
- Remove dead blendVec3, getHermite5, HERMITE_3
- buildRotationMatrixVal returns [16]f32; bone loop uses array ops instead of u32 pointer casts
- matMul4x4Local/matMul4x4InPlace variants for local array operands
- shortInterpToFloat takes pre-read stf parameter

3574 cycles (-14% vs 4176 baseline), parity PASS (9120 bytes).
2026-03-16 16:25:26 -07:00
MarcelineVQ d9d2e41eb4 bone_sse: return InterpResult in registers, eliminate store-forward latency
findInterpIdx now returns {idx0, idx1, t} as a struct instead of writing
all three to the output buffer. Only output[0] is written for next-frame
cache persistence. All 29 call sites updated to use returned values.

3574 cycles (-14% vs 4176 baseline), was 3841 (-8%). Parity PASS.
2026-03-16 15:37:28 -07:00
MarcelineVQ 0419f39833 bench: 2M iterations, baseline 4176 cycles, SSE 3841 (-8%), parity PASS 2026-03-16 15:25:00 -07:00
MarcelineVQ 99b4c8b2d0 bone_sse: f32 findInterpIdx t division, non-inline sections with quota — 3697 cycles (-6%) 2026-03-16 14:12:53 -07:00
MarcelineVQ 502fd9ba04 bone_sse: remove align(1) for naturally-aligned game data — 3641 cycles (was 3744) 2026-03-16 12:43:51 -07:00
MarcelineVQ 88fec305be bone_sse: aggressive inlining + skip identity init + fused rotateByQuaternion — 3744 cycles (was 4040) 2026-03-16 12:41:54 -07:00
MarcelineVQ 22f9e712f4 bench: 100% code path coverage — multi-track ranges, FloatTrack12 mode=0, 9120 bytes parity PASS 2026-03-16 12:34:07 -07:00
MarcelineVQ 4776a979c9 bench: 94% code path coverage, 8988 bytes parity check — all interp modes, billboards, particles, crossfade, attachments 2026-03-16 12:29:31 -07:00
MarcelineVQ 5baa49d97a bench: fix frame_ctr=0 so section functions execute — 3249 cycles, parity PASS 2026-03-16 12:20:22 -07:00
MarcelineVQ 9c215a0bb1 bench: full parity check across all 13 output buffers (8136 bytes), dual BASELINE/SSE 2026-03-16 12:12:57 -07:00
MarcelineVQ 4ed5495219 bench: dual BASELINE/SSE with parity check — 2177 vs 2112 cycles, PASS 2026-03-16 12:10:51 -07:00
MarcelineVQ 4198a00c26 bench: full path coverage + determinism validation — 2065 cycles/call, PASS 2026-03-16 11:57:37 -07:00
MarcelineVQ 75bb38f5af bench: full coverage fixture — 1012 cycles/call (ribbon, particle, attach, billboard, crossfade, clamped, GS, time delta) 2026-03-16 11:50:38 -07:00
MarcelineVQ b6ab4ac59f bench: comprehensive fixture exercising all code paths — 636 cycles/call baseline 2026-03-16 11:46:16 -07:00
MarcelineVQ f119c15b37 bench: comprehensive transform44 fixture — 578 cycles/call baseline (12 bones + texAnim + colorAnim + wordAnim + boneKF) 2026-03-16 11:42:31 -07:00
MarcelineVQ 93100a1a7e bench: add transform44 SSE benchmark — 288 cycles/call baseline (8 bones, 4 animated) 2026-03-16 11:30:06 -07:00
MarcelineVQ 91444f8104 bone_sse: incremental bone pointer advancement, V4 copyMat4 2026-03-16 11:18:50 -07:00
MarcelineVQ 3a522fd1e4 bone_sse: cache frame_ctr, pass to all section functions — eliminates ~29 redundant reads 2026-03-16 11:14:56 -07:00
MarcelineVQ c80e63522b bone_sse: @mulAdd (FMA) for all lerp/blend paths across interp and section functions 2026-03-16 11:06:59 -07:00
MarcelineVQ c36352f4ff bone_sse: @mulAdd (FMA) for lerpVec3, applyTranslation, vec3SqMag 2026-03-16 11:00:01 -07:00
MarcelineVQ a7e2f978e2 bone_sse: simplify rotateByQuaternion to use matMul4x4 V4+FMA path 2026-03-16 10:56:48 -07:00
MarcelineVQ 5d41cd03a3 bone_sse: V4 FMA matmul, enable SSE4.1+FMA+AVX target 2026-03-16 10:54:25 -07:00