Commit Graph

77 Commits

Author SHA1 Message Date
MarcelineVQ b12fc820ef Publish source: Unlicense, public README, repo hygiene
The remote was previously a distribution-only point for pre-built DLLs.
This opens the source.

- LICENSE: Unlicense, with a GPL-3.0 carve-out for src/dpslog/WeirdDPSMate
  (a DPSMate fork that keeps its own license)
- README.md replaces the stale internal one with the user-facing docs from
  DLL_README.md, swapping the 'Why No Source Code?' section for build and
  layout notes. DLL_README.md is dropped; one README now serves both.
- RELEASING.md: drop the trim-the-README-per-release dance and the
  remote/WeirdUtils/ distribution clone, both obsolete now
- gitignore agent/editor scratch, build caches, the vendored WSBT addon,
  and the WeirdThreat/uwu-logs checkouts (separate upstream repos)
- Commit outstanding module work: superweirdo, clickthrough portal visuals,
  transform44 decompiles, worldmarkers demo presets, tools/
2026-07-27 21:47:55 -07:00
MarcelineVQ db675c7a7e fix: MSVC ABI, heap allocator, deferred timer - vanillafixes compat
- Restore MSVC ABI (was accidentally GNU since v0.6.0, broke .CRT section)
- Replace game allocator with Windows process heap for filecache and
  libdeflate malloc/free - game allocator not initialized during DllMain
  when injected via CreateRemoteThread
- Defer timer calibration (Sleep 500ms) to lateInit - blocks under
  loader lock during DllMain
- Remove exported malloc/free symbols from DLL
- Eliminate addObject compilation units for SSE files - direct @import
  with AVX target instead
- Heap-allocate filecache (was 9.3MB static BSS)
- Strip transform44 of performance/ externs, pure profiling only
- Rename performance/ to weirdperformance/ to match module convention
- Skip default-off modules in all-variants build step
- Remove dead debug vars and stride logging from particle_sse
2026-03-28 05:31:31 -07:00
MarcelineVQ bf5a7aa624 perf: GUID cache with destruction hook, remove glyph cache
GUID lookup cache: 4096-entry direct-mapped with proper invalidation.
MoveObjectToDeletedList (0x464920) hook evicts entries AFTER the
original runs (prevents re-caching from internal FindObjectByGUID call).
DestroyObjectManager (0x467700) hook flushes on zone change/logout.
93% hit rate, 1.15x speedup on FindObjectByGUID (10K+ calls/frame).

Glyph shadow cache removed: game has internal glyph cache at
GetOrCreateCharacterGlyph (0x5CA2D0). Our hook only saw cache misses
(~30/frame), providing no benefit. The 3.65% perf profile was the
game's own hash table work, not redundant computation.

Standalone guidcache module for isolated testing.
2026-03-26 04:14:34 -07:00
MarcelineVQ 1f318a8565 perf: GUID lookup cache (97% hit, 2.6x)
FindObjectByGUID (0x464890): 4096-entry direct-mapped cache with
validate-on-hit. 97% hit rate at 10K+ calls/frame, reducing from
0.8% to 0.3% frame time. Validates cached pointers by checking
GUID at obj+0x30/+0x34 on every hit.

AddToSpatialGrid (0x6816F0): SSE rewrite attempted, no measurable
gain (memory-bound linked list ops dominate). Not shipped.
2026-03-25 14:46:56 -07:00
MarcelineVQ ad33eb9b90 perf: SSE ray_tri_indexed_int (2.2x), new profiling hooks
rayTriIntersectIndexedInt (0x7C2C40): SSE Moller-Trumbore with deferred
divide, matching original's epsilon thresholds. Parity-tested against
original for edge hits, backfaces, parallel rays, and per-triangle t/uv.
JMP-patched in weirdperformance.

Added transform44 profiling hooks for ray_tri_indexed_int (0x7C2C40)
and ProcessStaticObjectsCulling (0x683BF0).

Bench: added rayTriIndexedInt bench with exhaustive parity tests.
2026-03-25 13:36:33 -07:00
MarcelineVQ 41ca30cfd5 perf: SSE spatial culling, inlined ray-tri, clickthrough CC fix
PerformSpatialCulling (0x6B8C60): Zig rewrite with SSE outcode
computation. 1.4x speedup, JMP-patched in weirdperformance.

performCollisionDetection (0x6B88E0): fully inlined SSE Moller-Trumbore
ray-triangle intersection, eliminating 4 SetVector3 calls and the
ray_tri function pointer call per triangle. ~22 cyc/tri in bench.

Both graduated from transform44 A/B testing to production JMP patches.

entity_sse.zig: reimplementations of UpdateEntityAndChunksPositions and
updateEntitiesInBounds (A/B tested, 1.2x bench, not shipped - memory
bound with negligible real-world gain).

clickthrough: fixed CheckObjectTypePermissions hook from fastcall to
thiscall (ECX preservation), fixed ClntObjMgrObjectPtr from fastcall(4)
to fastcall(5) with correct arg count.

bench: added performCollisionDetection bench with synthetic mesh data
and patched FindOrCreateHashEntry stub.
2026-03-25 12:34:57 -07:00
MarcelineVQ e78ebfd74d particle: fix tail texcoords and transformVec4 input size
Tail vertices V0/V1 read texture U/V from wrong addresses (no offset
from base instead of +8/+16 stride). Assembly reads 0x87D734/738 for
V0 and 0x87D73C/740 for V1, not the base 0x87D72C/730. This made both
base vertices share UVs, collapsing the texture and causing light beams
to taper to a point instead of maintaining width.

Also fix transformVec4 input: must be [4]f32 with w=0.0, not [3]f32.
The function reads all 4 components including [EDX+0xC]. Reading past
a 3-element array produced garbage w values that corrupted the velocity
transform, causing tail particles to extend wildly.

Fixed in both SSE and reference implementations.
2026-03-24 22:01:32 -07:00
MarcelineVQ 979e5729bc perf: consolidate filecache, timer calibration into performance module
Move filecache.zig and timer_fix.zig from standalone modules into
src/performance/. Filecache no longer has its own hooks/mutex — stats
are dumped by performance's worldupdate hook. Timer calibration runs
during performance install instead of transform44. Remove filecache
from build module list (enabled automatically with performance).
Merge DLL_README sections into single Performance entry.
2026-03-24 17:28:19 -07:00
MarcelineVQ 2f0e8a8bc0 fix: particle color byte masking, add reference version
Color channel packing missing & 0xFF after >> 14 extraction — upper
bits bled into adjacent channels causing broken particle fading.
Same class of bug as the alpha output fix earlier.

Added particle_sse_reference.zig (faithful recreation from commit
574f96f) as a separate compilation unit for correctness comparison.
2026-03-24 01:08:01 -07:00
MarcelineVQ 112246e687 perf: make all verified optimizations permanent, remove A/B toggles
- Glyph cache: unconditional (was A/B toggled)
- RenderParticleSprites SSE: unconditional (was A/B toggled)
- transform44 bone SSE: already permanent (teardown guard only)
- processLinkedListCollision: already permanent (JMP patch)
- frustumCullBoundingBox: already permanent (JMP patch)
- silicon functions (ftol, normalize, matmul, etc.): already permanent

Also adds decompilation of UpdateEntityAndChunksPositions — analyzed
but not optimizable (game function calls dominate, math is ~50 cycles
of the ~375 cycle total).
2026-03-24 00:44:18 -07:00
MarcelineVQ 3d872fdb1e spritequad: recreation + analysis (disabled, no improvement)
Faithful recreation of RenderSpriteQuads (0x5A0F50) with hoisted
invariant division and inlined DisplayMode_CalculateOffset. No
measurable improvement — cost is dominated by getAdapterInfo (7
sub-calls to D3D device per invocation × 3203 calls/frame) and
DrawPrimitive/DrawIndexedPrimitive virtual dispatch.

Added decompilation and assembly dumps for future reference.
2026-03-24 00:36:37 -07:00
MarcelineVQ 1046c0a4e3 particle: WIP setupParticleRendering recreation (disabled)
Faithful recreation attempt of SetupParticleRendering (0x7B3D20).
All game function CCs verified from assembly. Vertex buffer setup
works (8 verts produced), but D3D draw submission doesn't produce
visible output. Disabled pending investigation of GfxDeviceMethod
param struct layout. The function remains as timing-only pass-through.

Fixes found during work:
- max_particle_sprites global: 0xCF58F4 → 0xCF5B60
- billboard_matrix global: 0xCF5898 → 0xCF5888
- index_buffer_6/12: 0xCF58D0/D4 → 0xCF5BAC/0xCF5AF4
- BuildIndexBuffer takes renders_count not field_28
- Identity matrices must be mutable (game writes to them)
2026-03-24 00:25:12 -07:00
MarcelineVQ a73b5273fd research: Ghidra decompilation of SetupParticleRendering (326 lines)
10% of frame time, ~275 calls/frame, ~5800 cyc/call. Builds identity
matrices (32 stores of dead work), does 1-5 matmul calls, copies to
g_worldMatrix global. Translation matrix is identity+offset — matmul
chain can be simplified to direct computation.
2026-03-23 22:50:49 -07:00
MarcelineVQ 5373686bd4 particle: V4 vertex store, inline for unroll — ~42% total reduction
- V4 store (vmovups) writes xyz+color as one 16-byte op instead of 4
  scalar stores. Unaligned but still 1 μop on modern CPUs.
- inline for unrolls 4-vertex loops, letting LLVM schedule stores
  across vertices and fill pipeline bubbles.
- Hoisted world_pos to locals to prevent array re-reads.
- A/B: BASELINE ~433ms → CUSTOM ~299ms (~31% per-period reduction).
  Total from original: 520ms → 299ms = 42% reduction.
2026-03-23 22:47:50 -07:00
MarcelineVQ 379fcb62eb particle: contiguous vertex writes, skip normals, stride logging
- Detected interleaved 24-byte vertex layout: xyz(12)+color(4)+uv(8).
  Fast path writes 6 sequential u32s instead of scattered stores.
- Normal stride=0 (shared global) — write once in writeback, not 4x.
- Added stride_info export + logging for vertex layout analysis.
- A/B: ~30% peak reduction, baseline also improved due to less overhead.
2026-03-23 22:41:16 -07:00
MarcelineVQ 0d994db4b3 particle: VBState all paths, cached setupRender, @mulAdd — ~25% speedup
- VBState caching on all 5 vertex paths (was only 2D and sin/cos).
  Eliminates pointer re-reads: load once, emit 4 vertices, writeback.
- Cache setupRender() result per-frame via static + resetParticleCache()
  called from worldUpdateDetour. Saves ~7K function calls/frame.
- @mulAdd throughout for FMA codegen on vertex position and texcoord.
- A/B verified: BASELINE ~521ms → CUSTOM ~387ms (~25% reduction).
2026-03-23 22:29:12 -07:00
MarcelineVQ 1dc1350645 particle: inline calcColor + mat*vec3, VBState caching — ~20% speedup
- Inline calcColor: eliminates function call, allows OoO overlap of
  cache misses on colorCtx with vertex math. Pow path falls back to
  game function.
- Inline mat*vec3 transform: V4 FMA chain replaces call to 0x7BCA80.
- VBState: cache VB pointers/strides in locals, write back once after
  4 vertices. Eliminates ~80 pointer re-reads per particle.
- @mulAdd throughout vertex loops for FMA codegen.
- A/B verified: BASELINE ~470ms → CUSTOM ~382ms (~20% reduction).
2026-03-23 22:22:31 -07:00
MarcelineVQ 574f96f661 particle: faithful RenderParticleSprites recreation, A/B verified
Full recreation of RenderParticleSprites (0x7B2A50, 2688 bytes) in
particle_sse.zig. All 5 code paths: 2D billboard, 3D billboard,
2D+rotation (sin/cos), 3D+rotation (axis-angle matrix), tail particles.

Key fixes during verification:
- colorCtx address: removed Ghidra's spurious -0x12 offset
- calcColor arg2: pass raw u32 from emitter+0x1A8, not truncated
- Texture coord lookups: +8 offset to match assembly's eax increment
  between position and texcoord reads in the vertex loop

Verified in-game: particles render identically in CUSTOM vs BASELINE.
Next: optimize with SSE (inline calcColor, V4 vertex math).
2026-03-23 22:12:17 -07:00
MarcelineVQ d205693dbb particle: Ghidra decompilations, particle_sse.zig scaffold, bench
- Ghidra C decompilation of RenderParticleSprites (422 lines) and 5
  helper functions (calculateColorValues, matVec3Transform, etc.)
- particle_sse.zig with calcColorValues_SSE (10.7x bench but cache-miss
  bound in-game — needs inlining into full function replacement)
- Bench harness for calcColorValues with correctness check
- build.zig: particle_sse as separate ReleaseFast compilation unit
- colorDetour reverted to pass-through (SSE has no in-game effect due
  to L1 cache misses on scattered ColorCtx structs)
2026-03-23 21:56:57 -07:00
MarcelineVQ fec9ac9952 research: particle system SSE optimization analysis
Assembly dumps and research doc for RenderParticleSprites (1.73% CPU),
calculateColorValues (0.63%), SetupParticleRendering, and
ProcessActiveParticles. Identifies SSE opportunities: V4 color interp,
billboard vertex math, rotation block. Plan for particle_sse.zig.
2026-03-23 21:23:19 -07:00
MarcelineVQ 4b5fcd6fa6 perf: glyph cache A/B testing, fast hash, frustumCull rcpss+cvtss2si
- glyph cache: add A/B toggle so BASELINE/CUSTOM periods alternate
  between original function and shadow cache. Swap Murmur2 hash for
  fast golden-ratio integer mix (3 insns vs multi-step).
- frustumCullBBox: fastRecip (vrcpss+NR) and cvtss2si replace vdivss
  and @round bloat. 2.0x speedup (was 1.6x).
- Fix all Zig operator precedence bugs in silicon_sse: & and | bind
  looser than == in Zig, so (flags & 0x8 == 0) was always false.
  Affected frustumCullBBox and processLinkedListCollision.
2026-03-23 21:15:15 -07:00
MarcelineVQ 7819d6d914 perf: findInterpIdx dedup, processLinkedListCollision SSE, permanent t44
- bone_sse: deduplicate findInterpIdx calls in bone loop — rotation's
  search result reused for scale/translation when tracks share temporal
  structure (canReuseInterp guard). Est. ~23% bone loop cycle reduction.
- silicon: SSE replacement for processLinkedListCollision (0x6ABC40,
  1.57% CPU). V4 AABB overlap test replaces 6 x87 FCOMP/FNSTSW.
  Benched at 3.2x speedup (378→115 cyc/call, 8 nodes).
- transform44: remove A/B toggle, always use SSE path (teardown guard
  kept). A/B infrastructure remains for other hooks.
- bench: add processLinkedListCollision benchmark with fake linked list
  test fixture and stubbed addGeometryToBuffer.
2026-03-23 20:27:26 -07:00
MarcelineVQ 61ee4f48e1 bone_sse: f32 callFtol, fastMod conditional subtract for looping anims
callFtol: use f32 multiply instead of f64 intermediate. Parity holds --
the delta*scale product is well within f32 precision range.

fastMod: replace integer modulo (idiv, ~25 cycles) with conditional
subtract (~2 cycles) for looping animation frame computation. Falls
back to real modulo for large time skips (alt-tab, etc).

3609 cycles (-14% vs 4176 baseline), parity PASS.
2026-03-16 17:06:20 -07:00
MarcelineVQ dd43f6aca3 bone_sse: return interp values in registers, hoist runtime constants, value-based local matrix
- interpAnimKF returns [4]f32, interpVec3Track returns [3]f32, interpFloatTrack returns f32
- Internal crossfade blends stay in registers instead of writing then re-reading from memory
- Bone loop uses returned values directly for buildRotationMatrix/scaleMatrix3x3/translation
- Hoist getShortToFloat() reads to function entry in section functions (Proposal D)
- Remove dead blendVec3, getHermite5, HERMITE_3
- buildRotationMatrixVal returns [16]f32; bone loop uses array ops instead of u32 pointer casts
- matMul4x4Local/matMul4x4InPlace variants for local array operands
- shortInterpToFloat takes pre-read stf parameter

3574 cycles (-14% vs 4176 baseline), parity PASS (9120 bytes).
2026-03-16 16:25:26 -07:00
MarcelineVQ d9d2e41eb4 bone_sse: return InterpResult in registers, eliminate store-forward latency
findInterpIdx now returns {idx0, idx1, t} as a struct instead of writing
all three to the output buffer. Only output[0] is written for next-frame
cache persistence. All 29 call sites updated to use returned values.

3574 cycles (-14% vs 4176 baseline), was 3841 (-8%). Parity PASS.
2026-03-16 15:37:28 -07:00
MarcelineVQ 0419f39833 bench: 2M iterations, baseline 4176 cycles, SSE 3841 (-8%), parity PASS 2026-03-16 15:25:00 -07:00
MarcelineVQ 99b4c8b2d0 bone_sse: f32 findInterpIdx t division, non-inline sections with quota — 3697 cycles (-6%) 2026-03-16 14:12:53 -07:00
MarcelineVQ 502fd9ba04 bone_sse: remove align(1) for naturally-aligned game data — 3641 cycles (was 3744) 2026-03-16 12:43:51 -07:00
MarcelineVQ 88fec305be bone_sse: aggressive inlining + skip identity init + fused rotateByQuaternion — 3744 cycles (was 4040) 2026-03-16 12:41:54 -07:00
MarcelineVQ 4ed5495219 bench: dual BASELINE/SSE with parity check — 2177 vs 2112 cycles, PASS 2026-03-16 12:10:51 -07:00
MarcelineVQ 91444f8104 bone_sse: incremental bone pointer advancement, V4 copyMat4 2026-03-16 11:18:50 -07:00
MarcelineVQ 3a522fd1e4 bone_sse: cache frame_ctr, pass to all section functions — eliminates ~29 redundant reads 2026-03-16 11:14:56 -07:00
MarcelineVQ c80e63522b bone_sse: @mulAdd (FMA) for all lerp/blend paths across interp and section functions 2026-03-16 11:06:59 -07:00
MarcelineVQ c36352f4ff bone_sse: @mulAdd (FMA) for lerpVec3, applyTranslation, vec3SqMag 2026-03-16 11:00:01 -07:00
MarcelineVQ a7e2f978e2 bone_sse: simplify rotateByQuaternion to use matMul4x4 V4+FMA path 2026-03-16 10:56:48 -07:00
MarcelineVQ 5d41cd03a3 bone_sse: V4 FMA matmul, enable SSE4.1+FMA+AVX target 2026-03-16 10:54:25 -07:00
MarcelineVQ 08d1992eaf bone_sse: replace IsParticleBufferEmpty with pure Zig — 1 game call remains (atexit) 2026-03-16 10:45:06 -07:00
MarcelineVQ 1987dc5622 bone_sse: pure Zig — only 2 game calls remain (atexit init + particle buffer check)
Replaced all remaining game function calls:
- findInterpIdx (0x713D50): full temporal-coherence search reimplementation
- interpAnimKF in boneKeyframeLoop: reuses existing pure Zig version
- applyTranslation/rotateByQuaternion/scaleMatrix3x3 in boneKeyframeLoop
- getInterpolatedFloat (0x71AF20): replaced with interpFloatTrack (identical)
- extractByte (0x71AE90): findInterpIdx + direct byte read
- Child recursion: direct call to transformImpl_SSE instead of 0x714260 hook

Only 2 game calls remain (cannot be replaced):
- 0x409AEF: one-time atexit registration in boneKeyframeLoop
- 0x7B5F60: IsParticleBufferEmpty (reads game particle state)
2026-03-16 10:38:34 -07:00
MarcelineVQ 16c913b9b7 bone_sse: pure Zig findInterpIdx — last major game call in interpolation path 2026-03-16 10:33:40 -07:00
MarcelineVQ 800adbe185 bone_sse: pure Zig bone loop — interpAnimKF, buildRot, scaleMat, matMul all replaced 2026-03-16 10:27:33 -07:00
MarcelineVQ 0aab0d3678 bone_sse: replace getIndexOffset/setShortValue with direct ri16 reads 2026-03-16 10:22:06 -07:00
MarcelineVQ 8736e0b3da bone_sse: clean A/B testing, remove diagnostic comparison code
Remove the double-call REF/SSE bone output comparison diagnostic.
Clean detour: REF baseline, SSE custom, simple toggle.
2026-03-16 10:18:35 -07:00
MarcelineVQ 008db88403 bone_sse: replace matMul/ftol/vec3SqMag with pure Zig, fix visual artifacts
Root cause of billboard visual artifacts: callVec3SqMag used inline asm
to call game's x87 vec3SqMag (0x4549F0) with fstps to capture ST0.
With SSE2 codegen, the x87/SSE state interaction caused corrupted float
values in billboard bone matrices (mat[0][2] wildly wrong).

Replaced with pure Zig: x*x + y*y + z*z — no x87, no inline asm.

Also replaced:
- matMul (0x74A7C0): pure Zig f64 scalar matmul, no alignment needs
- callFtol (0x40A2B0): f64 intermediate + @intFromFloat (cvttsd2si)

Architecture: bone_sse.zig is REF code compiled with SSE2, called as
cdecl from thiscall wrapper in transform44.zig (cross-object to prevent
LLVM inlining AND ESP alignment into thiscall frame).

A/B: other hooks gated behind AB_OTHER_HOOKS=false for isolated testing.
2026-03-16 10:10:02 -07:00
MarcelineVQ 252853dd8f bone_sse: fix findInterpIdx from assembly, runtime constants, interpAnimKF stride
Assembly-verified fixes from t44_helpers_asm.txt:

findInterpIdx (0x713D50):
- Range format is [start, last] not [start, count] (DEC EDI pattern)
- Backward scan entry: delta >= 0xFFFFFE0C not > (JC = unsigned below)
- 500-tick threshold (0x1F4) for forward/backward vs binary search
- GS check is CMP AX,0xFFFF (word compare), not >= 0
- t computation: FILD qword (i64 numer) / FIDIV dword (i32 denom)

interpAnimKF (0x713EA0):
- Keyframe stride is 16 bytes (SHL EAX,0x4), NOT 8 (CompQuat)
- Values are raw floats, no short-to-float conversion needed

Runtime constants — all now read from game memory:
- 0x80297C (3.0) and 0x802990 (6.0) for Hermite/Bezier basis
- 0x80C5C8 for billboard squared magnitude threshold
- 0x811610 and 0x8029D4 already read at runtime

Wrapper pattern: thiscall export delegates to normal fn for AVX alignment.
2026-03-15 18:39:27 -07:00
MarcelineVQ 3a031803ef bone_sse: pure Zig SSE/FMA reimplementation, zero game function calls
Replace bone_sse.zig with a complete pure Zig implementation compiled
with SSE4.1 + FMA + AVX. All 18 game function calls replaced:

- findInterpIdx (0x713D50): temporal-coherence keyframe search
- interpAnimKF (0x713EA0): CompQuat lerp for rotation keyframes
- extractByte (0x71AE90): byte keyframe extraction
- getInterpolatedFloat (0x71AF20): float track with direct blend read
- callFtol (0x40A2B0): @intFromFloat replaces x87 __ftol
- callVec3SqMag (0x4549F0): inline FMA dot product
- callGetIndexOffset/callSetShortValue (0x71AFF0/0x71B010): direct ri16
- matMul (0x74A7C0): V4 FMA matmul (broadcast + 3 @mulAdd per row)
- buildRotFn (0x74B6B5): inline quat→matrix
- rotateQuat (0x7BDDB0): quat→matrix then FMA matmul
- scaleMat (0x7BDCA0): inline scale from vec3 ptr
- applyTrans (0x7BDC40): inline FMA dot product translation

Only 2 game calls remain:
- 0x409AEF: one-time atexit init (boneKeyframeLoop)
- 0x7B5F60: IsParticleBufferEmpty (reads game particle state)

Child recursion calls transformMatrix4x4_SSE directly instead of
going through the hook at 0x714260.

Detour cleaned up: REF is baseline, SSE activates via ab_use_custom
toggle. Diagnostic/bisect/FPU-comparison scaffolding removed.
build.zig: bone_sse gets dedicated target with sse4_1+fma+avx features.
2026-03-15 18:16:07 -07:00
MarcelineVQ 8fb1ece2b5 bone_sse_ref: fix world entry crash — 6 bugs found via full asm stepthrough
Full 5317-instruction walkthrough of t44_full_asm.txt vs bone_sse_reference.zig.

Crash fix (Issues 1-2): When anim_start >= anim_end in the looping animation
path, assembly always writes prim_time/sec_time = anim_start as fallback.
REF skipped the write, leaving garbage in bone_rt timing fields. On newly
loaded world SceneObjects this propagated through findInterpIdx → extractByte
→ ACCESS_VIOLATION at 0x71AEBC with ECX=0x7FFFFFFF (self-reinforcing bad
cached index).

Time clamp fix (Issues 3-4): Clamped-not-passed animation path now clamps
cur_time to sec_start when sec_start > cur_time, matching assembly at
0x7145EB/0x71474B.

Crossfade fix (Issues 5-6): interpVec3Track36 and interpFloatTrack12 had
'else return' for unknown interp modes. Assembly's JNZ skips primary interp
but falls through to crossfade check. Changed to 'else {}' fallthrough.

Also includes prior uncommitted fixes: particle crossfade blend_weight
restoration, bw>0→bw!=0, attach_count==0 early-return removal.
2026-03-15 17:38:14 -07:00
MarcelineVQ 6f8b5ea0e7 bone_sse_ref: disable particle crossfade, fix cross product z, guard cleanup
- Particle interpVec3Track/interpFloatTrack: pass 0.0 blend_weight to
  disable crossfade, matching original which has no crossfade in particle
  sections (only bone loop and boneKeyframeLoop have crossfade).
- interpFloatTrack: add explicit blend_weight parameter instead of
  reading from bone_rt internally, allowing callers to control crossfade.
- Revert extractByte guard (was added then removed during investigation).

Known crash: extractByte (0x71AE90) crashes at 0x71AEBC with idx=0x7FFFFFFF
on world entry. findInterpIdx reads output[0] as cached search position;
if hierarchy buffer contains stale 0x7FFFFFFF, search overflows and
self-reinforces. Investigation ongoing — REF's attachment section matches
original assembly instruction-for-instruction.
2026-03-15 16:44:43 -07:00
MarcelineVQ 582b0dc132 bone_sse_ref: add crossfade blending, word animation section, fix cross product z
- texAnimLoop alpha: add crossfade blend for mode != 0 (mode 0 skips
  crossfade per original assembly JMP at 0x715B49). Fix alpha output
  base to output+0x30 matching original ESI.
- colorAnimLoop: add crossfade blend, same mode 0 skip pattern.
  Mode 0 uses direct short->float copy matching original.
- New wordAnimLoop: implements model_hdr+0x6C/0x70 word animation
  section (assembly 0x715E46-0x715F25). Word copy with crossfade,
  no float blending. Data stride 0x1C, output stride 0x20.
- Fix billboard cross product z-component for types 0x10/0x20:
  was +cross.z, should be -cross.z (r0y*r1x - r0x*r1y).
- Extract shortInterpToFloat helper shared by alpha/color crossfade.

Known: particle emitter crash (pre-existing, idx=0x7FFFFFFF in
secondary findInterpIdx) — exposed by corrected colorAnimLoop count.
2026-03-15 16:21:33 -07:00
MarcelineVQ 36ce1a05ce bone_sse_ref: fix M2 black screen — 5 bugs found via asm comparison
Assembly-level comparison of compiled REF against original 0x714260 revealed:

1. Billboard cross product sign error (types 0x10/0x20): computed +cross
   instead of -cross for components 0/1, corrupting billboard bone matrices
2. colorAnimLoop wrong count field: read model_hdr+0x6C instead of +0x64
3. colorAnimLoop wrong gate offset: checked anim_data+0x04 instead of +0x0C
4. Timestamp delta guard inverted: REF guarded on cur_ts!=0 and always
   wrote to this+0x4C; original guards on this+0x4C!=0 first and never
   seeds the field (something else initializes it)
5. Section 5 emitter_ctx cached instead of re-read after matMul call

Also: build REF with x87-only target (subtract SSE/SSE2 features) to
match original's FLD/FMUL/FSTP codegen, and use callVec3SqMag for all
magnitude computations instead of inline SSE math.

Remaining known issues (not yet fixed):
- texAnimLoop alpha track missing crossfade blend
- colorAnimLoop missing crossfade blend
- Missing word animation section (model_hdr+0x6C/0x70)
- Bisect infrastructure and diagnostic code still present (test scaffolding)
2026-03-15 15:04:05 -07:00
MarcelineVQ 89e49cedf2 bone_sse_ref: replace reimplemented game funcs with actual calls, fix runtime constants
- Replace all reimplemented game functions with actual game calls:
  vec3_sqmag (0x4549F0), __ftol (0x40A2B0), getIndexOffset (0x71AFF0),
  setShortValue (0x71B010) — matching assembly exactly
- Fix 3 wrong hardcoded constants that differ at runtime from Ghidra static values:
  SHORT_TO_FLOAT: 0x38000000→0x38000100 (1/32767 not 1/32768)
  BILLBOARD_EPSILON: 0x3727c5ac→0x34800000
  HERMITE_5: 5.0→6.0
  All now read from game memory at runtime
- Fix timestamp delta guard (this+0x4C): was guarding on stored value,
  assembly guards on anim_ctx pointer — prevents first-frame initialization
- Change REF calling convention to thiscall matching original
- Add comprehensive memory comparison diagnostic (original vs REF)
- Disable interpKfDetour hook (was pure passthrough)
2026-03-15 13:20:05 -07:00