Add module control API, bigcursor with fractional scaling, shared mutex

Module control API:
- Three cdecl exports: WeirdUtils_IsModuleActive, WeirdUtils_DisableModule, WeirdUtils_DisableAll
- C header include/weirdutils_api.h for runtime DLL discovery
- All modules gain pub module_name, isActive(), shared mutex via src/mutex.zig
- build.zig refactored to module_list array of ModuleDesc

Bigcursor:
- D3D9 vtable hooks on SetCursorProperties/ShowCursor
- Scale2x/Scale3x pixel-art upscalers in pure Zig (replaces 12K-line hqx C)
- Bilinear resampler for fractional scales (1.0-4.0, default 1.2)
- Win32 cursor via CreateIconIndirect bypasses D3D9 32x32 limit
- FNV-1a hash cache for ~16 cursor bitmaps
- Lua API: SetCursorScale(n) / GetCursorScale()
- CVar cursorScale for Config.wtf persistence (tenths)

Other:
- Add dpslog module stub
- Docs and README updates
This commit is contained in:
MarcelineVQ
2026-03-06 13:27:35 -08:00
parent 0783f3b0dd
commit 2aaa089c76
44 changed files with 1904 additions and 1022 deletions
@@ -56,7 +56,7 @@ Also review the CLI Ghidra skill:
## Execution Plan
## Phase 1 — Identify canonical native registration path
## Phase 1 - Identify canonical native registration path
### Task 1.1: Recover true iterator and linkage semantics
- Disassemble around `0x00683F80` (and nearby helpers) in detail.
@@ -80,7 +80,7 @@ Also review the CLI Ghidra skill:
---
## Phase 2 — Build proper registration wrapper in markers module
## Phase 2 - Build proper registration wrapper in markers module
### Task 2.1: Implement native-path wrapper
- Add a wrapper that uses discovered native creator/registration APIs.
@@ -101,7 +101,7 @@ Also review the CLI Ghidra skill:
---
## Phase 3 — Validation protocol (must pass)
## Phase 3 - Validation protocol (must pass)
Run and capture logs for:
+13 -13
View File
@@ -1,12 +1,12 @@
# Outline Rendering Design — Requirements & Approach Analysis
# Outline Rendering Design - Requirements & Approach Analysis
## Date: 2026-02-24
## Requirements (Exact)
1. **Walls, terrain, game objects (doors, pillars)**: MUST occlude outlines
2. **Players and NPCs**: MUST NOT occlude outlines — outlines are specifically for making enemies easier to see in combat, so they must always show over other units
3. **Dead friendly players**: Outlines visible through EVERYTHING (including walls) — for finding corpses to resurrect
2. **Players and NPCs**: MUST NOT occlude outlines - outlines are specifically for making enemies easier to see in combat, so they must always show over other units
3. **Dead friendly players**: Outlines visible through EVERYTHING (including walls) - for finding corpses to resurrect
Any architectural approach is acceptable. The existing 3-pass stencil code is proof-of-concept, not a constraint. Efficiency and cleanliness matter more than preserving existing code.
@@ -18,7 +18,7 @@ The depth buffer doesn't distinguish "wall pixel" from "player pixel." After a f
- If you skip depth testing → players don't occlude (✓) but walls also don't occlude (✗)
- If you composite in EndScene after all rendering → outlines are on top of everything, including walls (✗)
**You need depth information from BEFORE players/NPCs have rendered, but AFTER terrain/WMOs have rendered.** This is only available at a specific point during the frame — when M2 model batches begin processing.
**You need depth information from BEFORE players/NPCs have rendered, but AFTER terrain/WMOs have rendered.** This is only available at a specific point during the frame - when M2 model batches begin processing.
## Approach Analysis
@@ -46,7 +46,7 @@ The depth buffer doesn't distinguish "wall pixel" from "player pixel." After a f
- Without batch reordering, player depth may already be present → players occlude mask pixels
- With batch reordering, the mask captures the right silhouette, but compositing in EndScene draws OVER walls that rendered later
The composite happens at the wrong time — after everything has rendered, including walls that should occlude.
The composite happens at the wrong time - after everything has rendered, including walls that should occlude.
**Could work if combined with batch reordering** and a depth-aware composite, but this reintroduces the batch reordering requirement and adds render target + fullscreen quad overhead on top.
@@ -61,16 +61,16 @@ Copy the depth buffer at the start of M2 rendering (after terrain/WMOs, before p
Same core idea as Approach A but optimized:
1. **Pass 1 (stencil + outline):** Set stencil to mark body. Draw the expanded outline geometry with custom VS, using stencil to exclude body pixels. Write `STENCIL_BIT_OUTLINE` where outline draws.
2. **Pass 2 (normal):** Draw model normally (game's original state). This is the draw that would have happened anyway — just done after the outline.
2. **Pass 2 (normal):** Draw model normally (game's original state). This is the draw that would have happened anyway - just done after the outline.
Wait — this is still 3 DIP calls (stencil mark needs the normal geometry first). The passes can't easily be collapsed because the stencil mark (body silhouette) must exist before the outline can exclude it.
Wait - this is still 3 DIP calls (stencil mark needs the normal geometry first). The passes can't easily be collapsed because the stencil mark (body silhouette) must exist before the outline can exclude it.
### Approach E: Inverted Hull with Game's Depth (Refined Stencil)
Same 3-pass stencil, but:
- **Use the game's own VS for pass 1 and 3** (no custom shader needed for stencil mark + normal draw)
- **Custom VS only for pass 2** (outline expansion) — this is the only pass that needs modified geometry
- **Minimize state changes** — only save/restore what we actually modify
- **Custom VS only for pass 2** (outline expansion) - this is the only pass that needs modified geometry
- **Minimize state changes** - only save/restore what we actually modify
- **Skip vertex declaration swap** if the game's declaration is compatible with our VS
This is what the current code already does, just cleaned up.
@@ -109,12 +109,12 @@ The Reset hook releases shaders. Ensure `shaders_attempted` is reset so they're
### 6. Consider vs_3_0 upgrade
The current VS uses vs_2_0. The game supports ps_3_0 (confirmed 0xFFFF0300), so vs_3_0 is available. Benefits: better precision for screen-space normal calculation, no instruction count limit. However vs_2_0 works and is simpler — this is optional.
The current VS uses vs_2_0. The game supports ps_3_0 (confirmed 0xFFFF0300), so vs_3_0 is available. Benefits: better precision for screen-space normal calculation, no instruction count limit. However vs_2_0 works and is simpler - this is optional.
## Implementation Order
1. Update `types.zig` — add any missing D3D9 constants
2. Update `d3d9_hook.zig` — add ZWRITEENABLE, DEPTHBIAS, stride check
3. Verify `model_hook.zig` — batch reordering and stencil flags are correct
1. Update `types.zig` - add any missing D3D9 constants
2. Update `d3d9_hook.zig` - add ZWRITEENABLE, DEPTHBIAS, stride check
3. Verify `model_hook.zig` - batch reordering and stencil flags are correct
4. Build and verify compilation
5. Test (login screen → in-game with live targets)
+24 -24
View File
@@ -1,4 +1,4 @@
# WoW 1.12.1 Calling Conventions — Ghidra Verified
# WoW 1.12.1 Calling Conventions - Ghidra Verified
All conventions verified against WoW.exe 1.12.1 build 5875 via Ghidra decompilation and raw byte analysis.
@@ -6,17 +6,17 @@ All conventions verified against WoW.exe 1.12.1 build 5875 via Ghidra decompilat
| # | Address | Function | Convention | Params | Prologue | RET | Status |
|---|---------|----------|------------|--------|----------|-----|--------|
| 1 | `0x0070b360` | CM2SceneRenderDraw | `__thiscall` | ECX=this, stack: viewMatrix, batchData, batchIndices, batchCount | `55 8B EC 81 EC 80 00 00 00` (9B) | — | CORRECT |
| 2 | `0x00710b90` | CM2Model_ManageRenderListNode | `__thiscall` | ECX=model, stack: addToList | `55 8B EC 8B 45 08` (6B) | — | CORRECT |
| 3 | `0x0070cb30` | CM2Scene_DrawBatchProjected | `__fastcall` | ECX=renderContext | `55 8B EC 83 EC 10` (6B) | — | CORRECT |
| 1 | `0x0070b360` | CM2SceneRenderDraw | `__thiscall` | ECX=this, stack: viewMatrix, batchData, batchIndices, batchCount | `55 8B EC 81 EC 80 00 00 00` (9B) | - | CORRECT |
| 2 | `0x00710b90` | CM2Model_ManageRenderListNode | `__thiscall` | ECX=model, stack: addToList | `55 8B EC 8B 45 08` (6B) | - | CORRECT |
| 3 | `0x0070cb30` | CM2Scene_DrawBatchProjected | `__fastcall` | ECX=renderContext | `55 8B EC 83 EC 10` (6B) | - | CORRECT |
## Game Function Wrappers (outline/wow.zig)
| # | Address | Function | Convention | Params | Prologue | RET | Status |
|---|---------|----------|------------|--------|----------|-----|--------|
| 4 | `0x00515970` | Script_UnitGUID | `__fastcall` | ECX=unitIdStr → EAX:EDX (64-bit) | `55 8B EC 51 56 68 90 00 00 00` | — | CORRECT |
| 5 | `0x00464870` | GetObjectByGUID | **`__stdcall`** | **stack: guidLow, guidHigh → EAX** | `55 8B EC 8B 45 08 8B 4D 0C` | **RET 8** | **FIXED** — was incorrectly using `hook.fastcall` |
| 6 | `0x006061E0` | CGUnit_C::UnitReaction | `__thiscall` | ECX=localPlayer, stack: unit → EAX (reaction int) | `53 8B DC 83 EC 08 83 E4 F8` | — | CORRECT |
| 4 | `0x00515970` | Script_UnitGUID | `__fastcall` | ECX=unitIdStr → EAX:EDX (64-bit) | `55 8B EC 51 56 68 90 00 00 00` | - | CORRECT |
| 5 | `0x00464870` | GetObjectByGUID | **`__stdcall`** | **stack: guidLow, guidHigh → EAX** | `55 8B EC 8B 45 08 8B 4D 0C` | **RET 8** | **FIXED** - was incorrectly using `hook.fastcall` |
| 6 | `0x006061E0` | CGUnit_C::UnitReaction | `__thiscall` | ECX=localPlayer, stack: unit → EAX (reaction int) | `53 8B DC 83 EC 08 83 E4 F8` | - | CORRECT |
### GetObjectByGUID Detail
@@ -36,11 +36,11 @@ E8 ... CALL FindObjectByGUID
C2 08 00 RET 8 ; callee cleans 8 bytes
```
The C++ reference declared this as `__fastcall(uint64_t)`. Under MSVC, `uint64_t` (8 bytes) is too large for a single 32-bit register, so `__fastcall` passes it on the stack — making it behave like `__stdcall`. The Zig code split it into two `u32` args and passed them in ECX/EDX via `hook.fastcall`, which was wrong.
The C++ reference declared this as `__fastcall(uint64_t)`. Under MSVC, `uint64_t` (8 bytes) is too large for a single 32-bit register, so `__fastcall` passes it on the stack - making it behave like `__stdcall`. The Zig code split it into two `u32` args and passed them in ECX/EDX via `hook.fastcall`, which was wrong.
The transmog addon (`transmogfix/src/main.zig:134`) and interact module (`weirdutils/src/interact.zig:50`) already had the correct push-to-stack implementation.
## Dead Overlay Functions (not yet ported — for future reference)
## Dead Overlay Functions (not yet ported - for future reference)
| # | Address | Function | Convention | Params | Status |
|---|---------|----------|------------|--------|--------|
@@ -55,23 +55,23 @@ All WoW 1.12.1 Lua C API functions use `__fastcall` with L (lua_State*) in ECX.
| # | Address | Function | Convention | Params | RET | Status |
|---|---------|----------|------------|--------|-----|--------|
| 11 | `0x00704120` | FrameScript::Register | `__fastcall` | ECX=name, EDX=funcAddr | — | CORRECT |
| 11 | `0x00704120` | FrameScript::Register | `__fastcall` | ECX=name, EDX=funcAddr | - | CORRECT |
| 12 | `0x006F3070` | lua_gettop | `__fastcall` | ECX=L → int | RET | CORRECT |
| 13 | `0x006F3080` | lua_settop | `__fastcall` | ECX=L, EDX=index | — | CORRECT |
| 14 | `0x006F3350` | lua_pushvalue | `__fastcall` | ECX=L, EDX=index | — | CORRECT |
| 13 | `0x006F3080` | lua_settop | `__fastcall` | ECX=L, EDX=index | - | CORRECT |
| 14 | `0x006F3350` | lua_pushvalue | `__fastcall` | ECX=L, EDX=index | - | CORRECT |
| 15 | `0x006F3400` | lua_type | `__fastcall` | ECX=L, EDX=index → int | RET | CORRECT |
| 16 | `0x006F3510` | lua_isstring | `__fastcall` | ECX=L, EDX=index → int | RET | CORRECT |
| 17 | `0x006F3690` | lua_tostring | `__fastcall` | ECX=L, EDX=index → char* | — | CORRECT |
| 18 | `0x006F39F0` | lua_pushboolean | `__fastcall` | ECX=L, EDX=bool | — | CORRECT |
| 19 | `0x006F3890` | lua_pushstring | `__fastcall` | ECX=L, EDX=string | — | CORRECT |
| 17 | `0x006F3690` | lua_tostring | `__fastcall` | ECX=L, EDX=index → char* | - | CORRECT |
| 18 | `0x006F39F0` | lua_pushboolean | `__fastcall` | ECX=L, EDX=bool | - | CORRECT |
| 19 | `0x006F3890` | lua_pushstring | `__fastcall` | ECX=L, EDX=string | - | CORRECT |
| 20 | `0x006F3810` | lua_pushnumber | `__fastcall` | ECX=L, stack: f64 (8 bytes) | RET 8 | CORRECT |
| 21 | `0x006F3920` | lua_pushcclosure | `__fastcall` | ECX=L, EDX=func, stack: nupvalues | — | **FIXED** — was 0x6F3B80 (wrong addr) |
| 22 | `0x006F4940` | luaL_error | `__cdecl` | stack: L, fmt, ... | — | CORRECT |
| 23 | `0x006F4DC0` | luaL_openlib | `__fastcall` | ECX=L, EDX=libname, stack: funcs, nup | — | CORRECT |
| 21 | `0x006F3920` | lua_pushcclosure | `__fastcall` | ECX=L, EDX=func, stack: nupvalues | - | **FIXED** - was 0x6F3B80 (wrong addr) |
| 22 | `0x006F4940` | luaL_error | `__cdecl` | stack: L, fmt, ... | - | CORRECT |
| 23 | `0x006F4DC0` | luaL_openlib | `__fastcall` | ECX=L, EDX=libname, stack: funcs, nup | - | CORRECT |
### lua_pushcclosure Detail
Ghidra search found `lua_pushcclosure @ 006f3920`. No function exists at the old address `0x6F3B80` — it falls mid-body of another function. The wrapper was unused (never called from current code) so no crash occurred.
Ghidra search found `lua_pushcclosure @ 006f3920`. No function exists at the old address `0x6F3B80` - it falls mid-body of another function. The wrapper was unused (never called from current code) so no crash occurred.
### lua_pushnumber Detail
@@ -81,14 +81,14 @@ Takes a `double` (8 bytes) which is too large for EDX, so it goes on the stack p
| # | Address | Function | Convention | Params | Prologue | RET | Status |
|---|---------|----------|------------|--------|----------|-----|--------|
| 24 | `0x0042a320` | ValidateFunctionPointer | `__fastcall` | ECX=addr | `55 8B EC 83 EC 40` (6B) | — | CORRECT (empty detour) |
| 24 | `0x0042a320` | ValidateFunctionPointer | `__fastcall` | ECX=addr | `55 8B EC 83 EC 40` (6B) | - | CORRECT (empty detour) |
| 25 | `0x00648620` | LoadFileWithTextureResourceFallback | `__stdcall` | 7 stack params | `55 8B EC 8B 4D 1C` (6B) | RET 0x1C | CORRECT |
| 26 | `0x00490250` | FrameScript_RegisterAllSystemCommands | `void(void)` | none | `56 E8 ...` (6B) | — | CORRECT (fixup at offset 1) |
| 27 | `0x0051F600` | LoadAddonsRecursively | `__fastcall` | ECX=error_handler | `53 8B 1D ...` (7B) | — | CORRECT |
| 26 | `0x00490250` | FrameScript_RegisterAllSystemCommands | `void(void)` | none | `56 E8 ...` (6B) | - | CORRECT (fixup at offset 1) |
| 27 | `0x0051F600` | LoadAddonsRecursively | `__fastcall` | ECX=error_handler | `53 8B 1D ...` (7B) | - | CORRECT |
| 28 | `0x006EDB90` | loadFileListWithIncludes | `__fastcall` | ECX=path, EDX=md5ctx, stack: error_handler | `55 8B EC 6A FF ...` | RET 4 | CORRECT |
| 29 | `0x004B6F70` | LoadUIBindingsFromFile | `__thiscall` | ECX=binding_mgr, stack: path, md5ctx, callback | `55 8B EC 81 EC 1C 04 00 00` | RET 0x0C | CORRECT |
| 30 | `0x0046a400` | GameEngine_MainInitialize | `void(void)` | none | `55 8B EC 83 EC 28` (6B) | — | CORRECT |
| 31 | `0x00490BD0` | World_HandlePlayerLogin | `void(void)` | none | `56 E8 ...` (6B) | — | CORRECT (fixup at offset 1) |
| 30 | `0x0046a400` | GameEngine_MainInitialize | `void(void)` | none | `55 8B EC 83 EC 28` (6B) | - | CORRECT |
| 31 | `0x00490BD0` | World_HandlePlayerLogin | `void(void)` | none | `56 E8 ...` (6B) | - | CORRECT (fixup at offset 1) |
### Note on #31
+10 -10
View File
@@ -8,7 +8,7 @@ Outlines rendered via the JFA pipeline have a "marching ants" pattern of missing
## What Was Ruled Out
### 1. Silhouette RT is clean (CONFIRMED)
Added `DEBUG_SHOW_SILHOUETTE` comptime flag in `d3d9_hook.zig` that skips JFA and composites the raw silhouette RT directly to the backbuffer. The silhouette was solid with no banding — the cached draw replay produces correct geometry. **Stale VB hypothesis is NOT the cause.**
Added `DEBUG_SHOW_SILHOUETTE` comptime flag in `d3d9_hook.zig` that skips JFA and composites the raw silhouette RT directly to the backbuffer. The silhouette was solid with no banding - the cached draw replay produces correct geometry. **Stale VB hypothesis is NOT the cause.**
### 2. JFA sentinel value (1.0, 1.0) → (-1.0, -1.0) (NO IMPROVEMENT)
Changed the JFA init shader sentinel from `(1.0, 1.0)` to `(-1.0, -1.0)` so unflooded pixels can't act as false seeds near the right screen edge. This did NOT fix the marching ants pattern. The sentinel is currently set to `(-1.0, -1.0)` in the code (uncommitted).
@@ -29,7 +29,7 @@ Batch reorder puts outline targets last so depth buffer has full scene geometry
- `DEBUG_SHOW_SILHOUETTE` comptime flag added (currently `false`), with `debug_sil_ps` shader
- Debug shader had a cmp operand inversion bug that was fixed (`c0.w, c0.x` not `c0.x, c0.w`)
## Remaining Investigation — JFA Pipeline Bug
## Remaining Investigation - JFA Pipeline Bug
Since the silhouette is clean, the bug is in the JFA shaders (Phase 2). Possible causes:
@@ -37,25 +37,25 @@ Since the silhouette is clean, the bug is in the JFA shaders (Phase 2). Possible
The propagation shader uses `cmp` (ps_3_0) which tests `>= 0` vs `< 0`. When `new_dist² == best_dist²` (exactly equal), `cmp` picks the OLD seed (`>= 0` branch). This tie-breaking might cause systematic bias where seeds from certain directions are always preferred, creating directional artifacts. Worth testing: swap `cmp` operands to prefer new seed on ties, or add a small epsilon.
### B. dp2add precision on DXVK
`dp2add r2.z, r2, r2, c1.x` computes `r2.x*r2.x + r2.y*r2.y + 0.0`. DXVK translates this to Vulkan — there may be precision differences vs native D3D9 that affect distance comparisons, especially for pixels equidistant from multiple seeds.
`dp2add r2.z, r2, r2, c1.x` computes `r2.x*r2.x + r2.y*r2.y + 0.0`. DXVK translates this to Vulkan - there may be precision differences vs native D3D9 that affect distance comparisons, especially for pixels equidistant from multiple seeds.
### C. Neighbor sampling at texture edges
When `mad r4.xy, offset, step_uv, v0.xy` goes outside [0,1], CLAMP addressing returns the edge texel. This could feed stale/wrong seed UVs into the comparison. Clamping the sample coordinate to valid range before comparison could help.
### D. The JFA algorithm itself may not suit this use case
Consider alternative approaches:
- **Screen-space dilation** (iterative morphological expand of silhouette) — simpler, no distance field needed
- **Gaussian blur difference** — blur silhouette, subtract original, threshold
- **Screen-space dilation** (iterative morphological expand of silhouette) - simpler, no distance field needed
- **Gaussian blur difference** - blur silhouette, subtract original, threshold
- **Sobel/edge detection** on the silhouette RT
- Docs in `/media/storage/projects/zig/weirdutils/docs/` describe these alternatives
- Reference articles: ameye.dev "5 ways to draw an outline", Ben Golus "Quest for Very Wide Outlines"
## Key Files
- `src/outline/d3d9_hook.zig` — D3D9 hooks, JFA pipeline, all shaders
- `src/outline/model_hook.zig` — batch reordering, rendering_outline flag
- `src/outline/tracker.zig` — per-frame model tracking
- `src/outline/types.zig` — D3D9 constants
- `reference/c_overlay/d3d9_hook.cpp` — C reference (uses 3-pass shell extrusion, not JFA)
- `src/outline/d3d9_hook.zig` - D3D9 hooks, JFA pipeline, all shaders
- `src/outline/model_hook.zig` - batch reordering, rendering_outline flag
- `src/outline/tracker.zig` - per-frame model tracking
- `src/outline/types.zig` - D3D9 constants
- `reference/c_overlay/d3d9_hook.cpp` - C reference (uses 3-pass shell extrusion, not JFA)
## Build / Environment
- `zig build` from `/media/storage/projects/zig/weirdutils/`
+13 -13
View File
@@ -14,9 +14,9 @@ Pure Zig reimplementation of the WoW 1.12.1 unit outline system, ported from the
**Prologue verification** (Ghidra, WoW.exe 1.12.1 build 5875):
- `0x0070b360`: `55 8B EC 81 EC 80 00 00 00` — `PUSH EBP; MOV EBP,ESP; SUB ESP,0x80`. Boundaries at +1, +3, +9. The `SUB ESP,0x80` is a 6-byte instruction (81 EC + imm32) spanning offset +3..+9, so 6-byte overwrite is **unsafe** — changed to 9.
- `0x00710b90`: `55 8B EC 8B 45 08` — `PUSH EBP; MOV EBP,ESP; MOV EAX,[EBP+8]`. Boundaries at +1, +3, +6. Clean 6-byte boundary.
- `0x0070cb30`: `55 8B EC 83 EC 10` — `PUSH EBP; MOV EBP,ESP; SUB ESP,0x10`. Boundaries at +1, +3, +6. Clean 6-byte boundary.
- `0x0070b360`: `55 8B EC 81 EC 80 00 00 00` - `PUSH EBP; MOV EBP,ESP; SUB ESP,0x80`. Boundaries at +1, +3, +9. The `SUB ESP,0x80` is a 6-byte instruction (81 EC + imm32) spanning offset +3..+9, so 6-byte overwrite is **unsafe** - changed to 9.
- `0x00710b90`: `55 8B EC 8B 45 08` - `PUSH EBP; MOV EBP,ESP; MOV EAX,[EBP+8]`. Boundaries at +1, +3, +6. Clean 6-byte boundary.
- `0x0070cb30`: `55 8B EC 83 EC 10` - `PUSH EBP; MOV EBP,ESP; SUB ESP,0x10`. Boundaries at +1, +3, +6. Clean 6-byte boundary.
All three use `buildFastcallToCdeclThunk` to bridge to `callconv(.c)` detour functions (since `__thiscall` is `__fastcall` with unused EDX).
@@ -34,9 +34,9 @@ Vtable obtained by creating a temporary `IDirect3DDevice9` via `Direct3DCreate9`
Three-pass stencil approach per outline model:
1. **Pass 1 — Mark body**: Draw original geometry to stencil buffer (bit 0), no colour write.
2. **Pass 2 — Draw outline**: Screen-space vertex shader expands vertices along normals. Stencil test rejects body pixels. Write outline bit 1. Dead players disable depth test (through-wall); targets/raid marks respect depth.
3. **Pass 3 — Normal draw**: Restore all state, draw model normally on top.
1. **Pass 1 - Mark body**: Draw original geometry to stencil buffer (bit 0), no colour write.
2. **Pass 2 - Draw outline**: Screen-space vertex shader expands vertices along normals. Stencil test rejects body pixels. Write outline bit 1. Dead players disable depth test (through-wall); targets/raid marks respect depth.
3. **Pass 3 - Normal draw**: Restore all state, draw model normally on top.
## Vertex Shader
@@ -56,7 +56,7 @@ Pixel shader: `ps_3_0`, outputs solid colour from `c0`.
| `+0xC0` | Local player GUID (from ObjMgr) |
| `0x00B71368` | Raid target GUID array (8 × 8 bytes) |
| `0x515970` | `UnitGUID(__fastcall, string_ECX→EAX:EDX)` |
| `0x464870` | `GetObjectByGUID(__stdcall, lo_stack, hi_stack→EAX)` — NOT fastcall! |
| `0x464870` | `GetObjectByGUID(__stdcall, lo_stack, hi_stack→EAX)` - NOT fastcall! |
| `0x6061E0` | `UnitReaction(__thiscall, player_ECX, unit_stack→int)` |
| model+`0x28` | Direct owner object pointer |
| model+`0x3C0` | Callback owner object pointer |
@@ -64,9 +64,9 @@ Pixel shader: `ps_3_0`, outputs solid colour from `c0`.
## Category Priority
1. **Target** (golden amber `#FFC800`) — current target, 2.25px outline
2. **Raid-marked** (per-icon colour) — units with raid icons 1-8, 1.5px
3. **Dead player** (cyan `#00FFFF`) — deceased friendly players, 2.5px, through walls
1. **Target** (golden amber `#FFC800`) - current target, 2.25px outline
2. **Raid-marked** (per-icon colour) - units with raid icons 1-8, 1.5px
3. **Dead player** (cyan `#00FFFF`) - deceased friendly players, 2.5px, through walls
## Per-frame Flow
@@ -76,8 +76,8 @@ Pixel shader: `ps_3_0`, outputs solid colour from `c0`.
## Differences from Reference
- No `std::unordered_set`/`std::unordered_map` — fixed arrays with linear search (max 64 dead GUIDs, 8 raid marks, 256 outline models).
- No `CriticalSection` — all hooks run on the main WoW thread; no synchronisation needed.
- No MinHook — uses the project's existing `libs/hook` inline hook library.
- No `std::unordered_set`/`std::unordered_map` - fixed arrays with linear search (max 64 dead GUIDs, 8 raid marks, 256 outline models).
- No `CriticalSection` - all hooks run on the main WoW thread; no synchronisation needed.
- No MinHook - uses the project's existing `libs/hook` inline hook library.
- Object scanning is frame-based (EndScene), not event-driven (no separate Idris runtime thread).
- Shader loading uses dynamic `d3dx9_43.dll` lookup; gracefully disabled if absent.
+33 -33
View File
@@ -36,7 +36,7 @@ Uses **depth-buffer + Sobel edge detection** tightly integrated into the renderi
- During the skinned mesh rendering pass, shaders write **scaled depth** to a secondary buffer via **Multiple Render Targets (MRT)**.
- Outlines are produced by running a **Sobel filter** on that scaled depth buffer. The Sobel filter finds discontinuities in depth corresponding to silhouette edges.
- The detected edge is rendered back over the skinned mesh — done **per-mesh individually**, not as a single full-screen post-process.
- The detected edge is rendered back over the skinned mesh - done **per-mesh individually**, not as a single full-screen post-process.
- For GPUs that do not support MRT, there is a **fallback using stencil buffers**.
- Rendering order places outlines as a dedicated stage between skinned meshes and grass/water in a 13-stage pipeline.
@@ -48,13 +48,13 @@ Sources:
### Valve Source Engine (Left 4 Dead / DOTA 2 / TF2)
Uses the **"L4D Glow Effect"** — a **stencil + render-to-texture + blur** approach. Used across Left 4 Dead, TF2, CS:GO, and DOTA 2 (pre-Source 2):
Uses the **"L4D Glow Effect"** - a **stencil + render-to-texture + blur** approach. Used across Left 4 Dead, TF2, CS:GO, and DOTA 2 (pre-Source 2):
1. **Stencil pass**: Draw the entity onto the Stencil Buffer. Creates a "cutout" mask of the entity's silhouette.
2. **Color pass**: Draw the entity with the desired glow color (flat/constant color) onto a separate Render Target ("GlowBuff1").
3. **Blur + composite**: Blur GlowBuff1 (using a second RT "GlowBuff2" for ping-pong blur passes), then render the blurred result to the screen **while respecting the stencil buffer**. The stencil test ensures only the blurred pixels that extend beyond the entity's silhouette are visible, producing a halo/outline effect.
The stencil cutout is the key innovation — it prevents the glow color from appearing inside the character, so you only see the outline fringe.
The stencil cutout is the key innovation - it prevents the glow color from appearing inside the character, so you only see the outline fringe.
Sources:
- https://developer.valvesoftware.com/wiki/L4D_Glow_Effect
@@ -94,10 +94,10 @@ Sources:
**How it works:**
Render each target's silhouette as flat color to an offscreen render target (depth-tested against terrain for alive, no depth for dead). Then run a pixel shader that samples an NxN neighborhood — if any sample is "on", the pixel is outline. Subtract the original mask to get just the ring. Composite over backbuffer.
Render each target's silhouette as flat color to an offscreen render target (depth-tested against terrain for alive, no depth for dead). Then run a pixel shader that samples an NxN neighborhood - if any sample is "on", the pixel is outline. Subtract the original mask to get just the ring. Composite over backbuffer.
```hlsl
// SM3.0 pixel shader — fixed-size box dilation
// SM3.0 pixel shader - fixed-size box dilation
sampler2D SilhouetteTex;
float2 TexelSize; // (1.0/screenW, 1.0/screenH)
@@ -131,7 +131,7 @@ if (length(float2(x, y)) <= OutlineRadius) {
**Cons:**
- Square corners at large radii (box kernel artifact)
- Cost grows as O(N^2) with outline width — impractical beyond ~8px
- Cost grows as O(N^2) with outline width - impractical beyond ~8px
- Needs 2-3 render targets
**Performance:** 49 texture samples per pixel at 1024x768 = ~38M samples. On modern hardware: effectively free (<1ms). On 2004-era hardware: 2-4ms.
@@ -144,7 +144,7 @@ if (length(float2(x, y)) <= OutlineRadius) {
The JFA (Rong & Tan, 2006) computes an approximate 2D distance transform on the GPU using O(log N) pixel shader passes. This is the foundation of high-quality screen-space outlines in modern games.
**Step 1 — Seed initialization:**
**Step 1 - Seed initialization:**
Render unit silhouettes into a binary mask. An init shader reads this mask: "on" pixels output their own UV coordinates, "off" pixels get a sentinel value (e.g., `(9999, 9999)`). Output format: RG16F (two channels for x,y coordinates).
```hlsl
@@ -159,7 +159,7 @@ float4 SeedInitPS(float2 uv : TEXCOORD0) : COLOR0
}
```
**Step 2 — JFA propagation (iterative):**
**Step 2 - JFA propagation (iterative):**
Execute `ceil(log2(maxOutlineRadius))` passes. For pass k, step size = `2^(N-k-1)` (starts large, halves each pass). Each pixel samples itself and 8 compass neighbors at the step offset (9 total samples in a 3x3 grid with large spacing). Keep the seed coordinate nearest to the current pixel. Ping-pong between two render targets.
```hlsl
@@ -190,7 +190,7 @@ float4 JFAPassPS(float2 uv : TEXCOORD0) : COLOR0
}
```
**Step 3 — Distance readout and outline generation:**
**Step 3 - Distance readout and outline generation:**
After all passes, each texel holds the UV of the nearest seed. Convert to pixel-space distance and threshold:
```hlsl
@@ -210,12 +210,12 @@ float4 OutlinePS(float2 uv : TEXCOORD0) : COLOR0
}
```
**D3D9/SM3.0 compatibility:** Fully compatible. Each pass is a simple pixel shader with 9 texture samples and simple arithmetic. No gather, no integer bitops, no geometry/compute shaders required. Ping-pong between two textures is standard D3D9. The only requirement is that D3D9 does not allow reading and writing the same surface — alternate between two textures each pass.
**D3D9/SM3.0 compatibility:** Fully compatible. Each pass is a simple pixel shader with 9 texture samples and simple arithmetic. No gather, no integer bitops, no geometry/compute shaders required. Ping-pong between two textures is standard D3D9. The only requirement is that D3D9 does not allow reading and writing the same surface - alternate between two textures each pass.
**Pros:**
- Exact circular distance field — perfectly round outlines at any width
- Exact circular distance field - perfectly round outlines at any width
- Anti-aliasable (smoothstep on the distance)
- Cost is O(log2(N)) passes — a 32px outline costs only 5 passes
- Cost is O(log2(N)) passes - a 32px outline costs only 5 passes
- Enables soft glow, pulsing, gradient effects for free (just change threshold function)
- Nothing requires anything beyond SM2.0
@@ -227,7 +227,7 @@ float4 OutlinePS(float2 uv : TEXCOORD0) : COLOR0
**Performance:** 10 fullscreen passes at 9 samples each = 90M samples at 1024x768. On modern hardware: sub-millisecond. Can run at half resolution (512x384) to halve cost with minimal quality loss for outlines up to 5-6px.
**Occlusion:** Same as dilation — mask generation is independent of outline generation.
**Occlusion:** Same as dilation - mask generation is independent of outline generation.
**JFA quality:** Approximation error bounded at sqrt(2)/2 pixels at jump step boundaries. For outlines up to ~20px, visually imperceptible. Results are smooth, rotationally symmetric, and anti-aliasable.
@@ -247,7 +247,7 @@ Sources:
**How it works:**
Render each target with a unique ID value into an R8 render target (depth-tested). Run a 3x3 Sobel filter — pixels where neighboring IDs differ are edges.
Render each target with a unique ID value into an R8 render target (depth-tested). Run a 3x3 Sobel filter - pixels where neighboring IDs differ are edges.
```hlsl
float4 SobelEdgePS(float2 uv : TEXCOORD0) : COLOR0
@@ -276,7 +276,7 @@ float4 SobelEdgePS(float2 uv : TEXCOORD0) : COLOR0
- No variable-thickness artifacts
**Cons:**
- Produces only 1-2px outlines — can't thicken without adding dilation anyway
- Produces only 1-2px outlines - can't thicken without adding dilation anyway
- Detects unit-to-unit boundaries too (unwanted internal edges between overlapping characters)
- Alone, not sufficient for controllable-width outlines
@@ -291,7 +291,7 @@ float4 SobelEdgePS(float2 uv : TEXCOORD0) : COLOR0
Render silhouette to RT. Separable Gaussian blur (H pass + V pass). Subtract original from blurred → outline. Can downsample to 1/4 res for performance (like retail WoW does for bloom).
```hlsl
// Gaussian blur 5-tap (separable — run horizontal then vertical)
// Gaussian blur 5-tap (separable - run horizontal then vertical)
float weights[5] = {0.0625, 0.25, 0.375, 0.25, 0.0625};
float4 GaussianBlurPS(float2 uv : TEXCOORD0) : COLOR0
@@ -320,13 +320,13 @@ float4 OutlineExtractPS(float2 uv : TEXCOORD0) : COLOR0
- Downsampling to 1/4 res makes a 4px kernel act like a 16px outline
**Cons:**
- Soft/gradient edges, not crisp — looks like a glow, not a hard outline
- Soft/gradient edges, not crisp - looks like a glow, not a hard outline
- Width control is imprecise (tied to blur sigma)
- Can't produce a hard-edged outline without thresholding (which re-introduces aliasing)
**Performance:** 2 fullscreen passes with 5 samples each = 10 samples total. Very cheap.
**Occlusion:** Same as dilation — mask generation is independent.
**Occlusion:** Same as dilation - mask generation is independent.
### 5. Normal Extrusion (Current Approach, Refined)
@@ -371,12 +371,12 @@ This makes outline thickness uniform in screen pixels at any depth.
All screen-space techniques (1-4) share a two-phase architecture that differs fundamentally from the current normal-extrusion approach:
**Phase A — Silhouette mask generation (in DIP hook, per-unit)**
**Phase A - Silhouette mask generation (in DIP hook, per-unit)**
- For alive targets: render unit geometry with depth test ON against scene depth → writes to RT_Silhouette
- For dead targets: render with depth test OFF → writes to RT_Dead
- This is where the 3 occlusion requirements are enforced
**Phase B — Outline generation + composite (in EndScene, once per frame)**
**Phase B - Outline generation + composite (in EndScene, once per frame)**
- Dilate/JFA/blur the mask → extract outline ring → alpha-blend over backbuffer
- This is purely 2D, knows nothing about depth
@@ -397,9 +397,9 @@ Requirement 2 (other units don't occlude outlines) is the hardest to satisfy. It
**JFA with batch reordering for Req 2.**
Rationale:
- Batch reordering already solves Req 2 without needing a secondary depth buffer — outline targets render when only terrain depth exists
- Batch reordering already solves Req 2 without needing a secondary depth buffer - outline targets render when only terrain depth exists
- JFA gives the best outline quality (uniform, circular, anti-aliased, any width) at O(log2(N)) cost
- The outline width of 2-3px only needs ~2 JFA passes — nearly free
- The outline width of 2-3px only needs ~2 JFA passes - nearly free
- Glow/pulse effects come for free if desired
- Everything is SM2.0 compatible, let alone SM3.0
- The silhouette mask pass replaces the current pass 1+2 (stencil body + normal extrusion) with a simpler "render flat color to RT"
@@ -409,9 +409,9 @@ Rationale:
**Resources to create at device creation/reset:**
```
RT_Silhouette: RGBA8, screen size — mask for all outline targets
RT_JFA_A: RG16F, screen size — JFA ping-pong buffer A
RT_JFA_B: RG16F, screen size — JFA ping-pong buffer B
RT_Silhouette: RGBA8, screen size - mask for all outline targets
RT_JFA_A: RG16F, screen size - JFA ping-pong buffer A
RT_JFA_B: RG16F, screen size - JFA ping-pong buffer B
```
**Hook intercept points:**
@@ -454,15 +454,15 @@ Restore all after outline composite.
## Additional References
- "Inking the Cube" (GPU Gems 1, Chapter 11, Everitt) — screen-space dilation
- "Advanced Techniques in Real-Time Rendering" (GDC 2011, de Carpentier) — screen-space outlines
- "Post-Processing Effects in Games" (GDC 2013, Wihlidal) — Sobel ID-buffer approach
- "Inking the Cube" (GPU Gems 1, Chapter 11, Everitt) - screen-space dilation
- "Advanced Techniques in Real-Time Rendering" (GDC 2011, de Carpentier) - screen-space outlines
- "Post-Processing Effects in Games" (GDC 2013, Wihlidal) - Sobel ID-buffer approach
- Unreal Engine 4 custom depth/stencil outline documentation
- https://ameye.dev/notes/rendering-outlines/ — "5 Ways to Draw an Outline"
- https://linework.ameye.dev/soft-outline/ — soft outline documentation
- https://ameye.dev/notes/rendering-outlines/ - "5 Ways to Draw an Outline"
- https://linework.ameye.dev/soft-outline/ - soft outline documentation
- https://www.codeproject.com/Articles/128527/Stencil-Buffer-Glows-Part-1
- https://www.codeproject.com/Articles/156323/Stencil-Buffer-Glows-Part-2
- https://www.tomlooman.com/unreal-engine-soft-outline/
- https://aras-p.info/texts/D3D9GPUHacks.html — D3D9 GPU hacks reference
- https://ameye.dev/notes/edge-detection-outlines/ — edge detection outlines
- https://www.videopoetics.com/tutorials/pixel-perfect-outline-shaders-unity/ — pixel-perfect outlines
- https://aras-p.info/texts/D3D9GPUHacks.html - D3D9 GPU hacks reference
- https://ameye.dev/notes/edge-detection-outlines/ - edge detection outlines
- https://www.videopoetics.com/tutorials/pixel-perfect-outline-shaders-unity/ - pixel-perfect outlines
+2 -2
View File
@@ -64,7 +64,7 @@ Verified 2026-02-23 using Ghidra MCP against WoW.exe 1.12.1 (build 5875, 4,907,0
### Method
Raw bytes read from each hook target address via `get_bytes`. Instruction boundaries decoded to confirm the `prologue_size` parameter passed to `hook.prepare()` lands on a clean instruction boundary. A mid-instruction cut would corrupt the trampoline — the copied bytes would decode as a different instruction when followed by the trampoline's JMP.
Raw bytes read from each hook target address via `get_bytes`. Instruction boundaries decoded to confirm the `prologue_size` parameter passed to `hook.prepare()` lands on a clean instruction boundary. A mid-instruction cut would corrupt the trampoline - the copied bytes would decode as a different instruction when followed by the trampoline's JMP.
### Results
@@ -81,4 +81,4 @@ Raw bytes read from each hook target address via `get_bytes`. Instruction bounda
### No other changes needed
- All three prologues contain only register/memory instructions (no `E8 CALL` or `E9 JMP`), so `rel32_fixups` remains `&.{}`.
- The 4-byte NOP padding (bytes 5-8) at `0x0070b360` after the 5-byte `E9 JMP` is harmless — it is never executed (execution jumps to the detour thunk).
- The 4-byte NOP padding (bytes 5-8) at `0x0070b360` after the 5-byte `E9 JMP` is harmless - it is never executed (execution jumps to the detour thunk).