BLESS

system

t--:--:--

grid--

status

t+0.000shapevanilla: 1 generate + 90 chain dispatches. ours: 1 generate + 1
t+0.001costchain 0.40 to 0.42 ms → 0.14 ms, in the seat's synthetic harness
t+0.002exactin game: 8 of 8 checks bit-identical (lever-verify-1)
t+0.003framec24, whiterun ultra 1080p: 193.0 → 201.0 fps, +4.1%
t+0.004openunder community shaders it likely never engages. not demonstrated

scope

blessed-skyrim is my dxvk fork tuned for skyrim se, i wanna yank one tech out of it. the longer writeup is in past native. this is just looking at one of the mechanisms: collapsed compute dispatch

the tech qualifies as image exact so it gets to run on both profiles, strict included; strict does not permit temporal reuse and must match vanilla precisely.

the material came from notes in blessed-dxvk (blessed-notes/vol-collapse-notes.md), from campaign c24 in the tuning report, and from the audits of rounds six and eight.

what does skyrim even do every frame…

…it’s visible in a frame dump of whiterun (whiterun-dump-1, frame 1920). one generate dispatch opens shader ab674eb1, running 10x6x90 thread groups. it fills two volumes; a and b, each R16_FLOAT at 320x192x90.

chain follows: 90 dispatches of shader 1c4ebb62 (10x6x1 groups apiece) alternating constantly. dispatch reads a and writes b, then the next reads b and writes a.. with nothing between them other than constant-buffer maps (later in the frame pass 138 [c480e36e] and the lens flare [69c92902] read volume a).

the gpu is left launching 90 micro-jobs over and over, each blocked by the one before it.

[fig. 01]the volumetric chain, vanilla against collapsed (whiterun, frame 1920)
 vanilla                                collapsed
 ┌──────────────────────────────┐       ┌──────────────────────────────┐
 │ generate ab674eb1  10x6x90   │       │ generate ab674eb1  as before │
 ├──────────────────────────────┤       ├──────────────────────────────┤
 │ chain 1c4ebb62 #1  a → b     │       │ collapse, one dispatch:      │
 │ chain 1c4ebb62 #2  b → a     │       │  each thread walks its own   │
 │ chain 1c4ebb62 #3  a → b     │       │  (x, y) column through the   │
 │   …   90 dispatches, 10x6x1  │       │  whole chain, same loads and │
 │                              │       │  stores, same order          │
 │                              │       │ chain #2 .. #90 skipped      │
 └──────────────┬───────────────┘       └──────────────┬───────────────┘
                └──→ pass 138 c480e36e, lens flare 69c92902 read a ←──┘

trick or treat

each thread in the chain touches its own (x, y) column alone // threads are restricted from any neighbor’s column.

first frame, the fork studies the chain, records the first k, the count n, then the two volumes and the group size. when a first chain dispatch arrives after a generate, the fork sends one compute shader in its place. that single shader reads the chain column by column and replays every store, using typed loads and stores on both confirmed textures // the fork drops the remaining n-1 dispatches, only when they still match what was recorded. any mismatch, or a chain that stops early goes to the log + flips the collapse off.

is it exact?

well, each collapse thread carries out the same loads & stores.. in the same order.. against the same two textures. every store rounds to r16f along the same store path the chain itself would take // hardware performs its rounding on both sides.

vanilla samples the volume with SampleLevel at (k + 0.5) / d, and that must land exactly on the texel itself, either a point sampler or a linear one sitting precisely at a texel centre. the synthetic harness verifies with a linear sampler + the in-game check covers the real sampler. when k runs out of range, vanilla’s extra reads just feed a store past the end // both paths discard that store.

0th base

the synthetic harness mirrors the game’s layout (320x192x90 in r16f) with 90 vanilla chain dispatches at k = 1..90 & a constant buffer mapped for individual dispatch.

  • timing put the vanilla chain at 0.40 to 0.42 ms, against 0.14 ms with the collapse.
  • full read of both textures matched vanilla bit-for-bit, under a linear sampler + under a point sampler.
  • outcome at 160x90x64, 256x144x63 and 100x50x7.

ingame, BLESSED_VOL_COLLAPSE_VERIFY=1 leaves the first 8 chains running as vanilla. the collapse instead runs on copies captured before each chain, and when the last dispatch finishes the fork diffs the two. the daemon’s run came back 3/3 checks bit-identical, with 5,529,600 voxels touched per texture & 0 different // lever-verify-1, a check later, reached 8/8.

fast != truth

a running sum held in a register, rounded with f32tof16, ran 9% faster but review caught that it wasn’t exact. 5.26 million voxels drifted 1 ulp off, because the r16f store path rounds differently than f32tof16 does.

a round-six audit reopened the matter with care: find a register conversion that matches the typed store across halfway cases, neighbors, subnormals, overflow & live sums // after that move the sum into registers. it was noted that neither this nor new group shapes for the generate were enough to close strict.

bananathread

in synthetic harness, per collapse dispatch:

shapems
8x80.143
16x40.147
32x20.170
32x40.205
32x10.453

8x8 came out fastest.

ig ppl want framemaxxing huh

campaign c24 (old; out of c80), whiterun, ultra 1080p, 3 interleaved repeats:

rowfps1% lowgpu busy
perf (cb ring + threaded front end)193.0159.45.16 ms
perf + vol collapse201.0163.34.94 ms
perf + vol half-rate203.6158.55.11 ms

solo the collapse gained us +4.1% & half-rate volumetrics on its own got +5.5%. combined they reached +7.7%, short of the sum since the two overlap: a half-rate frame never runs the chain the collapse accelerates.

open items:

  • with community shaders it prob does nothing. community shaders swaps out the volumetric chain the collapse looks for. during c79 the round-eight audit found no log line of the chain being learned. activation could not be demonstrated // the tech likely rests cold.
  • the 0.14 ms doesn’t have an in-game timer yet. both figures, 0.40 to 0.42 ms and 0.14 ms, originate in the seat’s synthetic harness. round eight explored a reproducible source inside the named evidence.
  • the rounding question, open, as above.

report: closed.