// nonfiction
- doc
- report
- 013
- filed
- 2026.09.28
- read
- 005 passes
- state
- [filed]
ninety to one: volumetric collapse
one of the stranger tricks in blessed-skyrim :3
status
| t+0.000 | shape | vanilla: 1 generate + 90 chain dispatches. ours: 1 generate + 1 |
| t+0.001 | cost | chain 0.40 to 0.42 ms → 0.14 ms, in the seat's synthetic harness |
| t+0.002 | exact | in game: 8 of 8 checks bit-identical (lever-verify-1) |
| t+0.003 | frame | c24, whiterun ultra 1080p: 193.0 → 201.0 fps, +4.1% |
| t+0.004 | open | under community shaders it likely never engages. not demonstrated |
scope
blessed-skyrim is my dxvk fork tuned for skyrim se, i wanna yank one tech out of it. the longer writeup is in past native. this is just looking at one of the mechanisms: collapsed compute dispatch
the tech qualifies as image exact so it gets to run on both profiles, strict included; strict does not permit temporal reuse and must match vanilla precisely.
the material came from notes in blessed-dxvk (blessed-notes/vol-collapse-notes.md), from campaign c24 in the tuning report, and from the audits of rounds six and eight.
what does skyrim even do every frame…
…it’s visible in a frame dump of whiterun (whiterun-dump-1, frame 1920). one generate dispatch opens shader ab674eb1, running 10x6x90 thread groups. it fills two volumes; a and b, each R16_FLOAT at 320x192x90.
chain follows: 90 dispatches of shader 1c4ebb62 (10x6x1 groups apiece) alternating constantly. dispatch reads a and writes b, then the next reads b and writes a.. with nothing between them other than constant-buffer maps (later in the frame pass 138 [c480e36e] and the lens flare [69c92902] read volume a).
the gpu is left launching 90 micro-jobs over and over, each blocked by the one before it.
vanilla collapsed
┌──────────────────────────────┐ ┌──────────────────────────────┐
│ generate ab674eb1 10x6x90 │ │ generate ab674eb1 as before │
├──────────────────────────────┤ ├──────────────────────────────┤
│ chain 1c4ebb62 #1 a → b │ │ collapse, one dispatch: │
│ chain 1c4ebb62 #2 b → a │ │ each thread walks its own │
│ chain 1c4ebb62 #3 a → b │ │ (x, y) column through the │
│ … 90 dispatches, 10x6x1 │ │ whole chain, same loads and │
│ │ │ stores, same order │
│ │ │ chain #2 .. #90 skipped │
└──────────────┬───────────────┘ └──────────────┬───────────────┘
└──→ pass 138 c480e36e, lens flare 69c92902 read a ←──┘trick or treat
each thread in the chain touches its own (x, y) column alone // threads are restricted from any neighbor’s column.
first frame, the fork studies the chain, records the first k, the count n, then the two volumes and the group size. when a first chain dispatch arrives after a generate, the fork sends one compute shader in its place. that single shader reads the chain column by column and replays every store, using typed loads and stores on both confirmed textures // the fork drops the remaining n-1 dispatches, only when they still match what was recorded. any mismatch, or a chain that stops early goes to the log + flips the collapse off.
is it exact?
well, each collapse thread carries out the same loads & stores.. in the same order.. against the same two textures. every store rounds to r16f along the same store path the chain itself would take // hardware performs its rounding on both sides.
vanilla samples the volume with SampleLevel at (k + 0.5) / d, and that must land exactly on the texel itself, either a point sampler or a linear one sitting precisely at a texel centre. the synthetic harness verifies with a linear sampler + the in-game check covers the real sampler. when k runs out of range, vanilla’s extra reads just feed a store past the end // both paths discard that store.
0th base
the synthetic harness mirrors the game’s layout (320x192x90 in r16f) with 90 vanilla chain dispatches at k = 1..90 & a constant buffer mapped for individual dispatch.
- timing put the vanilla chain at 0.40 to 0.42 ms, against 0.14 ms with the collapse.
- full read of both textures matched vanilla bit-for-bit, under a linear sampler + under a point sampler.
- outcome at
160x90x64,256x144x63and100x50x7.
ingame, BLESSED_VOL_COLLAPSE_VERIFY=1 leaves the first 8 chains running as vanilla. the collapse instead runs on copies captured before each chain, and when the last dispatch finishes the fork diffs the two. the daemon’s run came back 3/3 checks bit-identical, with 5,529,600 voxels touched per texture & 0 different // lever-verify-1, a check later, reached 8/8.
fast != truth
a running sum held in a register, rounded with f32tof16, ran 9% faster but review caught that it wasn’t exact. 5.26 million voxels drifted 1 ulp off, because the r16f store path rounds differently than f32tof16 does.
a round-six audit reopened the matter with care: find a register conversion that matches the typed store across halfway cases, neighbors, subnormals, overflow & live sums // after that move the sum into registers. it was noted that neither this nor new group shapes for the generate were enough to close strict.
bananathread
in synthetic harness, per collapse dispatch:
| shape | ms |
|---|---|
| 8x8 | 0.143 |
| 16x4 | 0.147 |
| 32x2 | 0.170 |
| 32x4 | 0.205 |
| 32x1 | 0.453 |
8x8 came out fastest.
ig ppl want framemaxxing huh
campaign c24 (old; out of c80), whiterun, ultra 1080p, 3 interleaved repeats:
| row | fps | 1% low | gpu busy |
|---|---|---|---|
| perf (cb ring + threaded front end) | 193.0 | 159.4 | 5.16 ms |
| perf + vol collapse | 201.0 | 163.3 | 4.94 ms |
| perf + vol half-rate | 203.6 | 158.5 | 5.11 ms |
solo the collapse gained us +4.1% & half-rate volumetrics on its own got +5.5%. combined they reached +7.7%, short of the sum since the two overlap: a half-rate frame never runs the chain the collapse accelerates.
open items:
- with community shaders it prob does nothing. community shaders swaps out the volumetric chain the collapse looks for. during c79 the round-eight audit found no log line of the chain being learned. activation could not be demonstrated // the tech likely rests cold.
- the 0.14 ms doesn’t have an in-game timer yet. both figures, 0.40 to 0.42 ms and 0.14 ms, originate in the seat’s synthetic harness. round eight explored a reproducible source inside the named evidence.
- the rounding question, open, as above.
report: closed.