|
|
| Page 781 of 781 |
|
|
Posted: Tue, 1st Sep 2026 16:17 Post subject: |
|
 |
You can try the optiscaler implementation: https://github.com/Dagherbou/OptiScaler_DLSSNR/releases/tag/v0.1.1.5-dlssnr
It let's you change the NR resolution from 100% to 75%, 50% and 33%
But it drastically decreases the fidelity of the NR, imho it's already a heavy but usable impact at 75%, then becomes virtually useless below, even though it also runs way better.
I think the problem with the current implementations is that they can only run on the final frame
But the ideal way would be to run it together with something like DLSS, lets say render a 4k output at 1080p or so, apply the NR to that image, and then let DLSS do it's upscaling magic
That should dramatically improve performance while only marginally decreasing fidelity
Pretty sure that is how nvidia intended it, but obviously requires proper game implementation or a more in-depth game specific mod like PureDarks Skyrim upscaler
|
|
| Back to top |
|
 |
vurt
Posts: 14672
Location: Sweden
|
Posted: Tue, 1st Sep 2026 16:23 Post subject: |
|
 |
DLSS 5 launches on September 3 with NBA 2K27, gonna be interesting to see what the actual benchmark results are instead of how its implemented as a mod and with the pre-release version (though its of course possible nothing has changed).
| Amadeus wrote: |
What's that gonna run at, 5 fps and look weird and unstable af? |
No way, its a 23 year old game so a 50% or more drop in FPS is nothing if you get like 400 FPS or whatever, which would pretty reasonable for such an old game as CoD. not sure what you mean with more passes, its with DLSS 5 on and off, not dual DLSS 5..
I have a hard time thinking the performance drop can be this much for actual DLSS 5 when its implemented for games. it would appeal to less than 1% of gamers (5090 users are 0.4% of Steam users), so its practically for no one, or its something for next generation of cards and to hype those up. We'll see.
Edit: talked a little to chapgpt and this is what it predicts in terms of performance
Most likely: 5070 Ti works quite well at 1440p, perhaps with a noticeable ~15–30% neural-rendering cost. DLSS SR + FG compensates for it. RTX 5090 gets the spectacular 4K showcase numbers.
Less pleasant: 5070 Ti works, but DLSS 5 is expensive enough that you really need aggressive upscaling + FG/MFG. NVIDIA can truthfully advertise support, while RTX 60 later gets much faster neural inference.
Worst case: current Blackwell is effectively a preview platform and DLSS 5 only becomes genuinely attractive on RTX 60. I think this is possible, but less likely.
|
|
| Back to top |
|
 |
harry_theone
Posts: 11609
Location: The Land of Thread Reports
|
Posted: Tue, 1st Sep 2026 16:58 Post subject: |
|
 |
@Amadeus Haha no chance with a 4090
|
|
| Back to top |
|
 |
|
|
Posted: Wed, 2nd Sep 2026 01:04 Post subject: |
|
 |
| vurt wrote: | DLSS 5 launches on September 3 with NBA 2K27, gonna be interesting to see what the actual benchmark results are instead of how its implemented as a mod and with the pre-release version (though its of course possible nothing has changed).
| Amadeus wrote: |
What's that gonna run at, 5 fps and look weird and unstable af? |
No way, its a 23 year old game so a 50% or more drop in FPS is nothing if you get like 400 FPS or whatever, which would pretty reasonable for such an old game as CoD. not sure what you mean with more passes, its with DLSS 5 on and off, not dual DLSS 5..
I have a hard time thinking the performance drop can be this much for actual DLSS 5 when its implemented for games. it would appeal to less than 1% of gamers (5090 users are 0.4% of Steam users), so its practically for no one, or its something for next generation of cards and to hype those up. We'll see.
Edit: talked a little to chapgpt and this is what it predicts in terms of performance
Most likely: 5070 Ti works quite well at 1440p, perhaps with a noticeable ~15–30% neural-rendering cost. DLSS SR + FG compensates for it. RTX 5090 gets the spectacular 4K showcase numbers.
Less pleasant: 5070 Ti works, but DLSS 5 is expensive enough that you really need aggressive upscaling + FG/MFG. NVIDIA can truthfully advertise support, while RTX 60 later gets much faster neural inference.
Worst case: current Blackwell is effectively a preview platform and DLSS 5 only becomes genuinely attractive on RTX 60. I think this is possible, but less likely. |
That 50% number is only a relative estimate for modern games
Of course the compute cost of something like DLSS 5 is many magnitudes larger than rendering a 23 year old game on a modern GPU, which will hit engine, driver and CPU limits first
So rather than going from 400 fps to 200 fps you're going to maybe 50-60 fps
There really is a limit on how many NR frames of a given resolution a RTX gpu can output
For example Deus Ex runs at like 600 fps but with NR only runs at around 70 fps, that's on a 5080
Because while the GPU can output many more rasterized frames than that, it can't produce more NR frames than that
PureDark managed to implement a way to select a different GPU for framegen and NR
So he let his 5090 run the raster on Skyrim
And a 4070ti running DLSS 5 / NR
It managed to do around 30fps
But I am sure official implementations will run better, esp. since they will likely run NR on frames before upscaling, so at a lower resolution
That Call of duty (2003) screenshot definitely ran multiple passes of DLSS 5 thanks to special mods that let you run it up to 10 times on each frame, which ofc will result in more extreme results, more artifacts and far less performance.
|
|
| Back to top |
|
 |
couleur
[Moderator] Janitor
Posts: 14982
|
Posted: Wed, 2nd Sep 2026 09:26 Post subject: |
|
 |
| Amadeus wrote: | ...
PureDark managed to implement a way to select a different GPU for framegen and NR
So he let his 5090 run the raster on Skyrim
And a 4070ti running DLSS 5 / NR
It managed to do around 30fps
... |
Pretty cool though, that you can run it that way. But in this specific case wouldn't it be better to use the 5090 for the DLSS 5 part? Surely that would result in more FPS?
"Enlightenment is man's emergence from his self-imposed nonage. Nonage is the inability to use one's own understanding without another's guidance. This nonage is self-imposed if its cause lies not in lack of understanding but in indecision and lack of courage to use one's own mind without another's guidance. Dare to know! (Sapere aude.) "Have the courage to use your own understanding," is therefore the motto of the enlightenment."
|
|
| Back to top |
|
 |
vurt
Posts: 14672
Location: Sweden
|
Posted: Wed, 2nd Sep 2026 21:09 Post subject: |
|
 |
| Amadeus wrote: |
That Call of duty (2003) screenshot definitely ran multiple passes of DLSS 5 thanks to special mods that let you run it up to 10 times on each frame |
i checked and you're right, it was a 3x pass. still, really impressive result.. looking forward to test it.
|
|
| Back to top |
|
 |
couleur
[Moderator] Janitor
Posts: 14982
|
|
| Back to top |
|
 |
|
|
Posted: Thu, 3rd Sep 2026 09:08 Post subject: |
|
 |
It depends on the game (enigine) a lot.
Another example: https://www.reddit.com/r/nvidia/s/pGvamXhSQF
Ryzen 9850X3D CO Per Core ~-29 | Light Loop 360 & 3x Phanteks T30 | ProArt X870E-CREATOR WIFI | MSI GeForce RTX 5090 Ventus OC | Fury Renegade RGB 64GB (2x 32GB) DDR5 6400MHz C32 @ 6000MHz C28 | FURY Renegade G5 4TB PCIe 5.0 | 40" G75F | S.M.S.L RAW-MDA1 & HiFiMAN Arya Organic | Lancool III Snow White + 4x be quiet! Silent Wings Pro 4 140mm | HX1200i | G Pro X SUPERLIGHT 2 & POWERPLAY | Win 11 Pro | Logitech MX MECHANICAL
Sometimes I publish YouTube videos: https://www.youtube.com/@RandomTechChannel
|
|
| Back to top |
|
 |
couleur
[Moderator] Janitor
Posts: 14982
|
Posted: Thu, 3rd Sep 2026 09:12 Post subject: |
|
 |
Yeah, that looks awesome.
"Enlightenment is man's emergence from his self-imposed nonage. Nonage is the inability to use one's own understanding without another's guidance. This nonage is self-imposed if its cause lies not in lack of understanding but in indecision and lack of courage to use one's own mind without another's guidance. Dare to know! (Sapere aude.) "Have the courage to use your own understanding," is therefore the motto of the enlightenment."
|
|
| Back to top |
|
 |
vurt
Posts: 14672
Location: Sweden
|
Posted: Thu, 3rd Sep 2026 10:36 Post subject: |
|
 |
It's how it looks and performs modded with a pre-release... how its applied in an official pipeline might be different.
Trails in the Sky looks like several layers of it btw, someone did it to make some stupid point and to get viral for the anti-AI bros.
|
|
| Back to top |
|
 |
|
|
|
| Back to top |
|
 |
|
|
Posted: Mon, 7th Sep 2026 12:15 Post subject: |
|
 |
|
|
|
| Back to top |
|
 |
LeoNatan
☢ NFOHump Despot ☢
Posts: 75349
Location: Israel
|
Posted: Mon, 7th Sep 2026 22:42 Post subject: |
|
 |
That's more akin to Ageia PhysX coprocessor than SLI.
|
|
| Back to top |
|
 |
Slizza
Posts: 2414
Location: Bulgaria
|
Posted: Mon, 7th Sep 2026 22:54 Post subject: |
|
 |
I was surprised Nvidia were not trying to flog a second GPU for ray tracing.
No more multi GPU for me. Not with today electricity prices.
|
|
| Back to top |
|
 |
Frant
King's Bounty
Posts: 24948
Location: Your Mom
|
|
| Back to top |
|
 |
Frant
King's Bounty
Posts: 24948
Location: Your Mom
|
Posted: Mon, 14th Sep 2026 23:34 Post subject: |
|
 |
| Amadeus wrote: | https://github.com/maohgad-web/Neural-coprocessor
SLI is back.. or something? |
Interesting.
So GPU-1 renders the game frame and sends it to GPU-2 via PCIe, GPU-2 performs the DLSS5 processing and sends the resulting frame, again over PCIe, to GPU-1 where it's displayed on the screen.
That feels like at least a ~2 fps latency introduced. I don't know though, during my quick glance at the github text files I didn't see any clear mention about added latency, just that there is (obviously) latency in this method.
* note: technical rant incoming
I dropped a thorough description to the AI bot in Brave and got this:
| Quote: | Added Latency: ~1.3–1.5 ms per frame
The dominant new cost is the PCIe 4.0 x16 round trip:
Component Time
Frame size (1440p RGBA) 2560 × 1440 × 4 ≈ 14.7 MB
PCIe 4.0 x16 practical bandwidth ~25 GB/s per direction
One-way transfer 14.7 MB ÷ 25 GB/s ≈ 0.59 ms
Round trip (GPU-1 → GPU-2 → GPU-1) ~1.18 ms
TLP headers, credit-based flow control, sync ~0.1–0.3 ms
Total added latency ~1.3–1.5 ms
This is on top of the latency the frame already had. The DLSS processing time itself doesn't add new latency — it was always going to happen — but it now sits between two PCIe transfers on the critical path. |
It was still missing something so added some thoughts about the overall expected performance cost of the mod itself. That changed things quite considerably.
| Quote: | There are several layers of cost beyond the raw PCIe transfer that I under-weighted:
1. Driver Hook / Interception Overhead
The modified driver must intercept the normal render→DLSS→present pipeline. Every frame, the CPU-side driver code now does additional work.
Signal render complete on GPU-1 (fence/event)
Queue the outbound transfer to GPU-2
Schedule the DLSS kernel on GPU-2's command queue
Wait for GPU-2 completion (poll or interrupt)
Queue the return transfer
Signal present to the display
2. Cross-Device Synchronization Cost
The multi-GPU synchronization research (NVIDIA V100/P100 study) shows that cross-device sync over PCIe is significantly more expensive than intra-device sync. Even with hardware-assisted P2P, each synchronization point (fence signal → wait → release) carries:
Ring buffer polling on the CPU or GPU side
Completion signaling (interrupt or MSI-X)
Cache coherency flushes if the frame metadata crosses the boundary
3. The "Modified Driver" Quality Factor — This Is the Big One
This is where your instinct is strongest. A community-modified driver (like the DLSS5-Feeder project on GitHub) introduces costs that NVIDIA's own driver would never have.
Memory allocation Stock driver uses pinned, pre-allocated, cache-line-aligned buffers. A mod may use generic allocations that aren't zero-copy, adding a CPU memcpy hop
Command batching Stock driver batches hundreds of draw calls per submission. A mod that injects per-frame sync points fragments the batch, increasing submission frequency
Polling vs. interrupts A well-written mod uses GPU-side fence polling (no CPU wake). A less-optimized one polls from the CPU, burning a core and adding jitter
Ring buffer management Two GPUs means two ring buffers the CPU must keep filled. If the mod doesn't pre-allocate deep enough rings, the CPU stalls waiting for GPU-2 to drain before it can submit the next frame
Error handling / fallback paths A mod that checks for GPU-2 availability, thermal state, or fallback-to-single-GPU every frame adds branch overhead to the hot path
4. Revised Total Latency Budget
Component Time
PCIe round trip (transfer) ~1.2 ms
Driver hook / extra API calls ~0.3–0.5 ms
Cross-device sync (fences, interrupts) ~0.2–0.4 ms
Pipelined handoff (1-frame delay) ~6.9 ms (at 144 FPS)
Total added latency ~8.6–9.0 ms
This aligns with Digital Foundry's measurement of the DLSS 5 dual-GPU mod adding 8–12 ms to the pipeline. The PCIe transfer is the small part — the synchronization architecture and the pipelining requirement dominate. |
It's an added latency of at LEAST ~1+ frame(s) in a game rendering at 1440p @ 144 Hz just from the process of the mod itself as far as deriving a plausible penalty for the whole process.
It's not really much of an issue though. The main benefit is that GPU-1 can render higher in-game quality settings at a good framerate since it doesn't have to deal with the DLSS5 processing.
Ph'nglui mglw'nafh Cthulhu R'lyeh wgah'nagl fhtagn!
|
|
| Back to top |
|
 |
| Page 781 of 781 |
All times are GMT + 1 Hour |
|
You cannot post new topics in this forum You cannot reply to topics in this forum You cannot edit your posts in this forum You cannot delete your posts in this forum You cannot vote in polls in this forum
|
Powered by phpBB 2.0.8 © 2001, 2002 phpBB Group
|
|
 |
|