![]() |
| You'll have to forgive my use of AI, aside from it being "appropriate", I haven't reinstalled my editing software after recently reinstalling the OS... |
It feels like an eternity since the ill-fated DLSS 5 reveal, back in March(?!!) but time has moved on and we a) now have the first implementation in a game, and b) have a slew of modded games and modding community integrations for inserting DLSS 5 into almost any and all games. We also have testing from the reviewer industry and a whitepaper from Nvidia (which is a little light on details). So, let's take a look at what I wrote in my first impressions, last time, and what all current data tells us, now...
Looking at What was Right and Wrong...
In my blogpost, I wrote the following:
"If you look at the actual technical information provided by Nvidia in their release and the short summary image above, you can see that DLSS neural rendering is simply a more advanced neural net guided lighting filter. It's a more advanced version of ENB and Reshade because it is fed the data from the frame (and prior frames) just like the DLSS upscaler and ray reconstruction. It's generating data as much (or maybe slightly less) than those two technologies do and that's why it's stable - that's why it's not mush that is uncontrollable, which real Gen AI is.DLSS 5 is Machine Learning (ML)."
In most ways, I was correct. In one way, I was wrong.
- DLSS 5 is a more advanced form of ENB/ReShade
- DLSS 5 is a more advanced form of lighting filter and is a more advanced form of neural net**
- DLSS 5 does use data from the frame but not prior frames*
- DLSS 5 is stable because it uses the trained ML model to interpret each frame and enhance the colour of each pixel
- DLSS 5 is not Gen AI
*Looking back, I'm not sure why I wrote that, as it was stated at the time to not be using prior frame data...
**It's a transformer model - which is an evolution in implementation of the original neural networks.
The evidence for all of these is provided by Nvidia's own whitepaper, though they still like to confuse the situation by mis-using terms. But the most succinct summary can be found in section 5.4 "Input Image Quality constraint": (highlights are mine)
"DLSS 5 produces an RGB image and does not modify the engine’s underlying geometry, scene state, or G-buffer contents. Its output is therefore fundamentally constrained by the information available in the rendered input.Although DLSS 5 can enrich local appearance, it cannot reliably recover scene information that is absent, incorrect, or severely degraded in the input. Missing, aliased, or otherwise corrupted scene evidence may be preserved or amplified. In such cases, the rendered observation does not provide enough evidence to support a faithful enhancement for DLSS 5. Low-quality scene rendering is particularly challenging: if object boundaries or structural details are weakly represented in the input, DLSS 5 cannot be expected to reconstruct their intended form reliably. The model is therefore better understood as an appearance-enhancement model rather than a mechanism for correcting arbitrary upstream rendering errors. It can improve the perceived realism of a valid and sufficiently informative input, but it cannot guarantee the reconstruction of missing geometry, object identity, or authored details. Figure 11 illustrates an example in which the aliasing and degraded detail in the window frames limit the quality of the corresponding region in the DLSS 5 output.This behavior reflects a deliberate trade-off. Where the input evidence is defective, a generative model could invent a plausible replacement, but an invented replacement risks changing what the image depicts. DLSS 5 is trained to take the safer side of that trade: it preserves the geometric structure of the input exactly — including defects such as aliasing (Sec. 2.4) — rather than risk altering the story or meaning of the image."
These things were pretty clear to me in the initial reveal but people really freaked out because Jensen lied and used "marketing/investor speak" instead of being accurate. If you remember my original blogpost, I pointed out the absurdity of his comments.
![]() |
| Lies, damn lies, and... graphs? |
Moving on to Today's Misdirections...
The preview we were shown of this technology back in March used two RTX 5090s showcase it. But the horsepower of both graphics cards was never required to actually run the technology - the demo was just set up like that for ease of demonstration. This was confirmed multiple times by Nvidia employees but most outlets did not report correctly. WCCFTech is one of the ones that I found that did. However, I remember Adam, of PC World fame, speaking about this on one of The Full Nerd Podcasts. So, we can safely remove the first plotted point in this graph (separately, it was pointed out to me by the venerable Osvaldo Pinali Doederlein that this first point is actually below 1x!!).
The second line says that it needs the RTX 5090 to run. I don't believe that, at all! It probably just describes the fact that they validated the technology on the 5090.
The unlabelled point is indeterminate - no idea what it represents. But the last point is just that they have also now validated the technology to run on the entire RTX 50 series consumer stack. How can I come to that conclusion? Well, it's a combination of the prior logic and also the fact that the RTX 5090 is approximately 420-500% stronger than the RTX 5090! So, you know, ~5x the improvement (rounding up)...
| Sure, it's not 5x but the RTX 5050 is compared at 1080p whereas the 5090 is compared at 4K, so we can extrapolate a little...[TechPowerUp] |
Coincidentally, the RTX 5090 is 2.24x the RTX 5070 (half-way down the stack)...
There's another reason for my skepticism of this claim of improvement. The metric that we should be counting here is the frametime cost of the technology - something which Nvidia do, themselves, in their whitepaper. That's a number in ms. Currently, the one number they are claiming is 8 ms at 4K resolution on an RTX 5090. According to their claims here, they were, at one point, (and my maths may be failing me right now!) at 40 ms for this same test (8/40 = 0.2 => 1/0.2 = 5x improvement).
The problem I have with such a claim is that the engineers didn't start from a blank canvas - they tell us that they built off of the already existing DLSS 4 transformer architecture. That means they already had a pre-existing, fast "backbone" to work from. I find it difficult to accept that their starting point was 25 fps or lower on two RTX 5090s... although, saying that, there was overhead from rendering a frame, transferring across the PCIe bus to the CPU and then back out to the second GPU (presumably on the same controller, otherwise, if they had the second GPU on the motherboard controller, that's probably quite a latency hit!).
So, to put it bluntly, I don't believe this graph for one minute!
Beam Me Up, Mr Scott...
Secondly, Nvidia's claim of 8 ms for a 4K frame is meaningless. It was already meaningless when it was officially revealed because it depends on the amount of tensor resources on the specific GPU - so this claim only applies to the RTX 5090 and this has been verified by various outlets performing initial testing on NBA 2K27.
It's doubly meaningless because it turns out that running this workload on the tensor cores is incredibly power consuming! Every single GPU in the stack is hitting the board power limit. That means power gating, frequency throttling, etc. It means that the DLSS 5 workload is not completing as fast as it can do on any card. Plus, for the 8 GB VRAM cards which are NOT power limited, there is a big issue with underutilisation due to exceeding the framebuffer...
![]() |
| Listing which tests run up to the listed TBP of the card at each resolution. Tests marked with an asterisk exceed the VRAM framebuffer, meaning that the GPU remains underutilised as data is shuttled back and forth to main system memory... - [Tom's Hardware] |
This was already partially shown by HardwareLuxx, but more careful testing performed by Tom's Hardware shows the effect of this by comparing the MSI Lightning Z RTX 5090 with two 12vhpwr connectors and an unlocked VBIOS with the stock/spec RTX 5090.
While I've looked at the data from Hardware Unboxed's and HardwareLuxx's testing, I'm discarding the obtained results because there appears to be a problem with DLSS upscaling in NBA 2K27. You can observe this in the RTX 5060 Ti 16 GB testing over in HUB's table, showing a plateau at 165 fps (DLSS 5 off!) but higher fps at native 1080p, along with a weird result for the RTX 5070 Ti at 1440p. Similarly, there are strange numbers reported from HWLuxx for the RTX 5070 Ti at 1440p, 4K, and the 5070 at 4K (could be a labelling typo but it's not corrected at the time I'm publishing this blogpost). In comparison, Tom's testing without DLSS upscaling shows no such inconsistencies - everything scales as expected, with the exception of the RTX 5060 Ti 8GB and RTX 5060/5050 where VRAM comes into play.
![]() |
| DLSS 5 frametime cost per card, per resolution. Red labelled values denote VRAM effect... - [Tom's Hardware] |
So, working with Tom's data, I've extracted the frametime cost of DLSS 5 for each card at each resolution and also cross referenced this with the power draw and frametime per tensor core. Additionally, I've summarised the tech specs of all RTX 40 & 50 series desktop cards, with some additional pertinent information in order to perform the following analysis.
![]() |
| My hierarchy of DLSS 5 performance for RTX 40 & 50 series cards. Red labelled values denote potential for switching of places depending on card behaviour... - [TechPowerUp] |
While I explored other relationships, the strongest correlation of DLSS 5 performance lay with Total Board Power (TBP) and the amount of matrix calculation potential in TFLOPS (since FP8 and FP16 are a function of each other, it doesn't matter which I plot - I chose FP8 since Nvidia stated that this is the primary format used for DLSS 5).
Frametime values which do not make sense (i.e. the highlighted red values above) are not included in the analysis and removed as inaccurate data).
What we observe is a non-linear relationship between the frametime cost in milliseconds for DLSS 5 and the power limit (Watts) and the matrix operations per second (TFLOPS).
So, in the best case scenario and without using framegen, the RTX 5060 Ti really is the hard cut-off at 1080p at about 150 fps base (8.6 ms frametime cost) - assuming VRAM and power isn't an issue.
![]() |
| Frametime plotted against Total Board Power at 4K... |
![]() |
| Frametime plotted against Total Board Power at 1440p... |
![]() |
| Frametime plotted against Total Board Power at 1080p... |
Returning to the table above which shows which resolutions and cards are power limited, we see that the 1080p has the most values which are not power limited. Thus, removing the power limited values, we obtain the following graph which shows a more accurate relationship between FP8 and DLSS 5 frametime cost.
From this, we can observe that (not unexpectedly) the curve is trending towards exponential in nature, never approaching zero. This also tells us that increasing power, the available TFLOPs of FP8, decreasing the data format (e.g. FP4), or increasing power efficiency, or some combination of these will all result in better performance in the DLSS 5 calculation.
![]() |
| Removing the power limited values, we obtain this plot... |
Conclusion...
I actually quite like the potential look of DLSS 5 in various games (where it makes sense). Ironically, I believe, based on these results, it's a technology that is ultimately better suited to being implemented on old games. The reason is that the more power-constrained a GPU is in any given game, the worse the frametime cost of DLSS 5 will be for that given card.
This means modding older games with DLSS 5 will always give the greatest power headroom for the technology to work within as they tend to challenge the hardware less than newer AAA titles. It also means that DLSS 5 will never perform the same across any game title. The claim of any specific frametime cost is going to be highly dependent on whether the GPU is power constrained or not... Which means it's not really even a metric we should be quoting...
![]() |
| A graphical representation of the frametime cost for DLSS 5 versus the base framerate (fps) to determine your output framerate (fps) - with a 60 fps cut-off... [Found here] |
We can also work out the theoretical relationship between required base framerate (before enabling DLSS 5) and the frametime cost (ms) of DLSS 5. Plotting that out from 30 fps up to 210 fps we get the above graph. Assuming people don't want to play below 60 fps before adding framegen*, we get the following approximate cut-off points:
*Or refusing to use it!
| Base framerate to achieve 60 fps with DLSS 5 enabled... |
So, in the best case scenario and without using framegen, the RTX 5060 Ti really is the hard cut-off at 1080p at about 150 fps base (8.6 ms frametime cost) - assuming VRAM and power isn't an issue.
For 1440p, the RTX 5070 Ti needs about 150 fps (7.4 ms cost) - assuming power isn't an issue.
For 4K, the RTX 5090 is the only viable option at 150 fps (8.4 ms cost) - once again, assuming power isn't an issue.
This brings me to the conclusion that the optimal base framerate for using DLSS 5 is above 150 fps, and the optimal DLSS 5 cost per resolution is around 8 ms. This means that the RTX 4060 Ti, 4060, 5060, and 5050 are not really ideal for using DLSS 5, except in very slow-paced games where framegen could be used without much fuss on top of ~30 fps base framerate.
Additionally, this technology lays the groundwork for Nvidia to prioritise matrix calculations which, based on my limited understanding of a large proportion of modern AI models (LLMs, Image Gen, and Machine Learning) can correspond to an "appropriate" technology investment in Nvidia's future GPU architectures. i.e. If Nvidia increase matrix calculations disproportionately to floating point and integer calculations (CUDA cores), along with ray tracing, per SM (Streaming Multiprocessor) unit then they will be able to point to this technology as being uplifted over the previous generation of graphics cards.
Finally, despite what Nvidia says, DLSS 5 is still not Generative AI. It's not generating anything. It's effectively highlighting and adapting the colour space - in the same way that DLSS upscaling does in a slightly different way. It's still Machine Learning and is based on the Machine Learning implemented as part of the DLSS 4 suite of technologies.










No comments:
Post a Comment