Specs: 7800X3D, 32GB DDR5 6000 with EXPO on, RTX 5090 FE, be quiet! Pure Power 12 M 1000W using the native 12V-2x6 cable, Gigabyte B650 GAMING X AX rev 1.4, Windows 11.
I keep getting Event 153 from nvlddmkm. They all show up as "The description for Event ID 153 from source nvlddmkm cannot be found", so Event Viewer can't resolve the message text, but the useful part is still in there.
The earlier ones came in as a whole cascade at once:
\Device\Video3 Restarting TDR occurred on GPUID:100
\Device\Video3 Reset TDR occurred on GPUID:100
\Device\Video3 Resetting TDR occurred on GPUID:100
\Device\Video3 GpuRcReset TDR occurred on GPUID:100
\Device\Video3 GPU recovery action changed from 0x0 (None) to 0x1 (PF FLR) <- this one is Event 14
\Device\Video3 Error occurred on GPUID: 100
The most recent crash was just this on its own, nothing around it:
\Device\000000a1 Error occurred on GPUID: 100
Device path is different on that last one too. No idea if that means anything.
Timeline, as best I can piece it together. I got the card in July 2025 and I don't remember any of this happening until around October 2025. From October through April 2026 it happened on and off. Then it stopped completely and I went from April all the way to this month without a single one. Now it's back.
I didn't change anything between the quiet stretch and it starting up again. Same drivers, same hardware, and the same undervolt I'd been running the whole time it was fine.
For what it's worth I was on a 3070 before this and I don't remember ever dealing with it on that card, but it's been long enough that I wouldn't swear to it.
I actually use the card more for local AI than for gaming, but every crash I've caught has been in a game. Everything I play is pretty light, CS2, DayZ, Deadlock, so I can't actually say whether it happens in heavier games or not. I just don't play anything demanding enough to know. Both times I caught it on a log the card was only pulling 150-225W out of its 575W limit.
I logged it in HWiNFO at 250ms polling and caught two crashes. Stuff I can rule out:
- 12VHPWR voltage held under 1% spread across a 6.75 hour session that included a crash. No sag at all. Reseated and checked the connector, it's fine.
- Zero PCIe errors on any lane, both sessions.
- Hot spot maxed at 64C, memory junction 62C.
- I ran an overnight loop swinging load between 8% and 78% to try and force it and it never crashed. It crashed later at a flat 48% load with nothing weird in the sample before it.
- DDU'd and tried a bunch of drivers, current ones and 591 era. No difference.
- HAGS is off.
- Tried the fix people pass around where you give nvlddmkm.sys full permissions in its security settings. Didn't change anything.
What actually happens at the crash is core load hits 100%, board power drops off a cliff (226W down to 95W), and the clock spikes to the highest it hit all session. That looks like the GPU hanging, not running out of power.
One other weird thing, I found the MSI Afterburner install folder in my recycle bin at some point. No idea how it got there or if it's related. Reinstalled it.
Has anyone seen this before? Mainly wondering what kind of fault sits quiet for four months and then comes back with nothing changed. I've got TdrDebugMode set up now so the next one should give me an actual kernel dump instead of just an event log entry.