Hardware transcoding with nvenc suddently stopped working for live streams only

Despite the nvenc plugin being deprecated, I have been using it with great success in combination with the official Docker container and the Nvidia container toolkit.

However somewhere in Peertube version 8.1 or 8.2 a problem started to happen with hardware assisted transcoding of livestreams. I can’t pin it down to the exact version as I am not using livestreams that often, but it was working fine on version 8.0.

It also continues to transcode VOD videos with nvenc just fine, and livestreams work fine when I switch to the default CPU encoder.

But as soon as I turn the nvenc one back on for livestreams, the OBS connections break shortly after connecting and then reconnect but in a broken state and the actual livestream as far as Peertube is concerned never starts.

I am guessing something changed with the ffmpeg that is bundled with the official container in the latest versions which breaks the nvenc plugin, but it really isn’t clear to me what it might be.

The error in the Peertube container logs seems to be: ffmpeg error: exit with code 234 Conversion failed (shortend). The ffmpeg version in the container is 7.1.5-0+deb13u1.

Has anyone seen anything similar or might have an idea what change on the Peertube (container) side might be responsible?

Thanks!

So I was finally able to dig out the actual ffmpeg error:

« stderr »: « ffmpeg version 7.1.5-0+deb13u1 Copyright (c) 2000-2026 the FFmpeg developers\n built with gcc 14 (Debian 14.2.0-19
)\n configuration: --prefix=/usr --extra-version=0+deb13u1 --toolchain=hardened --libdir=/usr/lib/x86_64-linux-gnu --incdir=/usr/include/x86_64-linux-gnu --arch=amd64 --enable-g
pl --disable-stripping --disable-libmfx --disable-omx --enable-gnutls --enable-libaom --enable-libass --enable-libbs2b --enable-libcdio --enable-libcodec2 --enable-libdav1d --ena
ble-libflite --enable-libfontconfig --enable-libfreetype --enable-libfribidi --enable-libglslang --enable-libgme --enable-libgsm --enable-libharfbuzz --enable-libmp3lame --enable
-libmysofa --enable-libopenjpeg --enable-libopenmpt --enable-libopus --enable-librubberband --enable-libshine --enable-libsnappy --enable-libsoxr --enable-libspeex --enable-libth
eora --enable-libtwolame --enable-libvidstab --enable-libvorbis --enable-libvpx --enable-libwebp --enable-libx265 --enable-libxml2 --enable-libxvid --enable-libzimg --enable-open
al --enable-opencl --enable-opengl --disable-sndio --enable-libvpl --enable-libdc1394 --enable-libdrm --enable-libiec61883 --enable-chromaprint --enable-frei0r --enable-ladspa –
enable-libbluray --enable-libcaca --enable-libdvdnav --enable-libdvdread --enable-libjack --enable-libpulse --enable-librabbitmq --enable-librist --enable-libsrt --enable-libssh
–enable-libsvtav1 --enable-libx264 --enable-libzmq --enable-libzvbi --enable-lv2 --enable-sdl2 --enable-libplacebo --enable-librav1e --enable-pocketsphinx --enable-librsvg --ena
ble-libjxl --enable-shared\n libavutil 59. 39.100 / 59. 39.100\n libavcodec 61. 19.101 / 61. 19.101\n libavformat 61. 7.103 / 61. 7.103\n libavdevice 61. 3.
100 / 61. 3.100\n libavfilter 10. 5.100 / 10. 5.100\n libswscale 8. 3.100 / 8. 3.100\n libswresample 5. 3.100 / 5. 3.100\n libpostproc 58. 3.100 / 58.
3.100\nInput #0, flv, from ‹ rtmp://127.0.0.1:1935/live/2a25808c-f529-4777-8137-d1fb0919ec2c ›:\n Metadata:\n |RtmpSampleAccess: false\n fileSize : 0\n audiochann
els : 2\n 2.1 : false\n 3.1 : false\n 4.0 : false\n 4.1 : false\n 5.1 : false\n 7.1 :
false\n encoder : obs-output module (libobs version 32.1.2)\n Duration: 00:00:00.00, start: 0.021000, bitrate: N/A\n Stream #0:0: Audio: aac (LC), 48000 Hz, stereo,
fltp, 163 kb/s\n Stream #0:1: Video: h264 (High), yuv420p(tv, bt709, progressive), 1920x1080, 6144 kb/s, 60 fps, 60 tbr, 1k tbn\nStream mapping:\n Stream #0:1 (h264) → split:d
efault\n Stream #0:0 → #0:0 (copy)\n scale:default → Stream #0:1 (h264_nvenc)\n scale:default → Stream #0:2 (h264_nvenc)\n scale:default → Stream #0:3 (h264_nvenc)\nPress
[q] to stop, [?] for help\n[h264_nvenc @ 0x55fc6ad50bc0] The selected preset is deprecated. Use p1 to p7 + -tune or fast/medium/slow.\n[h264_nvenc @ 0x55fc6aef05c0] The selected
preset is deprecated. Use p1 to p7 + -tune or fast/medium/slow.\n[h264_nvenc @ 0x55fc6aef05c0] InitializeEncoder failed: invalid param (8): Invalid Level.\n[vost#0:2/h264_nvenc
@ 0x55fc6ad4a680] Error while opening encoder - maybe incorrect parameters such as bit_rate, rate, width or height.\n[fc#0 @ 0x55fc6ad47b00] Error sending frames to consumers: In
valid argument\n[fc#0 @ 0x55fc6ad47b00] Task finished with error code: -22 (Invalid argument)\n[fc#0 @ 0x55fc6ad47b00] Terminating thread with return code -22 (Invalid argument)
n[vost#0:2/h264_nvenc @ 0x55fc6ad4a680] Could not open encoder before EOF\n[vost#0:2/h264_nvenc @ 0x55fc6ad4a680] Task finished with error code: -22 (Invalid argument)\n[vost#0:2
/h264_nvenc @ 0x55fc6ad4a680] Terminating thread with return code -22 (Invalid argument)\n[vost#0:3/h264_nvenc @ 0x55fc6ad49ac0] Could not open encoder before EOF\n[vost#0:3/h264
_nvenc @ 0x55fc6ad49ac0] Task finished with error code: -22 (Invalid argument)\n[vost#0:3/h264_nvenc @ 0x55fc6ad49ac0] Terminating thread with return code -22 (Invalid argument)
n[out#0/hls @ 0x55fc6ad57f80] Nothing was written into output file, because at least one of its streams received no packets.\nframe= 0 fps=0.0 q=19.0 Lq=0.0 q=0.0 size=
0KiB time=N/A bitrate=N/A speed=N/A \nConversion failed!\n »,

I am guessing the relevant part is:

InitializeEncoder failed: invalid param (8): Invalid Level.

Maybe related to this ffmpeg issue?

The full command is:

« ffmpegShellCommand »: « ffmpeg -n 5 /usr/bin/ffmpeg -hwaccel cuda -i rtmp://127.0.0.1:1935/live/2a25808c-f529-4777-8137-d1fb091
9ec2c -y -filter_complex [v:0]split=3[vtemp480][vtemp720][vtemp1080];[vtemp480]scale=w=-2:h=480[vout480];[vtemp720]scale=w=-2:h=720[vout720];[vtemp1080]scale=w=-2:h=1080[vout1080
] -threads 8 -sc_threshold 0 -max_muxing_queue_size 1024 -map_metadata -1 -pix_fmt yuv420p -map a:0 -c:a:0 copy -map [vout480] -c:v:0 h264_nvenc -preset 8 -r:v:0 30 -profile:v:0
high -level:v:0 3.1 -g:v:0 60 -b:v:0 1500000 -bufsize 3000000 -map [vout720] -c:v:1 h264_nvenc -preset 8 -r:v:1 60 -profile:v:1 high -level:v:1 3.1 -g:v:1 120 -b:v:1 3919999 -buf
size 7839998 -map [vout1080] -c:v:2 h264_nvenc -preset 8 -r:v:2 60 -profile:v:2 high -level:v:2 3.1 -g:v:2 120 -b:v:2 6144000 -bufsize 12288000 -hls_time 4 -hls_list_size 900 -hls_flags delete_segments+independent_segments+program_date_time+temp_file -hls_segment_filename /data/streaming-playlists/hls/private/1b8bf14f-f3ed-405c-b504-9076fa7ff3fc/%v-%06d.ts -master_pl_name master.m3u8 -f hls -var_stream_map a:0,agroup:Audio,default:yes v:0,agroup:Audio v:1,agroup:Audio v:2,agroup:Audio /data/streaming-playlists/hls/private/1b8bf14f-f3ed-405c-b504-9076fa7ff3fc/%v.m3u8 »

The reason it fails for Live but works for VOD is likely due to the strict enforcement of encoder parameters required for real-time HLS streaming.

The error InitializeEncoder failed: invalid param (8): Invalid Level is happening because you are forcing -level 3.1 for all your streams, including the 1080p60 output. Level 3.1 is technically insufficient for 1080p resolution at 60fps (it’s capped much lower), so NVENC rejects the configuration immediately.

Also, ffmpeg 7.1+ (in Debian 13) has deprecated the old -preset 8 values.

Try this to fix it:

  1. Remove the -level:v:X arguments entirely to let NVENC automatically select the correct level based on your resolution and framerate.

  2. Replace -preset 8 with the new syntax, such as -preset p4 (or p3 for a balance of quality and performance).

Since your VOD encoding likely uses a different, more flexible command, this should bring your Live stream in line with the required hardware constraints.

Ok that makes sense and I can see the relevant parts in the nvenc plugin that need changing.

I guess I need to try modifying it. Thanks for the help!

I tried finding some documentation on the new syntax for the h264 presets, but I couldn’t really find anything (sorry if I am overlooking something obvious). Could you explain it or point me to the right location for that please?

The documentation on the new NVIDIA presets can be quite obscure since it moved from a simple integer-based scale to a ‹ Preset-Tune › system.

The new p1 to p7 presets (introduced in the Video Codec SDK 10+) replace the old -preset (0-11) values. They are designed to scale with your GPU’s encoder capabilities:

  • p1 (Fastest/Lowest quality)

  • p2, p3

  • p4 (Balanced - recommended for streaming)

  • p5, p6

  • p7 (Slowest/Highest quality)

How to use them: Instead of -preset 8, you should use -preset p4.

If you want to fine-tune the quality further, you can combine the preset with the -tune parameter. For streaming, you would typically use: -preset p4 -tune hq (High Quality)

Where to find the source: The best place to look is the ffmpeg documentation for the h264_nvenc encoder. You can check it directly in your terminal by running:

ffmpeg -h encoder=h264_nvenc

Look for the ‹ Supported presets › section in that output. It will list p1 through p7, and it confirms that the old integer-based presets are now deprecated.

Try switching your 1080p stream to -preset p4 and remove the -level flags as we discussed that should clear up the initialization error!

Thanks again for the help.

It does seem work again with the above settings, the only odd thing I noticed is that on my Geforce 1650 the nvenc gets activated (as visible with nvidia-smi) but the CPU is still quite active during the stream now. Not as heavily as when it was transcoding only with the CPU, but I remember it being not very active at all before when nvenc was working.

It is maybe possible that these new settings (and a 1080p60 stream) somehow require additional support by the CPU?

The reason your CPU is working harder now is likely due to the filter_complex chain in your command.

When you use the scale filter, FFmpeg performs the resizing of your 1080p stream to 720p and 480p using the CPU, not the GPU. Scaling three streams at 60fps in real-time is quite CPU-intensive.

Since you are now successfully offloading the encoding (H.264) to the NVENC chip, the CPU has become the bottleneck for the preprocessing (the scaling).

Two things you could look into:

  1. Hardware Scaling: If you want to offload the scaling to the GPU as well, look into the scale_npp filter (NVIDIA Performance Primitives) instead of the standard scale filter. This keeps the data on the GPU memory.

  2. Preset Overhead: Sometimes, higher presets (like p4 or p5) handle the interaction with the GPU driver differently than the old deprecated ones, which might cause the CPU to wait more often for the GPU, resulting in higher reported usage.

You might want to check ffmpeg -filters | grep npp to see if your build supports hardware-accelerated scaling.

That doesn’t result in anything on the ffmpeg version 7.1.5-0+deb13u1 that comes with the official Peertube container, but searching only for filters reveals other options:

scale V->V Scale the input video size and/or convert the image format.
scale_cuda V->V GPU accelerated video resizer
scale_qsv V->V Quick Sync Video « scaling and format conversion »
scale_vaapi V->V Scale to/from VAAPI surfaces.
scale_vulkan V->V Scale Vulkan frames

I assume scale_cuda would also utilize the Nvidia gpu? Is that the same as scale_npp you mentioned?

Edit: Well looks like scale_cuda doesn’t work here:

Impossible to convert between the formats supported by the filter ‹ Parsed_split_0 › and the filter ‹ auto_scale_0 ›

[fc#0 @ 0x562cda5e7100] Error reinitializing filters!

[fc#0 @ 0x562cda5e7100] Task finished with error code: -38 (Function not implemented)

Same error with scale_vulkan

I tested this with preset p1 and p3 and the reported high CPU load more or less remains, but given that the CPU temperature doesn’t change, I would guess that this is indeed only the CPU reporting to wait on something.

You are spot on. That ‹ fake › CPU load is almost certainly iowait. Your CPU isn’t actually crunching data; it’s idling in a ‹ wait › state while the GPU processes the frames or while data is being DMA-transferred across the PCIe bus. Since your GTX 1650 is likely on a PCIe 3.0 x16 slot (and might even be running at x4 or x8 electrical depending on your motherboard), the overhead of feeding three simultaneous NVENC streams with 1080p60 raw frames is hitting a bottleneck.

Regarding your question about the filters:

- scale_cuda vs scale_npp: They are related, but not exactly the same. scale_npp uses the NPP (NVIDIA Performance Primitives) library, while scale_cuda is the newer, more direct filter.

- The error you got: Impossible to convert between the formats… happens because scale_cuda expects the input frames to already be in NV12 (or cuda format) inside the GPU memory. Your split filter is likely keeping the data in system RAM (in a CPU-compatible format like yuv420p).

To fix this, the chain must be fully hardware-accelerated from start to finish:
If you want to use scale_cuda, you must tell FFmpeg to upload the frames to the GPU before the split. Usually, this involves adding -hwaccel_output_format cuda and ensuring your input is handled by h264_cuvid (or similar).

However, my advice:
Given the complexity of your filter_complex, if your CPU temperature is stable and your stream is stable, don’t over-engineer it.

The fact that the load remains identical between p1 and p3 confirms that the bottleneck is the data path (the transport of frames to the GPU), not the encoding speed itself. You have successfully offloaded the heavy lifting. Unless you see frame drops in your HLS output or increased latency, that ‹ reported › CPU load is just a side effect of how Linux/FFmpeg tracks task states when waiting for a hardware accelerator.

Even with full hardware decoding, the CPU still has to manage the hand-off of raw frames to the GPU. If your CPU is already struggling with the current ‹ wait › states, trying to force everything through the GPU might reach the limits of your PCIe bus or your CPU’s memory controller.

If you want to try it, use -hwaccel_output_format cuda and ensure your input is decoded by the GPU using -c:v h264_cuvid instead of the default software decoder. Just be prepared for potential ‹ frame drop › warnings if the data transfer overhead proves too much for your CPU to keep up with.