Skip to content

[NVEncC 9.30] onnxruntime issues #791

Description

@rainman74

cmdline:

NVEncC64.exe -i "Test.mkv" -o "Test_realesrgan_x4.mkv" --avhw --vpp-onnx model="%CMDPATH:\=/%\bin\onnx_models\realesrgan\realesrgan_x4plus.onnx",provider=cuda,out_res=3840x2160 --codec hevc --preset p5 --vbr 8000 --audio-copy --sub-copy --chapter-copy --log-level debug --log logfile.txt -o "out.mkv"

Encoding is stuck:

2026-08-04 18:30:04.9918607 [W:onnxruntime:, transformer_memcpy.cc:111 onnxruntime::MemcpyTransformer::ApplyImpl] 4 Memcpy nodes are added to the graph main_graph for CUDAExecutionProvider. It might have negative impact on performance (including unable to run CUDA graph). Set session_options.log_severity_level=1 to see the detail logs before this message.
2026-08-04 18:30:05.0077373 [W:onnxruntime:, session_state.cc:1316 onnxruntime::VerifyEachNodeIsAssignedToAnEp] Some nodes were not assigned to the preferred execution providers which may or may not have an negative impact on performance. e.g. ORT explicitly assigns shape related ops to CPU to improve perf.
2026-08-04 18:30:05.0136550 [W:onnxruntime:, session_state.cc:1318 onnxruntime::VerifyEachNodeIsAssignedToAnEp] Rerunning with verbose output on a non-minimal build will show node assignments.
NVEncC (x64) 9.30 (r4008) by rigaya, Aug  4 2026 12:13:43 (VC 1944/Win)
OS Version     Windows 11 x64 (28000) [UTF-8]
CPU            11th Gen Intel Core i9-11900K @ 3.50GHz [TB: 4.92GHz] (8C/16T)
GPU            #0: NVIDIA GeForce RTX 3080 (8704 cores, 1800 MHz)[PCIe4x16][610.62]
NVENC / CUDA   NVENC API 13.0, CUDA 13.3, schedule mode: auto
Input Buffers  CUDA, 16 frames
Input Info     avcuvid: H.264/AVC, 882x540, 25/1 fps
Vpp Filters    cspconv(nv12 -> yv12)
               onnx: realesrgan_x4plus.onnx  882x540 -> 3528x2160 (x4)  io=rgb  backend=cuda matrix=bt601 range=tv [NVIDIA GeForce RTX 3080] prec=f32 -> out_res 3840x2160 (lanczos4)
               cspconv(yv12 -> nv12)
Output Info    H.265/HEVC main @ Level auto
               3840x2160p 0:0 25.000fps (25/1fps)
               avwriter: hevc, ac3, subtitle#1 => matroska
Encoder Preset P5
Rate Control   VBR
Multipass      none
Bitrate        8000 kbps (Max: 24000 kbps)
Target Quality auto
QP range       I:0-51  P:0-51  B:0-51
QP Offset      cb:0  cr:0
VBV buf size   auto
Split Enc Mode auto
Tuning Info    hq
Lookahead      off
GOP length     250 frames
B frames       3 frames [ref mode: middle]
Ref frames     5 frames, MultiRef L0:auto L1:auto
AQ             off
CU max / min   32 / auto
Others         mv:auto
2026-08-04 18:30:13.4408620 [W:onnxruntime:, transformer_memcpy.cc:111 onnxruntime::MemcpyTransformer::ApplyImpl] 4 Memcpy nodes are added to the graph main_graph for CUDAExecutionProvider. It might have negative impact on performance (including unable to run CUDA graph). Set session_options.log_severity_level=1 to see the detail logs before this message.
2026-08-04 18:30:13.4557730 [W:onnxruntime:, session_state.cc:1316 onnxruntime::VerifyEachNodeIsAssignedToAnEp] Some nodes were not assigned to the preferred execution providers which may or may not have an negative impact on performance. e.g. ORT explicitly assigns shape related ops to CPU to improve perf.
2026-08-04 18:30:13.4609146 [W:onnxruntime:, session_state.cc:1318 onnxruntime::VerifyEachNodeIsAssignedToAnEp] Rerunning with verbose output on a non-minimal build will show node assignments.
2026-08-04 18:33:19.0929371 [E:onnxruntime:, sequential_executor.cc:572 onnxruntime::ExecuteKernel] Non-zero status code returned while running Conv node. Name:'node_conv2d_350' Status Message: E:\_work\1\s\onnxruntime\core\framework\bfc_arena.cc:359 onnxruntime::BFCArena::AllocateRawInternal Failed to allocate memory for requested buffer of size 2042295808

onnx: onnx: CUDAゼロコピー経路を初期化できないためホスト経路を使用します: Non-zero status code returned while running Conv node. Name:'node_conv2d_350' Status Message: E:\_work\1\s\onnxruntime\core\framework\bfc_arena.cc:359 onnxruntime::BFCArena::AllocateRawInternal Failed to allocate memory for requested buffer of size 2042295808
onnx: onnx: inference failed: E:\_work\1\s\onnxruntime\core\providers\cuda\cuda_call.cc:129 onnxruntime::CudaCall E:\_work\1\s\onnxruntime\core\providers\cuda\cuda_call.cc:121 onnxruntime::CudaCall CUDA failure 400: invalid resource handle ; GPU=0 ; hostname=DESKTOP ; file=E:\_work\1\s\onnxruntime\core\providers\cuda\cuda_stream_handle.cc ; line=41 ; expr=cudaEventRecord(event_, static_cast<cudaStream_t>(GetStream().GetHandle()));
onnx: .
CUDA: Error while running filter "onnx".
Break in task CUDA: unknown error..
NVDEC: in 0, out 1, pending -1, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
NVDEC: state 1, decoderQueue 20, dataFlagQueue 26, hdr10plusQueue 0.
AUDIO: in 1, out 1, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CHECKPTS: in 1, out 1, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CUDA: in 0, out 0, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CUDA: frame release: 1.
NVENC: in 0, out 0, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
NVENC: encodeBuffer used=0, free=22.
NVDEC: in 0, out 1, pending -1, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
NVDEC: state 1, decoderQueue 20, dataFlagQueue 26, hdr10plusQueue 0.
AUDIO: in 1, out 1, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CHECKPTS: in 1, out 1, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CUDA: in 0, out 0, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
CUDA: frame release: 1.
NVENC: in 0, out 0, pending 0, outQeueue size: 0, queue first(ts=-9223372036854775808,dur=0,id=-1), last(ts=-9223372036854775808,dur=0,id=-1).
NVENC: encodeBuffer used=0, free=22.

avout: File header not written, unexpected error! **(~100 times!!!)**

avout: failed to write header for output file: Invalid data found when processing input

encoded 0 frames, 0.00 fps, 0.00 kbps, 0.00 MB
encode time 0:03:10, CPULoad: 9.1%

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions