[TensorRT] IOutputAllocator로 데이터 의존적 출력 버퍼를 안전하게 받는 법
TensorRT 10과 cuda.bindings에서 데이터 의존적 출력의 메모리 상한, bounded IOutputAllocator, callback shape를 검증하는 방법입니다.
대상 환경은 Linux, NVIDIA GPU, TensorRT 10 계열의 name-based Python API, cuda.bindings를 사용하는 추론 워커입니다. 입력 shape는 정상적으로 지정했지만 출력에 -1이 남거나, 출력 버퍼를 고정 크기로 잡은 뒤 enqueue 실패·잘못된 복사 크기·간헐적 OOM이 발생하는 증상을 다룹니다.
이 글은 TensorRT Dynamic Shape Profile 실전 글의 후속입니다. 이전 글은 입력의 min·opt·max shape와 latency를 기준으로 profile을 설계했지만, 입력 크기가 같아도 실제 데이터에 따라 결과 개수가 달라지는 출력 버퍼는 다루지 않았습니다. 운영에서는 이 빈틈을 큰 버퍼 하나로 덮으면 profile 상한이 커질 때 워커 전체 VRAM을 흔들 수 있습니다. 이번에는 상한을 먼저 확인하고, 바이트 예산 안에서만 버퍼를 등록한 뒤 실제 shape를 callback으로 받는 다음 단계를 구성합니다.
1. 입력 동적 shape와 데이터 의존적 출력을 구분한다
입력의 batch·높이·너비가 바뀌는 것은 optimization profile과 set_input_shape로 해결합니다. 반면 INonZeroLayer처럼 실제 값에 따라 출력 원소 수가 달라지면 enqueue 전에는 최종 shape를 알 수 없습니다. 이때는 다음 세 값을 따로 관리해야 합니다.
- profile과 입력 shape로 계산한 출력 메모리 상한
- allocator callback이 요청받은 실제 바이트 수
notify_shape가 알려 준 실제 출력 shape
최종 shape에 -1이 남아 있다는 이유만으로 엔진 오류로 판정하지 않습니다. 먼저 입력이 충분히 지정됐는지 확인하고, 데이터 의존적 출력이면 allocator 경로로 넘깁니다.
2. 엔진 I/O와 profile을 애플리케이션 밖에서 확인한다
배포 이미지와 같은 TensorRT 설치에서 엔진의 I/O, layer, profile 정보를 먼저 덤프합니다. images와 shape는 실제 이름과 profile 범위 안 값으로 바꿉니다.
trtexec --loadEngine=model.engine \
--shapes=images:1x3x640x640 \
--profilingVerbosity=detailed \
--dumpLayerInfo --exportLayerInfo=engine_layers.json기대 관찰값은 입력 이름과 dtype이 애플리케이션 계약과 같고, 문제 출력의 차원에 동적 표시가 있으며, 해당 출력을 만드는 layer를 추적할 수 있는 것입니다. trtexec 자체가 같은 shape에서 실패하면 Python 버퍼 코드보다 엔진·plugin·profile 호환성을 먼저 봅니다.
3. 상한을 VRAM 예산으로 바꾼다
TensorRT의 get_max_output_size는 현재 profile과 입력 shape를 기준으로 출력 바이트 상한을 제공합니다. 입력 shape를 지정하기 전에 호출하거나 출력 이름이 아니면 -1을 반환할 수 있습니다. 따라서 device 선택, runtime·engine·context 생성, profile 선택, 입력 shape·주소 지정, shape inference, 출력 상한 조회 순서를 지킵니다.
아래 검증기는 모든 출력 상한 합계가 --max-total-output-mib를 넘으면 메모리를 할당하기 전에 실패시킵니다. 상한을 그대로 허용할 수 없는 서비스에서는 profile을 줄이거나 모델의 후보 개수 상한을 낮추는 것이 우선입니다.
4. bounded IOutputAllocator 검증기
예시는 단일 DEVICE·LINEAR 입력과 하나 이상의 DEVICE·LINEAR 출력을 대상으로 합니다. shape inference I/O나 vectorized format은 별도 주소·stride 계약이 필요하므로 중단합니다. 최신 경로인 reallocate_output_async와 이전 10.x 호환 callback을 모두 구현하되, callback 안에서 새 메모리를 무제한 할당하지 않습니다.
import argparse
import json
import os
import numpy as np
import tensorrt as trt
from cuda.bindings import runtime as cudart
def check_cuda(result):
err, *values = result
if err != cudart.cudaError_t.cudaSuccess:
raise RuntimeError(f"CUDA error: {err}")
if not values:
return None
return values[0] if len(values) == 1 else tuple(values)
def tensor_nbytes(shape, dtype):
return int(np.prod(shape, dtype=np.int64)) * np.dtype(dtype).itemsize
class BoundedOutputAllocator(trt.IOutputAllocator):
def __init__(self, pointer, capacity_bytes):
trt.IOutputAllocator.__init__(self)
self.pointer = int(pointer)
self.capacity_bytes = int(capacity_bytes)
self.requested_bytes = 0
self.output_shape = None
def _accept(self, size):
self.requested_bytes = int(size)
if self.requested_bytes > self.capacity_bytes:
return None
return self.pointer
def reallocate_output(self, tensor_name, current_memory, size, alignment):
return self._accept(size)
def reallocate_output_async(
self, tensor_name, current_memory, size, alignment, stream
):
return self._accept(size)
def notify_shape(self, tensor_name, dims):
self.output_shape = tuple(int(dim) for dim in dims)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--engine", required=True)
parser.add_argument("--input", required=True)
parser.add_argument("--device", type=int, default=0)
parser.add_argument("--profile", type=int, default=0)
parser.add_argument("--max-total-output-mib", type=int, default=512)
args = parser.parse_args()
check_cuda(cudart.cudaSetDevice(args.device))
current_device = check_cuda(cudart.cudaGetDevice())
if current_device != args.device:
raise RuntimeError(
f"device mismatch: requested={args.device}, current={current_device}"
)
logger = trt.Logger(trt.Logger.WARNING)
runtime = trt.Runtime(logger)
with open(args.engine, "rb") as file:
engine = runtime.deserialize_cuda_engine(file.read())
if engine is None:
raise RuntimeError("engine deserialize failed")
context = engine.create_execution_context()
if context is None:
raise RuntimeError("execution context creation failed")
names = [engine.get_tensor_name(i) for i in range(engine.num_io_tensors)]
inputs = [
name for name in names
if engine.get_tensor_mode(name) == trt.TensorIOMode.INPUT
]
outputs = [
name for name in names
if engine.get_tensor_mode(name) == trt.TensorIOMode.OUTPUT
]
if len(inputs) != 1:
raise RuntimeError(f"expected one input, got {inputs}")
input_name = inputs[0]
for name in names:
if engine.is_shape_inference_io(name):
raise RuntimeError(f"shape inference I/O is unsupported: {name}")
if engine.get_tensor_location(name) != trt.TensorLocation.DEVICE:
raise RuntimeError(f"only DEVICE tensors are supported: {name}")
if engine.get_tensor_format(name) != trt.TensorFormat.LINEAR:
raise RuntimeError(f"only LINEAR tensors are supported: {name}")
host_input = np.ascontiguousarray(np.load(args.input))
input_dtype = np.dtype(trt.nptype(engine.get_tensor_dtype(input_name)))
if host_input.dtype != input_dtype:
raise RuntimeError(
f"input dtype mismatch: file={host_input.dtype}, engine={input_dtype}"
)
stream = None
device_pointers = []
allocators = {}
try:
stream = check_cuda(cudart.cudaStreamCreate())
stream_handle = int(stream)
if args.profile != 0:
if not context.set_optimization_profile_async(
args.profile, stream_handle
):
raise RuntimeError(f"cannot select profile {args.profile}")
declared = tuple(engine.get_tensor_shape(input_name))
if any(dim < 0 for dim in declared):
if not context.set_input_shape(input_name, host_input.shape):
raise RuntimeError(f"input shape rejected: {host_input.shape}")
elif declared != host_input.shape:
raise RuntimeError(
f"input shape mismatch: file={host_input.shape}, engine={declared}"
)
input_pointer = check_cuda(cudart.cudaMalloc(host_input.nbytes))
device_pointers.append(input_pointer)
if not context.set_tensor_address(input_name, int(input_pointer)):
raise RuntimeError(f"cannot bind input {input_name}")
unresolved = context.infer_shapes()
if unresolved:
raise RuntimeError(f"insufficient shape inputs: {list(unresolved)}")
capacities = {}
for name in outputs:
shape_before = tuple(context.get_tensor_shape(name))
dtype = np.dtype(trt.nptype(engine.get_tensor_dtype(name)))
if shape_before and all(dim >= 0 for dim in shape_before):
capacity = tensor_nbytes(shape_before, dtype)
else:
capacity = int(context.get_max_output_size(name))
if capacity <= 0:
raise RuntimeError(f"invalid output upper bound: {name}={capacity}")
capacities[name] = capacity
total_capacity = sum(capacities.values())
budget = args.max_total_output_mib * 1024 * 1024
if total_capacity > budget:
raise RuntimeError(
f"output upper bound exceeds budget: "
f"required={total_capacity}, budget={budget}"
)
for name, capacity in capacities.items():
pointer = check_cuda(cudart.cudaMalloc(capacity))
device_pointers.append(pointer)
allocator = BoundedOutputAllocator(pointer, capacity)
allocators[name] = allocator
if not context.set_tensor_address(name, int(pointer)):
raise RuntimeError(f"cannot bind output {name}")
if not context.set_output_allocator(name, allocator):
raise RuntimeError(f"cannot set output allocator: {name}")
check_cuda(cudart.cudaMemcpyAsync(
input_pointer,
host_input.ctypes.data,
host_input.nbytes,
cudart.cudaMemcpyKind.cudaMemcpyHostToDevice,
stream,
))
if not context.execute_async_v3(stream_handle):
raise RuntimeError("inference enqueue failed")
check_cuda(cudart.cudaStreamSynchronize(stream))
report_outputs = {}
for name, allocator in allocators.items():
if allocator.output_shape is None:
raise RuntimeError(f"notify_shape was not called: {name}")
if any(dim < 0 for dim in allocator.output_shape):
raise RuntimeError(
f"unresolved output after enqueue: {name}={allocator.output_shape}"
)
dtype = np.dtype(trt.nptype(engine.get_tensor_dtype(name)))
actual_bytes = tensor_nbytes(allocator.output_shape, dtype)
if actual_bytes > allocator.capacity_bytes:
raise RuntimeError(
f"actual output exceeds capacity: {name}, "
f"actual={actual_bytes}, capacity={allocator.capacity_bytes}"
)
host_output = np.empty(allocator.output_shape, dtype=dtype)
if actual_bytes:
check_cuda(cudart.cudaMemcpyAsync(
host_output.ctypes.data,
allocator.pointer,
actual_bytes,
cudart.cudaMemcpyKind.cudaMemcpyDeviceToHost,
stream,
))
check_cuda(cudart.cudaStreamSynchronize(stream))
report_outputs[name] = {
"shape": list(allocator.output_shape),
"dtype": str(dtype),
"requested_bytes": allocator.requested_bytes,
"capacity_bytes": allocator.capacity_bytes,
"finite": bool(np.isfinite(host_output).all()),
}
print(json.dumps({
"status": "PASS",
"pid": os.getpid(),
"logical_device": current_device,
"profile": args.profile,
"total_output_capacity_bytes": total_capacity,
"outputs": report_outputs,
}, ensure_ascii=False, indent=2))
finally:
if stream is not None:
cudart.cudaStreamSynchronize(stream)
for pointer in reversed(device_pointers):
cudart.cudaFree(pointer)
if stream is not None:
cudart.cudaStreamDestroy(stream)
if __name__ == "__main__":
main()5. 실행하고 관찰할 값
GPU UUID 하나만 노출하면 프로세스 내부의 논리 device는 0입니다. 출력 전체의 사전 할당 상한을 512 MiB로 제한한 예시는 다음과 같습니다.
CUDA_VISIBLE_DEVICES=GPU-xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx \
python verify_dynamic_outputs.py \
--engine model.engine \
--input sample_input.npy \
--device 0 --profile 0 \
--max-total-output-mib 512기대 관찰값은 JSON의 status가 PASS이고, 모든 출력의 shape가 0 이상의 정수이며, requested_bytes가 capacity_bytes 이하인 것입니다. 데이터 의존적 출력은 입력 내용에 따라 actual shape가 달라져도 상한과 서비스 예산 안에 머물러야 합니다.
최소·일반·최악 입력 세트로 반복해 상한 대비 실제 사용률과 GPU 메모리를 함께 기록합니다.
for sample in empty_case.npy normal_case.npy crowded_case.npy; do
python verify_dynamic_outputs.py \
--engine model.engine --input "$sample" \
--device 0 --profile 0 --max-total-output-mib 512
done
nvidia-smi --query-compute-apps=pid,gpu_uuid,used_memory \
--format=csv -l 1기대 관찰값은 빈 결과도 0차원을 포함한 정상 shape로 처리되고, 혼잡 입력에서도 예산 초과나 비대상 GPU 사용이 없으며, 같은 profile의 capacity는 입력 값에 따라 흔들리지 않는 것입니다.
6. 실패 위치별 진단 분기
| 실패 지점 | 우선 확인 | 판정 |
|---|---|---|
| infer_shapes가 이름 반환 | 입력 shape와 shape tensor 주소 | 입력 계약 누락 |
| 상한이 -1 또는 0 | profile 선택 순서, 출력 이름 | 상한 조회 시점 또는 I/O 식별 오류 |
| 상한 합계가 예산 초과 | profile max, 후보 개수 상한 | 할당 전 배포 거부, profile 재설계 |
| enqueue가 false | TensorRT 로그, callback 요청 바이트, plugin | 버퍼 부족 또는 실행 오류 |
| notify_shape 미호출 | allocator 객체 수명, 출력 연결 | callback 등록 오류 |
| 첫 실행 뒤 latency 변동 | CUDA memory pool과 동기화 구간 | 실행 중 재할당·pool release 가능성 |
7. 배포와 롤백 기준
배포 허용
- 최소·일반·최악 입력이 각각 3회 연속 PASS한다.
- 출력 상한 합계가 워커별 VRAM 예산 안이고 동시 worker 수를 곱한 값도 노드 여유 메모리를 넘지 않는다.
- 모든 callback 요청이 capacity 이하고, 실제 shape·dtype·유한값 검사와 후처리 결과가 기준 구현과 일치한다.
- PID와 GPU UUID가 배포 설정과 일치하며 반복 실행 latency가 기존 허용 범위 안이다.
즉시 롤백
출력 상한이 예산을 넘거나 callback이 capacity보다 큰 크기를 요구하면 해당 엔진·profile을 배포하지 않습니다. enqueue 실패, 잘못된 최종 shape, 비유한 출력, 다른 GPU에서의 PID 관찰 중 하나라도 발생하면 새 워커를 트래픽에서 제외하고 마지막으로 검증된 엔진과 profile로 되돌립니다. CUDA 오류가 난 프로세스는 예외를 삼켜 재사용하지 말고 종료해야 오류 상태가 다음 요청과 섞이지 않습니다.
결론
데이터 의존적 출력의 핵심은 최종 shape를 미리 맞히는 것이 아니라, profile 기반 바이트 상한을 서비스 예산으로 제한하고 TensorRT가 알려 주는 실제 요청 크기와 shape를 실행 후 검증하는 것입니다. 상한 조회, bounded allocator, callback 관측, 최악 입력 테스트를 하나의 readiness gate로 묶으면 고정 버퍼의 overflow와 무제한 재할당의 OOM을 모두 배포 전에 차단할 수 있습니다.