Overview
Deep learning models frequently encounter variable-size inputs, which traditionally lead to inefficiencies during batching due to the practice of padding. Standard dense batching mechanisms allocate a shared envelope for all elements within a batch, dedicating computational resources to these padded regions. This inefficiency escalates with an increasing number of varying axes; for instance, an explicit pair state might require $BN_{\max}^2$ positions instead of $\sum_i N_i^2$ for actual data. While packing can mitigate some of this waste, it often obscures the logical axes and sample boundaries, complicating the composition of subsequent packed operations.
DanLing NestedTensor addresses these challenges by introducing a PyTorch tensor abstraction that fundamentally integrates multi-ragged structure as an intrinsic property of the tensor itself. This system allows packed values to carry tensor-backed partitions and maintain logical dimension order. This design enables the creation of ragged axes through broadcasting, the preservation of these axes during feature transformations, and their consumption during reduction operations. The same representation strategy is maintained consistently across autograd, eager execution, and compiled execution environments.
Research Context
Variable-size inputs are a common characteristic across various deep learning applications. The conventional approach to handle such inputs involves dense batching, where all sequences or structures within a batch are padded to the length of the longest item. This method ensures uniform tensor dimensions, which are convenient for many deep learning frameworks and hardware accelerators. However, it results in computational overhead and increased memory allocation due to processing and storing the padding elements. The cost associated with this padding multiplies significantly when dealing with multiple varying axes; for example, processing explicit pair states can lead to an allocation of $BN_{\max}^2$ positions, even though the actual data might only sum up to $\sum_i N_i^2$ positions.
While techniques like packing can reduce this waste by storing non-padded elements contiguously, they often flatten the data, thereby losing explicit information about logical axes and individual sample boundaries. This loss of structural metadata can complicate the composition of subsequent operations that require an understanding of the original, variable structure. The challenge lies in developing a representation that is both memory-efficient and capable of preserving the essential structural information needed for complex deep learning operations without requiring manual offset management at each call site within the model code.
Approach
The DanLing NestedTensor system was developed as a PyTorch tensor abstraction. Its core mechanism involves making multi-ragged structure an inherent property of the tensor. This is achieved by having packed values within the tensor carry tensor-backed partitions. These partitions explicitly encode the logical dimension order and delineate sample boundaries, even within a packed buffer.
This approach facilitates several key functionalities:
- Broadcasting operations can inherently create ragged axes.
- Feature transformations are designed to retain the established ragged structures.
- Reduction operations are capable of consuming these ragged structures directly.
The system ensures that this consistent representation of multi-ragged structure is maintained throughout the entire deep learning workflow, encompassing autograd calculations, eager execution modes, and compiled execution environments. The tensor interface is designed to enable model code, built using its supported operators, to compose efficient variable-size computations. A central design goal was to eliminate the need for manual offset management at any call site within the model code, thereby simplifying development for variable-size inputs.
Findings
The implementation of DanLing NestedTensor demonstrated specific performance improvements across different deep learning models and execution modes:
- BERT Scales: On an A100 GPU, the system achieved a geometric-mean speedup of $2.74\times$ in eager execution and $3.39\times$ in compiled execution when compared to same-mode padding across four different BERT scales.
- FCN Backbones: For four FCN backbones, the system exhibited a $1.97\times$ geometric-mean speedup in eager execution.
- Pairformer-style Workload: A four-block Pairformer-style workload, when run in eager execution using native PyTorch kernels across various square length regimes, operated $2.40-4.32\times$ faster than a padded reference.
- Memory Allocation: For its high-variation batch, the peak memory allocation for the Pairformer-style workload decreased from $38.08$ GiB to $5.41$ GiB.
Why This Matters
The DanLing NestedTensor system provides a mechanism for deep learning practitioners to manage variable-size inputs more efficiently, potentially reducing the computational and memory overhead traditionally associated with padding. By embedding multi-ragged structure directly into the tensor abstraction, it enables the creation of model code that can compose efficient variable-size computation without manual offset management. The observed speedups and memory reductions suggest potential for more resource-efficient training and inference of deep learning models that process diverse input lengths or structures.
Potential Applications
The tensor interface presented allows model code, constructed from its supported operators, to compose efficient variable-size computation without requiring manual offset management at any call site. This implies that the system is directly applicable to deep learning models that inherently deal with inputs of varying sizes, such as those processing sequences, graphs, or sets, allowing for more streamlined and efficient development and execution.