Variable-size inputs are common across deep learning, but standard dense batching forces them into a shared envelope, spending compute on padding. According to a new arXiv paper, this overhead grows sharply when multiple axes vary at once. For example, an explicit pair state can require B times N_max squared positions, where the cost multiplies with each varying dimension.
The paper introduces DanLing NestedTensor, a composable multi-ragged tensor representation designed to avoid such padding waste. Rather than allocating a dense rectangular block, the approach composes ragged tensors to match the actual shape of each input. The abstract highlights this as a way to reduce the multiplicative cost of dense batching across varying axes.
While the full details of the implementation are not in the abstract, the motivation is clear: eliminating padding in variable-size inputs could cut both memory and computation in models that handle sequences, graphs, or other irregular data.