A new arXiv paper proposes model casting, a mid-training recipe designed to drastically sparsify activations inside the feed-forward network (FFN) layer of neural networks. The authors position this as a step toward more sparsely activated FFNs, which could reduce inference-time computation.
The method changes how inference is performed: rather than computing all FFN activations, the system first computes the output of a gating mechanism. The abstract indicates this gating output is then used to determine which parts of the FFN to activate, though the full details of the gating scheme and the resulting efficiency gains are not included in the available text.
Because only the abstract is available, this summary is limited to the paper's stated approach. There is no independent source to compare against, so the claims about sparsification and the gating strategy are presented as the authors' own description.