Mixture-of-experts models have become a common way to scale large language models without activating all parameters for every token. A token only routes to a small subset of experts, yet the full expert pool remains in memory. Expert pruning reduces that storage burden, and the new RAZOR method approaches it by asking which experts are actually replaceable.

At a fixed pruning budget, RAZOR's goal is to remove experts while keeping the original model's output distribution as intact as possible. The name captures the intuition that some experts are redundant—their contributions can be covered by others—so they are safe to cut. The abstract frames the problem as a distribution-preservation task rather than a simple accuracy benchmark.

The source is limited to the abstract, so it includes no experimental numbers, baselines, or detailed comparisons. The contribution described is the framing and method for identifying replaceable experts, and no competing methods are discussed in the provided text.