DPA-MoE: A Runtime Precision Control Framework for MoE Inference
Jaeyun Kim*, Hoonki Lee*, Susie Shin, Hyunmin Kim, Dahyun Lee, Sangyeob Kim†
Allocates precision across routed experts at runtime, so one MoE checkpoint covers a range of weight-size–quality trade-offs without retraining. To appear in Findings of EMNLP 2026.
Abstract and citation
To appear in Findings of EMNLP 2026.
DPA-MoE is a runtime precision control framework that dynamically allocates different precision levels to the routed experts of a mixture-of-experts model. Because the allocation happens at inference time, a single checkpoint supports flexible weight-size–quality trade-offs with no retraining and no per-precision copy of the weights.
Jaeyun Kim and Hoonki Lee contributed equally; Sangyeob Kim is the corresponding author.
@inproceedings{kim2026dpamoe,
title = {{DPA}-MoE: A Runtime Precision Control Framework for {M}o{E} Inference},
author = {Kim, Jaeyun and Lee, Hoonki and Shin, Susie and Kim, Hyunmin and
Lee, Dahyun and Kim, Sangyeob},
booktitle = {Findings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026},
url = {https://openreview.net/forum?id=BrcO2Gg9A9}
}
Jaeyun Kim*, Hoonki Lee*, Susie Shin, Hyunmin Kim, Dahyun Lee, and Sangyeob Kim†. DPA-MoE: A Runtime Precision Control Framework for MoE Inference. In Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026). To appear. (*equal contribution, †corresponding author)