Skip to content
Jaeyun Kim

Publications

Papers, preprints and reports.

2026

DPA-MoE: A Runtime Precision Control Framework for MoE Inference

Jaeyun Kim*, Hoonki Lee*, Susie Shin, Hyunmin Kim, Dahyun Lee, Sangyeob Kim

* Equal contribution · † Corresponding author

Findings of EMNLP 2026 · November 1, 2026

Allocates precision across routed experts at runtime, so one MoE checkpoint covers a range of weight-size–quality trade-offs without retraining. To appear in Findings of EMNLP 2026.

Abstract and citation

To appear in Findings of EMNLP 2026.

DPA-MoE is a runtime precision control framework that dynamically allocates different precision levels to the routed experts of a mixture-of-experts model. Because the allocation happens at inference time, a single checkpoint supports flexible weight-size–quality trade-offs with no retraining and no per-precision copy of the weights.

Jaeyun Kim and Hoonki Lee contributed equally; Sangyeob Kim is the corresponding author.

@inproceedings{kim2026dpamoe,
  title     = {{DPA}-MoE: A Runtime Precision Control Framework for {M}o{E} Inference},
  author    = {Kim, Jaeyun and Lee, Hoonki and Shin, Susie and Kim, Hyunmin and
               Lee, Dahyun and Kim, Sangyeob},
  booktitle = {Findings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026},
  url       = {https://openreview.net/forum?id=BrcO2Gg9A9}
}

Jaeyun Kim*, Hoonki Lee*, Susie Shin, Hyunmin Kim, Dahyun Lee, and Sangyeob Kim†. DPA-MoE: A Runtime Precision Control Framework for MoE Inference. In Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026). To appear. (*equal contribution, †corresponding author)