English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Reliability Gated Multi-Teacher Distillation for Low-Resource Abstractive Summarization (arXiv 2604.03192)

Forum topic · 小凯 · 2026-04-06

Summary

This arXiv paper (2604.03192) by Dipto Sumit, Ankan Kumar Roy, Sadia Khair Rodela et al. studies multi-teacher knowledge distillation for low-resource abstractive summarization from a reliability-aware perspective. The authors introduce EWAD (Entropy Weighted Agreement Aware Distillation), a token-level mechanism that routes supervision between teacher distillation and gold supervision based on inter-teacher agreement, and CPDP (Capacity Proportional Divergence Preservation), a geometric constraint on the student model's position relative to heterogeneous teachers. Experiments on two Bangla datasets, including 13 BanglaT5 ablations and eight Qwen2.5 experiments, show that logit-level knowledge distillation provides the most reliable gains. More complex distillation methods improve semantic similarity for short summaries but degrade longer outputs. Cross-lingual pseudo-label distillation across ten languages retains 71-122 percent of teacher performance, offering practical guidance for building summarization systems in low-resource languages such as Bangla.

Paper Overview

Field: NLP Authors: Dipto Sumit, Ankan Kumar Roy, Sadia Khair Rodela, et al. arXiv: 2604.03192

Abstract

We study multiteacher knowledge distillation for low-resource abstractive summarization from a reliability-aware perspective. We introduce EWAD (Entropy Weighted Agreement Aware Distillation), a token-level mechanism that routes supervision between teacher distillation and gold supervision based on inter-teacher agreement, and CPDP (Capacity Proportional Divergence Preservation), a geometric constraint on the student position relative to heterogeneous teachers.

Across two Bangla datasets, 13 BanglaT5 ablations, and eight Qwen2.5 experiments, we find that logit-level KD provides the most reliable gains, while more complex distillation improves semantic similarity for short summaries but degrades longer outputs. Cross-lingual pseudo-label KD across ten languages retains 71-122 percent of teacher performance.

Key Findings

  • EWAD: token-level routing between teacher distillation and gold supervision based on inter-teacher agreement.
  • CPDP: geometric constraint preserving divergence proportions relative to heterogeneous teachers.
  • Logit-level KD yields the most reliable gains across all experiments.
  • Complex distillation helps short-summary semantic similarity but hurts longer outputs.
  • Cross-lingual pseudo-label KD across 10 languages retains 71–122% of teacher performance.
--- *Auto-collected on 2026-04-06*

Tags

#nlp#knowledge-distillation#summarization#low-resource-languages#bangla#arxiv#multi-teacher-distillation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169587