Summary
SkMTEB is introduced as the first comprehensive MTEB-style text embedding benchmark for the Slovak language, consisting of 31 datasets spanning 7 task types. Building on this benchmark, the authors develop two Slovak-adapted models, e5-sk-small (45M parameters) and e5-sk-large (365M parameters), by applying vocabulary trimming and fine-tuning to Multilingual E5 models. Despite parameter size reductions of up to 62%, the adapted models achieve performance competitive with proprietary embedding APIs while remaining small enough for local deployment. The work addresses the gap in embedding evaluation resources for less-represented languages and demonstrates that language-specific adaptation can substantially reduce model size without sacrificing quality. The paper is authored by Marek Šuppa, Andrej Ridzik, Daniel Hládek, Natália Kňažeková, and Viktória Ondrejová, and is available on arXiv (2606.13647).
Overview
Field: NLP
Authors: Marek Šuppa, Andrej Ridzik, Daniel Hládek, Natália Kňažeková, Viktória Ondrejová
arXiv: 2606.13647
Abstract (translated from the forum post)
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, comprising 31 datasets across 7 task types. We develop e5-sk-small (45M) and e5-sk-large (365M) by applying vocabulary trimming and fine-tuning to Multilingual E5 models. Despite size reductions up to 62%, our models achieve competitive performance with proprietary APIs while remaining locally deployable.
Key points
- SkMTEB: the first comprehensive MTEB-style text embedding benchmark for Slovak.
- Covers 31 datasets across 7 task types.
- Two adapted models built from Multilingual E5 via vocabulary trimming and fine-tuning:
- e5-sk-small: 45M parameters
- e5-sk-large: 365M parameters
- Model size reductions of up to 62% while remaining competitive with proprietary embedding APIs.
- Models are locally deployable, removing dependence on external API services.
---
*Source: forum post auto-collected on 2026-06-14.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981287