Telenor Nordics Customer Service Self-Help Corpus

TL;DR AI
2 min readKey summary
Researchers released a multilingual Nordic telecom self-help dataset with 1,122 validated documents.
The public corpus contains over one million tokens from Finnish, Danish, Norwegian, and Swedish support pages.
It was cleaned for privacy with PII filtering and validated through LLM-assisted human review.
The dataset fills a major gap in high-quality Nordic customer service data for search, retrieval, and language-model research.
