XStoryCloze: A 10-Language Benchmark for Narrative Understanding

XStoryCloze: A 10-Language Benchmark for Narrative Understanding

TL;DR

XStoryCloze is a multilingual benchmark designed to evaluate narrative understanding across 10 non-English languages. Each language includes 360 training examples and 1,510 test examples, professionally translated to support reliable cross-lingual comparison.

What XStoryCloze Is

XStoryCloze, now listed on AIOZAI, is a benchmark for measuring how multilingual language models understand stories.

It tests whether a model can follow narrative flow, interpret context, and understand relationships between characters and events.

The dataset extends narrative evaluation beyond English through professionally translated examples in ten languages.

How the Benchmark Is Organized

XStoryCloze follows the same split structure for every supported language:

  • train: 360 examples for model preparation and experimentation
  • test: 1,510 examples for evaluating narrative understanding

Using consistent split sizes across all ten languages makes it easier to compare model performance without building a separate evaluation setup for each language.

What It Measures

Teams can use XStoryCloze to study:

  • Plot and event progression
  • Story context
  • Character interactions
  • Narrative coherence
  • Commonsense relationships between events
  • Zero-shot and few-shot multilingual performance
  • Cross-lingual transfer in narrative reasoning

XStoryCloze includes ten non-English languages: Russian, Chinese, Spanish, Arabic, Hindi, Indonesian, Telugu, Swahili, Basque, and Burmese.

Key Dataset Details

  • Dataset: XStoryCloze
  • Dataset type: Multilingual narrative-understanding benchmark
  • Languages: 10 non-English languages
  • Training examples: 360 per language
  • Test examples: 1,510 per language
  • Translation method: Professional translation
  • Evaluation focus: Plot, context, character interactions, coherence, and commonsense
  • Intended settings: Zero-shot and few-shot evaluation
  • License: Creative Commons Attribution 4.0 International (CC BY 4.0)

What Builders Should Know

XStoryCloze is designed for narrative understanding, so its results should not be treated as a complete measure of multilingual model quality. It does not directly evaluate factual knowledge, instruction following, coding, or other language capabilities.

Each language includes 360 training examples, making the train split suitable for preparation and experimentation rather than large-scale model training.

Access It on AIOZ AI

10 languages, professional translations, and one consistent evaluation structure: XStoryCloze provides a practical benchmark for comparing multilingual narrative understanding.

Access XStoryCloze on AIOZ AI and start evaluating story-level reasoning across languages.

FAQ

Q1: What is XStoryCloze used for?

It is used to evaluate how multilingual language models understand narrative flow, context, character interactions, coherence, and commonsense relationships between events.

Q2: Which languages does XStoryCloze include?

It includes Russian, Chinese, Spanish, Arabic, Hindi, Indonesian, Telugu, Swahili, Basque, and Burmese.

Q3: How many examples are included for each language?

Each language contains 360 training examples and 1,510 test examples.