TextVQA: A Benchmark for Reasoning About Text in Images

TL;DR
TextVQA is a visual question-answering benchmark built around text in real-world images. It contains 28,408 images and 45,336 question-image pairs from OpenImages.
Training and validation questions include 10 human-annotated answers, while test answers are withheld. The benchmark evaluates how well models combine optical character recognition, visual context, language understanding, and question answering.

What TextVQA Is
TextVQA, now available on AIOZ AI, is designed for models that need to read and reason about written information inside images.
It makes visual text essential to answering each question. Models must recognize text and understand how it relates to the surrounding scene.
The images come from OpenImages and cover varied real-world settings, layouts, and visual environments.
How the Dataset Is Organized
TextVQA includes:
- 28,408 images from OpenImages
- 45,336 question-image pairs
- 10 human-annotated answers per question for training and validation
- Test questions with answers withheld
Each example pairs an image with a question whose answer depends on reading and interpreting text in the scene.
What It Measures
TextVQA evaluates whether vision-language models can connect visible text with visual context to answer questions correctly.
- OCR with visual question answering
- Text understanding in natural images
- Visual-text reasoning
- Text detection and recognition
- Benchmarking vision-language models beyond object recognition
The challenge goes beyond reading text: models must interpret what the text means within the image and question.
Key Dataset Details
- Dataset: TextVQA
- Dataset type: Text-based visual question-answering benchmark
- Image source: OpenImages
- Images: 28,408
- Question-image pairs: 45,336
- Training and validation annotations: 10 human answers per question
- Test annotations: Answers are not provided
- Core tasks: Optical character recognition, visual question answering, and multimodal reasoning
- License: Creative Commons Attribution 4.0 International
What Builders Should Know
TextVQA measures a specific form of visual understanding in which written information is necessary to answer a question. Results should therefore be interpreted as evidence of visual-text reasoning rather than a complete measure of image understanding.
Because test answers are not released, builders who need direct answer labels should use the training and validation splits during development.
Access TextVQA on AIOZ AI
With 28,408 real-world images, 45,336 question-image pairs, and human answer annotations, TextVQA provides a practical benchmark for models that need to read and reason about text in visual scenes.
Access TextVQA on AIOZ AI and start evaluating visual-text understanding.
FAQ
Q1: What is TextVQA used for?
It is used to evaluate whether vision-language models can read text in real-world images, understand its context, and answer related questions.
Q2: How large is the TextVQA dataset?
It contains 28,408 images from OpenImages and 45,336 question-image pairs.
Q3: How are answers provided in TextVQA?
Training and validation questions include 10 human-annotated answers each, while test answers are not provided.