The Natural Questions corpus, a question answering data set, is presented, introducing robust metrics for the purposes of evaluating question answering systems; demonstrating high human upper bounds on these metrics; and establishing baseline results using competitive methods drawn from related literature.
We present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature.
Jakob Uszkoreit
7 papers
Chris Alberti
8 papers
T. Kwiatkowski
4 papers
J. Palomaki
3 papers
Olivia Redfield
1 papers
Michael Collins
7 papers
Ankur P. Parikh
5 papers
D. Epstein
1 papers
I. Polosukhin
4 papers
Jacob Devlin
8 papers
Kenton Lee
10 papers
Kristina Toutanova
7 papers
Llion Jones
5 papers
Ming-Wei Chang
11 papers
Slav Petrov
5 papers