Querying knowledge graphs in natural language

Shiqi Liang; Kurt Stockinger; Tarcisio Mendes de Farias; Maria Anisimova; Manuel Gil

doi:10.1186/s40537-020-00383-w

Querying knowledge graphs in natural language

J Big Data. 2021;8(1):3. doi: 10.1186/s40537-020-00383-w. Epub 2021 Jan 6.

Authors

Shiqi Liang¹, Kurt Stockinger², Tarcisio Mendes de Farias^{3

4}, Maria Anisimova^{2

3}, Manuel Gil^{2

3}

Affiliations

¹ ETH Swiss Federal Institute of Technology, Rämistrasse 101, 8092 Zurich, Switzerland.
² Zurich University of Applied Sciences, Obere Kirchgasse 2, 8400 Winterthur, Switzerland.
³ SIB Swiss Institute of Bioinformatics, Quartier Sorge-Bâtiment Amphipôle, 1015 Lausanne, Switzerland.
⁴ Department of Ecology and Evolution, University of Lausanne, Quartier Sorge-Bâtiment Biophore, 1015 Lausanne, Switzerland.

Abstract

Knowledge graphs are a powerful concept for querying large amounts of data. These knowledge graphs are typically enormous and are often not easily accessible to end-users because they require specialized knowledge in query languages such as SPARQL. Moreover, end-users need a deep understanding of the structure of the underlying data models often based on the Resource Description Framework (RDF). This drawback has led to the development of Question-Answering (QA) systems that enable end-users to express their information needs in natural language. While existing systems simplify user access, there is still room for improvement in the accuracy of these systems. In this paper we propose a new QA system for translating natural language questions into SPARQL queries. The key idea is to break up the translation process into 5 smaller, more manageable sub-tasks and use ensemble machine learning methods as well as Tree-LSTM-based neural network models to automatically learn and translate a natural language question into a SPARQL query. The performance of our proposed QA system is empirically evaluated using the two renowned benchmarks-the 7th Question Answering over Linked Data Challenge (QALD-7) and the Large-Scale Complex Question Answering Dataset (LC-QuAD). Experimental results show that our QA system outperforms the state-of-art systems by 15% on the QALD-7 dataset and by 48% on the LC-QuAD dataset, respectively. In addition, we make our source code available.

Keywords: Knowledge graphs; Natural language processing; Query processing; SPARQL.

Grants and funding

UL1 TR002014/TR/NCATS NIH HHS/United States