Abstract

RDF triplestores’ ability to store and query knowledge bases augmented with semantic annotations has attracted the attention of both research and industry. A multitude of systems offer varying data representation and indexing schemes. However, as recently shown for designing data structures, many design choices are biased by outdated considerations and may not result in the most efficient data representation for a given query workload. To overcome this limitation, we identify a novel three-dimensional design space. Within this design space, we map the trade-offs between different RDF data representations employed as part of an RDF triplestore and identify unexplored solutions. We complement the review with an empirical evaluation of ten standard SPARQL benchmarks to examine the prevalence of these access patterns in synthetic and real query workloads. We find some access patterns, to be both prevalent in the workloads and under-supported by existing triplestores. This shows the capabilities of our model to be used by RDF store designers to reason about different design choices and allow a (possibly artificially intelligent) designer to evaluate the fit between a given system design and a query workload.

Original languageEnglish
JournalThe VLDB Journal
Volume31
Issue number2
Pages (from-to)347-373
Number of pages27
ISSN1066-8888
DOIs
Publication statusPublished - 21 Jan 2022

Keywords

  • Data representation
  • Database
  • Knowledge graphs
  • Query
  • RDF
  • SPARQL

Fingerprint

Dive into the research topics of 'A design space for RDF data representations'. Together they form a unique fingerprint.

Cite this