Hi,
I have some experience using HDT with Jena. I think HDT is an amazing
technology and I've so far been happy with the performance, but as Rob
said, the use case matters a lot and benchmarking is recommended.
In my case I have a conversion pipeline [1] that converts a set of MARC
bibliographic records into a 30M triple NT file. Loading that into TDB
would probably take hours. Instead I'm converting it to HDT using the
hdt-cpp toolkit [2] (it's faster than the Java version and uses less
memory) and create an index file alongside the main HDT file. The
HDT+index files are a fraction of the size of the NT file (4GB NT file
vs. less than 500MB for the HDT+index).
I can then start up a version of Fuseki that exposes the data in the HDT
file as a read-only SPARQL endpoint. In my experience, query performance
is very reasonable, though I haven't benchmarked it against TDB. Since
the HDT file + index are rather small, they will soon be held mostly in
the disk cache, so although the technology is disk based, in practice
the disk will not be used very much unless you are extremely low on memory.
Running the conversion from NT to HDT, creating the index file, and
starting up Fuseki altogether take less than 5 minutes and SPARQL
queries can then be run immediately. In that time the TDB loader would
have barely started.
As an alternative to Fuseki, SPARQL queries can be run directly on the
HDT file using the hdtsparql command line tool from the hdt-jena toolkit.
-Osma
[1] https://github.com/NatLibFi/bib-rdf-pipeline
[2] https://github.com/rdfhdt/hdt-cpp
04.04.2017, 12:27, Rob Vesse kirjoitti:
HDT is primarily on disk. Whether it is query-able depends on the exact
encoding, there is one encoding designed primarily for transportation of data
and another designed for querying called HDT-FoQ aka focused on querying
In either case, there will be some memory usage as they do perform some
caching. They may also take advantage of memory mapped files similar to what
TDB does.
As far as comparisons with TDB I have never done any myself. For simplistic
queries, I would expect that HDT performs ok since from what I remember the
indexing is suitable for simple scans. However, for queries with any kind of
complexity i.e. Filters, negations, Joins etc. I would expect TDB to outperform
it and will scale far better.
But as I always point out on these kinds of questions your use case will
matter. If you think one solution will be better than the other for your use
case then you should benchmark that yourself. Generic benchmarking will only
tell you so much and give you a general indication of comparative performance.
Rob
On 04/04/2017 07:03, "Lorenz B." <[email protected]> wrote:
Well, I'm not that familiar with HDT, thus, I'm probably wrong. And I
saw right now that they also provide some kind of indexing concept.
Let's wait for response from Andy and/or Rob.
(In the meantime, I'll play around with HDT and Jena today to get some
more insights. )
>> Jena HDT is in-memory, right?
> Is it? I thought it was a on-disk, compressed, and query-able list of
quads...
>
--
Lorenz Bühmann
AKSW group, University of Leipzig
Group: http://aksw.org - semantic web research center
--
Osma Suominen
D.Sc. (Tech), Information Systems Specialist
National Library of Finland
P.O. Box 26 (Kaikukatu 4)
00014 HELSINGIN YLIOPISTO
Tel. +358 50 3199529
[email protected]
http://www.nationallibrary.fi