Hi,

I have some experience using HDT with Jena. I think HDT is an amazing technology and I've so far been happy with the performance, but as Rob said, the use case matters a lot and benchmarking is recommended.

In my case I have a conversion pipeline [1] that converts a set of MARC bibliographic records into a 30M triple NT file. Loading that into TDB would probably take hours. Instead I'm converting it to HDT using the hdt-cpp toolkit [2] (it's faster than the Java version and uses less memory) and create an index file alongside the main HDT file. The HDT+index files are a fraction of the size of the NT file (4GB NT file vs. less than 500MB for the HDT+index).

I can then start up a version of Fuseki that exposes the data in the HDT file as a read-only SPARQL endpoint. In my experience, query performance is very reasonable, though I haven't benchmarked it against TDB. Since the HDT file + index are rather small, they will soon be held mostly in the disk cache, so although the technology is disk based, in practice the disk will not be used very much unless you are extremely low on memory.

Running the conversion from NT to HDT, creating the index file, and starting up Fuseki altogether take less than 5 minutes and SPARQL queries can then be run immediately. In that time the TDB loader would have barely started.

As an alternative to Fuseki, SPARQL queries can be run directly on the HDT file using the hdtsparql command line tool from the hdt-jena toolkit.

-Osma

[1] https://github.com/NatLibFi/bib-rdf-pipeline

[2] https://github.com/rdfhdt/hdt-cpp



04.04.2017, 12:27, Rob Vesse kirjoitti:
HDT is primarily on disk. Whether it is query-able depends on the exact 
encoding, there is one encoding designed primarily for transportation of data 
and another designed for querying called HDT-FoQ aka focused on querying

In either case, there will be some memory usage as they do perform some 
caching. They may also take advantage of memory mapped files similar to what 
TDB does.

 As far as comparisons with TDB I have never done any myself. For simplistic 
queries, I would expect that HDT performs ok since from what I remember the 
indexing is suitable for simple scans. However, for queries with any kind of 
complexity i.e. Filters, negations, Joins etc. I would expect TDB to outperform 
it and will scale far better.

But as I always point out on these kinds of questions your use case will 
matter. If you think one solution will be better than the other for your use 
case then you should benchmark that yourself. Generic benchmarking will only 
tell you so much and give you a general indication of comparative performance.

Rob

On 04/04/2017 07:03, "Lorenz B." <[email protected]> wrote:

    Well, I'm not that familiar with HDT, thus, I'm probably wrong. And I
    saw right now that they also provide some kind of indexing concept.

    Let's wait for response from Andy and/or Rob.

    (In the meantime, I'll play around with HDT and Jena today to get some
    more insights. )

    >> Jena HDT is in-memory, right?
    > Is it? I thought it was a on-disk, compressed, and query-able list of 
quads...
    >
    --
    Lorenz Bühmann
    AKSW group, University of Leipzig
    Group: http://aksw.org - semantic web research center








--
Osma Suominen
D.Sc. (Tech), Information Systems Specialist
National Library of Finland
P.O. Box 26 (Kaikukatu 4)
00014 HELSINGIN YLIOPISTO
Tel. +358 50 3199529
[email protected]
http://www.nationallibrary.fi

Reply via email to