Not to detract from HDT in anyway but we routinely load 25M triple file sets (Turtle) to TDB in around 10mins on modest cloud VMs and rather faster on local desktops with modern SSDs. So HDT might still have some load speed benefits but at that scale it is less than 2x and not hours v.s. minutes.

Dave

On 04/04/17 10:56, Osma Suominen wrote:
Hi,

I have some experience using HDT with Jena. I think HDT is an amazing
technology and I've so far been happy with the performance, but as Rob
said, the use case matters a lot and benchmarking is recommended.

In my case I have a conversion pipeline [1] that converts a set of MARC
bibliographic records into a 30M triple NT file. Loading that into TDB
would probably take hours. Instead I'm converting it to HDT using the
hdt-cpp toolkit [2] (it's faster than the Java version and uses less
memory) and create an index file alongside the main HDT file. The
HDT+index files are a fraction of the size of the NT file (4GB NT file
vs. less than 500MB for the HDT+index).

I can then start up a version of Fuseki that exposes the data in the HDT
file as a read-only SPARQL endpoint. In my experience, query performance
is very reasonable, though I haven't benchmarked it against TDB. Since
the HDT file + index are rather small, they will soon be held mostly in
the disk cache, so although the technology is disk based, in practice
the disk will not be used very much unless you are  extremely low on
memory.

Running the conversion from NT to HDT, creating the index file, and
starting up Fuseki altogether take less than 5 minutes and SPARQL
queries can then be run immediately. In that time the TDB loader would
have barely started.

As an alternative to Fuseki, SPARQL queries can be run directly on the
HDT file using the hdtsparql command line tool from the hdt-jena toolkit.

-Osma

[1] https://github.com/NatLibFi/bib-rdf-pipeline

[2] https://github.com/rdfhdt/hdt-cpp



04.04.2017, 12:27, Rob Vesse kirjoitti:
HDT is primarily on disk. Whether it is query-able depends on the
exact encoding, there is one encoding designed primarily for
transportation of data and another designed for querying called
HDT-FoQ aka focused on querying

In either case, there will be some memory usage as they do perform
some caching. They may also take advantage of memory mapped files
similar to what TDB does.

 As far as comparisons with TDB I have never done any myself. For
simplistic queries, I would expect that HDT performs ok since from
what I remember the indexing is suitable for simple scans. However,
for queries with any kind of complexity i.e. Filters, negations, Joins
etc. I would expect TDB to outperform it and will scale far better.

But as I always point out on these kinds of questions your use case
will matter. If you think one solution will be better than the other
for your use case then you should benchmark that yourself. Generic
benchmarking will only tell you so much and give you a general
indication of comparative performance.

Rob

On 04/04/2017 07:03, "Lorenz B." <[email protected]>
wrote:

    Well, I'm not that familiar with HDT, thus, I'm probably wrong. And I
    saw right now that they also provide some kind of indexing concept.

    Let's wait for response from Andy and/or Rob.

    (In the meantime, I'll play around with HDT and Jena today to get
some
    more insights. )

    >> Jena HDT is in-memory, right?
    > Is it? I thought it was a on-disk, compressed, and query-able
list of quads...
    >
    --
    Lorenz Bühmann
    AKSW group, University of Leipzig
    Group: http://aksw.org - semantic web research center








Reply via email to