Hi,
I made an apples-to-apples comparison using my bibliographic data set.
The starting point is a NT file with 30M triples (unfortunately not yet
available to the public), gzipped into a 400MB file (uncompressed it
would be 4GB). I used my i3-2330M laptop with 8GB RAM and SSD.
Converting the dataset to HDT using rdf2hdt took 5 minutes and 15
seconds. Top memory usage was about 1.3GB. During the conversion the
rdf2hdt process used one CPU core for 100%, the gzip process an
additional 20-30% of another CPU core. The resulting HDT file is 250MB.
Creating the index file using hdtSearch took a further 25 seconds.
Memory usage was about 300MB with one CPU core at 100%. The index file
is 160MB.
I ran an example query that calculates the top 20 subjects (the ones
with most works about them). The query is included below. I ran it a few
times using hdtsparql and the execution time was 15.3-16.8 seconds.
Total wall clock time: 5:40 minutes
Total disk usage: 410MB
Fastest query: 15.3 seconds
Loading the same dataset to TDB using tdbloader2 took 11 minutes. CPU
usage was 110-180% and top memory usage was 1.3GB for the java process
that does the initial loading. Then came the sort processes that took
300% CPU and used up to 3.5GB memory. The resulting TDB directory size
is 2.7GB.
The example query took 12.9-14.7 seconds.
Total wall clock time: 11 minutes
Total disk usage: 2.7GB
Fastest query: 12.9 seconds
To summarize, generating a HDT with an index file is about twice as fast
as loading the data into TDB and uses less memory and CPU. Disk usage is
only 15% of what TDB uses. Query performance for this particular query
is about 20% slower with HDT than when using TDB.
-Osma
--example query--
PREFIX schema: <http://schema.org/>
SELECT ?ysoc (COUNT(DISTINCT ?w) AS ?count) WHERE {
?w schema:about ?ysoc .
FILTER(STRSTARTS(STR(?ysoc), 'http://www.yso.fi/onto/yso/'))
}
GROUP BY ?ysoc
ORDER BY DESC(?count)
LIMIT 20
--example query--
04.04.2017, 13:29, Osma Suominen kirjoitti:
04.04.2017, 13:10, Dave Reynolds kirjoitti:
Not to detract from HDT in anyway but we routinely load 25M triple file
sets (Turtle) to TDB in around 10mins on modest cloud VMs and rather
faster on local desktops with modern SSDs. So HDT might still have some
load speed benefits but at that scale it is less than 2x and not hours
v.s. minutes.
Right, sorry, I made a mistake in my estimate.
With my modest laptop (i3-2330M, SSD), loading the Geonames dataset
(173M triples NT file) into TDB using tdbloader2 takes about 70 minutes,
so the loading rate is about 40k triples per second. The size of the
resulting TDB is 16GB.
-Osma
--
Osma Suominen
D.Sc. (Tech), Information Systems Specialist
National Library of Finland
P.O. Box 26 (Kaikukatu 4)
00014 HELSINGIN YLIOPISTO
Tel. +358 50 3199529
[email protected]
http://www.nationallibrary.fi