News: 1678186807

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

The Great Graph Database Debate: Relational can't do everything

(2023/03/07)


Register Debate Welcome back to the latest [1]Register Debate in which writers discuss technology topics, and you the reader choose the winning argument.

The format is simple: we propose a motion, the arguments for the motion run on [2]Monday and Wednesday, and the arguments against run today and Thursday. During the week you can cast your vote on which side you support using the poll embedded below, choosing whether you're in favor or against. The final score will be announced on Friday, revealing which argument was most popular.

It's up to our writers to convince you to vote for their side.

[3]

This week's motion is:

[4]

[5]

[6]Graph databases – in which relationships are stored natively alongside the data elements – do not provide a significant advantage over well-architected relational databases for most of the same use cases.

It has been roughly 20 years since the first production deployment of Neo4j, one of the leading protagonists in the graph database story.

[7]

Strong market growth and interest from investors suggest it might be catching up with the rows and columns of RDBMSes, owing to its analysis of data according to networked relationships we see all around us: in business, media, society, medicine and science.

But detractors still have their doubts, suggesting that the benefits graph systems seem to offer can be created in relational systems, which have a longer history – and are arguably more mature and easier to manage – than their graph counterparts.

Jim Webber , our first contributor arguing AGAINST the motion, is Neo4j's chief scientist and a visiting professor of computer science at Newcastle University.

Relational cannot deal with every use case

We agree with many of the house's assertions about the usefulness of graph APIs. The crux of our disagreement is simply with the claim that some future "well-architected" relational database engine could render the use of today's useful, existing, in-production graph databases unnecessary. The adjective "well-architected" does a lot of heavy lifting here to make that speculation as barely credible as it is.

[8]

Jim Webber: The idea that relational is a universal DB option is spurious

We should also point out that the idea that relational can deal with any use case and is a universal database option, is spurious. The notion of "one size fits all" has been [9]dismissed [PDF] by one of the most influential theorists of the relational school, Michael Stonebraker. Stonebraker argues that distinct workloads, even within the purview of relational databases, require different data engines to work well. Within Neo4j we are happy to embrace this specialization as the graph database has a different engine from the graph analytics platform. In fact, there is a rich tradition of doing so in [10]Document , [11]Time series and, yes – [12]relational databases [PDF].

We cannot accept that relational can do anything. Neither can we accept "[graph] systems ignore many of the hard-learned lessons… from the last 50 years" as if graphs exist outside of day-to-day operational data management. Yet graph databases have consistently tackled transactions, query planning/execution, indexing, consensus, replication and concurrency control using a mixture of tried and tested techniques.

[13]

We're glad that the house sees "benefits from... graph APIs" and query-oriented query languages. This is why we built Cypher, [14]a declarative language [PDF] with formally described semantics. We are engaged with academia and industry to define GQL, an ISO- [15]standardized graph counterpart to SQL based on Cypher, by the same ISO committee that standardized SQL. GQL will allow huge numbers of new users to succinctly and humanely express graph queries that are cumbersome and error-prone in SQL.

We reject the assertion that graph databases cannot properly support views and migrations, and therefore lack (a narrow definition of) data independence. Migrations are frequently performed in practice (as in [16]this example ), while GQL includes native syntax for defining graph views. In fact, the schema-optional nature of many graph databases and the fuzzy pattern-matching abilities of their query languages means they are better off than others with respect to data independence.

But databases aren't just about theory. They are complex systems designed to perform some of the most demanding tasks in computing. It is not sufficient to offer vague notions of "well-architected" implying previous databases have been poorly architected. Many thousands of graph production systems in use today certainly haven't chosen graphs out of ignorance, but because the graph model and performance was the best engine for their specific requirements.

As to the claim that you can just write SQL that does graph work, heterogeneous graphs have relationships between nodes that are many, varied, and bi-directional. They are sometimes regular, sometimes not; they are sometimes sparse and sometimes dense. The house's notion that graphs should be modelled as a collection of tables isn't practical. We know this only too well because this is how Neo4j began: by using a graph API atop a relational database, we ended up going against the grain with exploding complexity and dwindling performance. It is that precise technological gap that forced us to build an engine that could process graphs natively, not a perverse will to build our own database!

The fact is that for graph use cases common in the wild, native storage engines often have [17]very significant run-time benefits , which include increased data locality, better concurrency support and reduced space usage to name just a few. Some cherry-picked studies from 5+ years ago which largely focus on analytical workloads over homogeneous graphs do not change this reality.

We have had 50 years to understand relational databases. We have come to respect their utility and understand their limitations. Accordingly, we argue that the "well-architected" processing engine proposed by the house is already here. And it's called a graph database. ®

Cast your vote below. We'll close the poll on Thursday night and publish the final result on Friday. You can track the debate's progress [18]here .

JavaScript Disabled Please Enable JavaScript to use this feature.

Get our [19]Tech Resources



[1] https://www.theregister.com/Debates/

[2] https://www.theregister.com/2023/02/27/great_graph_debate_monday

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZAdtvAFLLMizls9nljUp@AAAAEc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZAdtvAFLLMizls9nljUp@AAAAEc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZAdtvAFLLMizls9nljUp@AAAAEc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/Debates/2023/03/06/great_graph_debate/

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZAdtvAFLLMizls9nljUp@AAAAEc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://regmedia.co.uk/2023/02/22/jim_webber.jpg

[9] https://cs.brown.edu/~ugur/fits_all.pdf

[10] https://github.com/wiredtiger/wiredtiger

[11] https://docs.influxdata.com/influxdb/v1.8/concepts/storage_engine/

[12] https://vldb.org/pvldb/vol13/p3217-matsunobu.pdf

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZAdtvAFLLMizls9nljUp@AAAAEc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://homepages.inf.ed.ac.uk/libkin/papers/sigmod18.pdf

[15] https://www.iso.org/standard/76120.html

[16] https://neo4j.com/labs/neo4j-migrations/

[17] https://www.vldb.org/pvldb/vol15/p1173-fuchs.pdf

[18] https://www.theregister.com/Debates/2023/03/06/great_graph_debate/

[19] https://whitepapers.theregister.com/



Tom7

The crux of our disagreement is simply with the claim that some future "well-architected" relational database engine could render the use of today's useful, existing, in-production graph databases unnecessary

Not a great way to start, but mis-quoting (apparently deliberately) the other side. No-one has so far mentioned some future "well-architected relational database engine" but rather well-architected databases (ie schemas) within existing relational database engines.

This, and a big pile of thinly-veiled neo4j sales-speak, appears to be about the level this contribution to the debate is operating on.

stiine

He's also never heard of indexes to say nothing of multiple indexes. Perhaps its because a graph database is soley indexes.

Knock knock

elsergiovolador

*knock knock*

Someone: who is there?

Salesman: would you like to talk about our saviour graph database?

Someone: sure, would you like to come in?

Salesman: *confused* I am not sure I've never made it this far.

YAQL is not the solution

Charlie Clark

Much I dislike SQL, the presence of SPARQL, GQL, etc. which try to provide something like it, show the need for a query language. But all the attempts stem from building something for a different storage engine rather than a better query language, and that is an indication of a greater flaw in the approaches: confusing the model with the implementation. But there is nothing in the relational model that requires data be stored in row tuples. Hence, any conclusions that derived from that are spurious.

in fact, rather than spending even more on licences, I think a lot of problems could be solved by ensuring that people dealing with data are properly supported by RDBMS staff so that models provide the data they should, as fast as is necessary and without breaking the bank. The persistent fad for map/reduce (denormalise and then parallelise) because hardware is so cheap is just leading to massive bills for stuff running on clouds.

Use-case

andy 103

Consider these 2 (very common in the real world) use-cases:

1. People using relational databases with all the complexity of MySQL, Postgres, MariaDB etc.... for tiny little web applications like Wordpress or Magento.

2. Companies harvesting a metric fuckton of your data to do nothing except track / advertise shit to you. These tend to go more with the graph database option over relational.

The thing I'm struggling with though is for all the complexity afforded by these database platforms, the actual tangible output and use-case of the systems, is... boring.

If we look at some other apps and services the choice of database is quite moot when you actually consider how worthwhile the end application tends to be. It would be a bit like over-engineering a house to the point where it could withstand a war zone, whilst ignoring the fact nobody actually wants to live there.

TL/DR: it doesn't really matter which you pick because your application for that database is incredibly simple. The database platform used isn't the issue.

Re: Use-case

Tim99

For (1) - If you really need a small relational database SQLite is your friend. Whilst it is "a single user database", a web server with multiple clients is "a single user". With WAL (write ahead logging) I’ve just prototyped something with multiple related tables (using that) running a multiuser test at ~120 added rows per second on a 2GB Raspberry pi 4. Eventually it got to >5 million rows. I was pleasantly surprised to find that the sqlite3 command line shell could successfully back it up into another SQLite data file whist it was running at close to that rate…

Re: Use-case

elsergiovolador

For (1) I don't think you even need a database at all. You can just use the filesystem and construct filenames as if you were using something like dynamodb and call it a day.

Groo The Wanderer

The problem I've found in the industry over the years is companies with a "standard database" policy, or lead designers who have a "preferred" database that they always use.

It is best to keep an open mind with databases, and use the provider that is best suited to the task at hand. There is no one database model that "does it all" easily for all use cases; each of the main architectures (relational, graph, object store, and document) has advantages for certain types of data access. Trying to shoehorn out-of-band data into the wrong database engine is not only painful and inefficient, it is prone to bugs and breakage.

I'd rather go to the hassle of setting up two phased commits so I could use the right database for each type of data to be managed than try to stuff it all in one general purpose engine.

So you think that money is the root of all evil. Have you ever asked what
is the root of money?
-- Ayn Rand