News: 1679398215

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Ex-BigQuery exec and Motherduck CEO: For some users, the answer is to think small

(2023/03/21)


Interview Jordan Tigani made his name as an engineer leading the team behind BigQuery, Google’s data warehouse, which was among a group of systems to transform the market by separating storage and compute. Despite helping win global customers including Vodafone, and taking on big-name rivals Snowflake, Azure Synapse and AWS Redshift, he’s begun to think the approach has run out of “magic beans” for some users.

Now founder and CEO of [1]MotherDuck – which built a serverless analytics system based around DuckDB – he finds himself eschewing the virtues of scale-out systems for the benefits a lightweight in-process OLAP database affords. You can catch all of this detail and more in our interview with him below.

Tagani recalled coming across DuckDB while he was chief product officer at SingleStore — his first job after Google.

“We started seeing DuckDB appearing on performance reports and giving us a run for the money. Not beating us, but it was still surprising. Then I started poking at it and I realised it could do some stuff that we couldn't do in SingleStore, that Snowflake couldn’t do either, and that was pretty interesting,” he says.

The brainchild of academics at Amsterdam's Centrum Wiskunde & Informatica mathematical and theoretical computing research center, [2]DuckDB is embedded within a host process . There is no DBMS server software to install, update or maintain. For example, the DuckDB Python package can run queries directly on data in Python software library Pandas without importing or copying data. Written in C++, DuckDB is free and open source under the MIT License.

[3]

Tagani’s pitch to co-authors Hannes Mühleisen and Mark Raasveldt was to build a cloud-based serverless product around DuckDB.

[4]

[5]

The important thing about DuckDB is that it's a scale-up system, contrary to the received wisdom of the last 15 years of analytics, embodied in the first generation of so called big data systems based on the Hadoop Distributed File System. The scale-out approach led to the boom in cloud-based analytics systems like Snowflake, BigQuery, Synapse and Redshift.

[6]The nodes have it in the Great DB debate: Reg readers pick graph

[7]Cassandra 4.1 promises dev guardrails and pluggable storage

[8]MotherDuck scores $47.5m to prove scale-up databases are not quackers

[9]Google wants to copy-paste your mainframe applications into its cloud

“At some point, there are no magic beans left. There will be a convergence around the performance of all these systems. Snowflake, Redshift, BigQuery, and Synapse are all probably within a factor of two in performance right now. For the most part, that's not what should drive people to use one of these systems versus the other,” Tagani says.

Added to that, the server machinery and infrastructure around these systems can “be really quite crufty,” he adds.

“DuckDB has been able to kind of strip all that away by being an in-process database, and that means that you basically can marshal data in and out of your application, or your data frames, with the minimum of data movements,” he says.

[10]

DuckDB is not designed to replace company-wide enterprise data warehouse systems like BigQuery but — by courting data scientists, machine learning engineers and developers, it could still have plenty of space to fly. ®

Get our [11]Tech Resources



[1] https://www.theregister.com/2022/11/17/475_million_says_scaleup_databases/

[2] https://www.theregister.com/2022/09/09/duckdb_0_5_0/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZBnisvxCtI@SQX@I1enc0gAAAMQ&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZBnisvxCtI@SQX@I1enc0gAAAMQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZBnisvxCtI@SQX@I1enc0gAAAMQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2023/03/10/great_graph_debate_friday/

[7] https://www.theregister.com/2022/12/09/cassandra_41_promises_dev_guardrails/

[8] https://www.theregister.com/2022/11/17/475_million_says_scaleup_databases/

[9] https://www.theregister.com/2022/10/12/google_cloud_dual_run_service/

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_specialfeatures/spotlightondatabases&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZBnisvxCtI@SQX@I1enc0gAAAMQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[11] https://whitepapers.theregister.com/



Embarassing enough to be American already...

wub

...but does the complete Americanization of ElReg have to include typos in the subheads? I'm pretty sure that was supposed to be "big data" not "bid data".

Ready, and frankly anxious to be found wrong on this one.

I'm starting to think the gene pool could use a little chlorine.