News: 1656433808

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Databricks promises cheap cloud data warehousing at an eighth of the cost of rivals

(2022/06/28)


Databricks, the company born out of the Apache Spark boom, has let loose a raft of updates at its San Francisco conference, including an elastic compute option for analytics.

Databricks SQL Serverless, available in preview on AWS, has been designed to improve query performance and concurrency of BI and analytics workloads on messy data lake repositories.

The move is part of the company's plan to bring data lakes and data warehouses together on one system: the proverbial “lakehouse”, the coinage du jour achieving currency among vendors and commentators alike.

[1]

Databricks is also announcing an update to Photon, its query engine for lakehouse systems, making it available in Databricks Workspaces — the environment where users view their Databricks assets. It is releasing Open source connectors for Go, Node.js, and Python to simplify to access a lakehouse from operational applications. Meanwhile, query federation in Databricks SQL is set to let users query data in PostgreSQL, MySQL, AWS Redshift from Databricks.

[2]

[3]

Announced [4]last year , Databricks SQL Serverless is designed to provide instant compute to users for their BI and SQL workloads. The company promised “minimal management” and capacity optimizations to lower overall cost by an average of 40 per cent.

Joel Minnick, Databricks marketing VP, told The Register that SQL Serverless will allows users coming into Databricks to “go from start up to query in three seconds”.

[5]

It also makes economic sense in the cloud, he said. “You truly only pay for what you use on a data warehouse and workloads. That is a real game-changer in terms of the cost of getting these workloads done.”

Minnick said the nearest cloud data warehouse rival would be eight times more expensive than Databricks for these types of workloads.

[6]NoSQL player Aerospike links up with Starburst for SQL-based access to edge data

[7]MongoDB wants to grab work from other databases

[8]Google Cloud previews new BigLake data lakehouse service

[9]Tiny Uber offshoot tries to do for data lakes what Snowflake did for data warehousing

The serverless option would also help address the issue of user concurrency on analytics queries, where [10]data lakes have attracted criticism .

Minnick said Databricks had already made progress on this issue, and the vast majority of enterprise data warehousing concurrency needs would be met by Databricks SQL, he said.

Hyoun Park, CEO and chief analyst at Amalgam Insights says Databricks' Serverless SQL makes it easier to support large amounts of distributed data in a cost-effective manner. “It is a response to other vendors providing serverless SQL offerings such as Azure SQL or CockroachDB, but this should also allow Databricks customers to more easily support multi-region and hybrid multi-cloud environments. From a practical perspective, this move makes it easier for potential Databricks customers to use a lakehouse without the significant challenges of manual resource management that can potentially occur as data gets bigger and faster from many different sources to multiple different destinations.”

[11]

Park says Serverless SQL could also help users solve the concurrency challenge, but more work would be required to address it in the future.

“Realistically, this concurrency issue will probably require a bit of compromise: Databricks customers should ideally partition and structure data across multiple instances to avoid concurrency issues and Databricks will need to continue working on methods to accelerate queries, duplicate resources, cache results, and provide other resource and analytic workarounds as Databricks is used on smaller amounts of data.”

However he cautions that while Databrick's claims on price-performance compared with rivals could stand up, it would depend on a lot of different variables. "It's hard to tell if they're making an apples-to-apples comparison without license and hardware and labor specs. I'd advise users to do their own total cost of ownership analysis to support this claim."

Since it [12]introduced the lakehouse concept in 2020 , Databricks has seen some competition spring up.

In April, [13]Google announced a preview on Google Cloud of BigLake, a data lake storage service that it claims can remove data limits by combining data lakes and data warehouses.

In Feb this year, tiny [14]Californian startup Onehouse won $8m in seed funding with hopes to grow a business worthy of taking on the giants of data engineering. The other aim is to make data lake projects faster, cheaper and easier than before. And not to be outdone, Snowflake too has announced support for [15]unstructured data in its data warehouse platform.

Park said the challenge for Databricks is quickly scaling in areas where there was already massive competition. “For instance, shifting to build apps on Databricks when embedded BI and analytics have been around for decades is a significant process shift,” he said.

“Although the new generation of analytics is often reduced to a "Databricks vs. Snowflake" matchup, this simplistic view ignores the practical use cases for each vendor differ based on their historical approaches to data, open-source, analytic processing, and semi-structured data.

“Although enterprise data demands are pushing Databricks and Snowflake product maps closer together, Databricks is a platform fundamentally better suited to the future of real-time analytic data across multiple varieties of data structures and formats while Snowflake is well structured for fast adoption based on the current era of rapid use of structured data for data marts and warehouses,” he said.

Park said the market for software vendors addressing analytic data problems was in a “Cambrian explosion” phase. “As these shifts occur, it would not be surprising to see new vendors arise to take on both Databricks and Snowflake by solving problems associated with hybrid cloud, multi-modal data, networking, storage, compute, low-code application development, or administration,” he says. ®

Get our [16]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Yrt6AD1bkb@xaQ5xC3@mYgAAAE8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yrt6AD1bkb@xaQ5xC3@mYgAAAE8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yrt6AD1bkb@xaQ5xC3@mYgAAAE8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[4] https://databricks.com/blog/2021/08/30/announcing-databricks-serverless-sql.html

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yrt6AD1bkb@xaQ5xC3@mYgAAAE8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2022/06/15/nosql_realtime_aerospike_links_up/

[7] https://www.theregister.com/2022/06/13/mongodb_foray_into_analytics_gets/

[8] https://www.theregister.com/2022/04/06/google_previews_biglake_data_lakehouse/

[9] https://www.theregister.com/2022/02/08/onehouse/

[10] https://www.theregister.com/2021/05/24/data_lakes_struggle_with_sql_gartner/

[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yrt6AD1bkb@xaQ5xC3@mYgAAAE8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[12] https://www.theregister.com/2020/02/25/databricks_touts_new_ingestion_files_sources/

[13] https://www.theregister.com/2022/04/06/google_previews_biglake_data_lakehouse/

[14] https://www.theregister.com/2022/02/08/onehouse/

[15] https://www.theregister.com/2021/11/16/snowflake_python_support/

[16] https://whitepapers.theregister.com/



Anonymous Coward

"Park said the market for software vendors addressing analytic data problems was in a “Cambrian explosion” phase. “As these shifts occur, it would not be surprising to see new vendors arise to take on both Databricks and Snowflake..."

Surprising? It'd be totally shocking mate. The market's crashed and the PE money fountain is over. Nobody is going to be seriously challenging Databricks or Snowflake on matters of scale or breadth. Databricks have closed over $3.5Bn in pre-IPO funding, Snowflake closed something like $2.1Bn then did their probably-never-to-be-matched IPO. Both of them are multiple years of development and cash burn away from their desired steady state in growth and feature terms. Nobody's raising the kind of capital today they'd need to compete with either one of them, let alone both of them *and* AWS, Azure and GCP all at the same time.

There will be an explosion of vendors running on those platforms or alongside them, but not competing with them. Not for a long time. Running cloud control planes and "serverless" infrastructure and building comprehensive data platform feature sets is just too expensive.

I pity the bit barn owners

Mellipop

True. Competing on price is a sure indication that a market is saturated. The product has become a commodity.

It's also incredibly wasteful of resources.

It is possible the "cambrian explosion" will not happen in data centres, but where the data is being generated; IoT.

Beauty:
What's in your eye when you have a bee in your hand.