News: 1600088409

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Database consolidation is a server gain. Storage vendors should butt out

(2020/09/14)


Register Debate Welcome to The Register Debate in which we pitch our writers against each other on contentious topics in IT and enterprise tech, and you – the reader – decide the winning side. The format is simple: a motion is proposed, for and against arguments are published today, then another round of arguments on Wednesday, and a concluding piece on Friday summarizing the brouhaha and the best reader comments.

During the week you can cast your vote using the embedded poll, choosing whether you're in favor or against the motion. The final score will be announced on Friday, revealing whether the for or against argument was most popular. It's up to our writers to convince you to vote for their side.

For this debate, the motion is: [1]Consolidating databases has significant storage benefits, therefore everyone should be doing it .

And so, kicking off this week's debate and arguing AGAINST the motion, is CHRIS MELLOR...

The idea that consolidating databases has significant storage benefits and therefore everyone should be doing it is missing the point.

A database consolidation exercise could involve a migration from a disk array to an all-flash array. This could well deliver storage benefits, such as significantly lower data centre power and cooling needs, reduced rack space take-up and faster performance. But this should not be confused with the actual database consolidation - it’s database storage acceleration. All-flash array suppliers may well say it is worth doing. DRAM-caching disk array vendors could have a different view. But this is a database storage issue and quite separate from database consolidation.

Database consolidation is fundamentally about database server consolidation. And it’s made possible by server CPUs having many more cores than before.

Look at this progression on HPE ProLiant DL360 core counts; Gen 7 DL360s supported just 6 cores, Gen 8 DL360 has increased this to 12. Gen 9 took the core total to 22 and the latest Gen 10 DL360s supports up to 28 cores.

That means in general that a single later generation DL360 can do the same processing work as several earlier DL360s.

It’s not just cores that are multiplying. Storage devices are getting bigger. Disk drives are bulking up with 10, 12, 14, 16 and 18TB drives, and 20TB models are getting ready for prime time. SSDS are getting larger with increasing layers in 3D NAND; 64 moving to 96 layers with 128 layers now appearing, while the number of bits per cell increases from three (TLC) to four (QLC).

The PCIe bus which hooks up storage devices to memory is doubling speed from PCIe Gen 3 to Gen 4. Memory speed and capacity is increasing with a DDR4 to DDR5 transition. Individual servers are turning into far faster and more capable data processing powerhouses than those in use five years ago, with 10-year old servers possibly having less CPU power than the latest mobile phone CPUs.

Imagine a business’ relational database system running on a group of servers which access a shared storage area network (SAN) array across am 8Gbit/s Fibre Channel network. It could be an iSCSI network but the ideas we’re bringing out will stay the same.

The servers run the database application code and read and write database records fetched from or sent to the SAN.

Let’s suppose you have 20 four-core database servers accessing a 200TB SAN and consolidate them to eight 10-core servers. You have got rid of 12 servers, with their power draw, their cooling needs and their rack space. The eight left do all the work. Great. But they still access the 200TB SAN. The database storage resource has not changed its capacity at all. There’s no saving there. But there could well be added cost because it may well need to change its storage networking infrastructure.

If our consolidation exercise results in just eight servers doing the same work as the original twenty then they will each need 2.5 times the Fibre Channel bandwidth of a single original server. That means upgrading the SAN and server Fibre Channel HBAs and the Fibre Channel switches in the storage network from 8Gbit/s to faster 16Gbit/s. In this aspect of database consolidation there is likely to be increased storage networking cost.

Consolidation benefits

So database consolidation isn’t affecting storage capacity and may increase storage costs. The benefit is concentrated on the servers. This has been well known for years.

A July 2020 “Best Practices for Database Consolidation” Oracle paper

[2]PDF

states: “The primary purpose of database consolidation is to reduce the cost of database infrastructure by taking advantage of more powerful servers that have come onto the market that include dozens of processor cores in a single server. … Database consolidation allows more databases to run on fewer servers, which reduces infrastructure, but also reduces operational costs.”

The paper asserts: “Cost reduction is the primary business driver of database consolidation,” followed by simplified administration, improved security and isolation.

What the Oracle paper does not say is that consolidating databases such as Oracle’s or Microsoft’s SQL Server on fewer servers also saves licensing costs where database instances are licensed per server.

Database consolidation onto fewer servers saves server cost because you need fewer servers, and also saves database instance licensing expense as you need fewer per-server instance licenses.

There is no storage benefit here but the potentially significant server-based benefits make database consolidation an attractive idea that can serve you right. ®

Cast your vote below. You can track the progress of the debate [3]right here .

JavaScript Disabled Please Enable JavaScript to use this feature.

Get our [4]Tech Resources



[1] https://www.theregister.com/Debates/2020/09/14/database_storage_consolidation

[2] https://www.oracle.com/technetwork/database/availability/maa-consolidation-5648225.pdf

[3] https://www.theregister.com/Debates/2020/09/14/database_storage_consolidation

[4] https://whitepapers.theregister.com/

Consolidation is putting all your eggs in one basket

alain williams

Any breakage and nothing works.

It also means that the one database/cluster/... does all the work ie a higher workload than when it is distributed. One humongous machine might be more costly than several smaller ones -- maybe.

Have to agree wtih Chris...

Mike the FlyingRat

Here's a guy who actually knows what he's talking about.

But here's the thing... I don't know if the question is being framed properly.

When you say 'database consolidation, what do you mean exactly.

Yes, its a strange question, but think about it.

You have databases that are OLTP transaction processing systems of truth. Then you have Data Warehouses (OLAP) that are used to drive analytics.

Then you have Data Lakes which in itself is a Data Warehouse consolidation by removing the silos. (Here the number of DWs goes down, but the storage requirements go up. )

And its not just the CPUs getting better, or storage, but also networking. 40GbE is becoming Cisco's norm. 100GbE is also there...

But at 40GbE you can start to consider data fabric as your storage layer. The issue is cost versus density and performance has to be evaluated on a case by case basis.

The networking also allows for a segregation of COTS and specialty hardware to get the most bang for your buck. You can weave a GPU appliance into your data fabric and then consolidate compute servers using K8s to allow distributed OLTP RDBMs to take better advantage of the hardware. (This is where the network can be a bottleneck. )

What's interesting and a side note... when you look at this... its in *YOUR* Data Center. Not on the cloud. (Although it could be in the Cloud too.)

These advances will spell a down turn in the cloud over the next 5 years. Thats not to say that there won't be a reason for cloud but more of a hybrid approach.

Just some random thought from someone who's been around this for far too long but too broke to retire. ;-)

OK, but factor in the Redmond tax

chivo243

database ...as you need fewer per-server instance licenses, but as MS has also made this calculation, it costs more to run windows on these cost saving servers. Save in one column, pay in another.

Database consolidation? More like a Cambrian explosion.

PeterCorless

I'll agree with Mike the FlyingRat that "I don't know if the question is being framed properly." That this whole question seemed fuzzy so that "it depends" is the only right answer.

For specialized processing of data you are seeing a practical Cambrian explosion of different databases. Sure, you have your standard RDBMS for backoffice and operations work (ERP). But aside from that, a number of different SQL and/or NoSQL systems are being spun out into a constellation of special-purpose transactions and analytics platforms.

Just an example:

1. An IIoT NoSQL time series database (OTLP) that tracks all of a company's products in the field. You have that just because the RDBMS cannot deal with the direct rate of ingestion.

2. An Apache Spark analytics cluster that draws quality assurance results (OLAP) from the above time series database.

Now, this second database *might* be consolidated with #1 above if you use something like ScyllaDB, which allows workload prioritization. It requires a slightly larger consolidated cluster. It also means that you have to use the right tools to be drawing information and inferences out of your NoSQL database. Many data scientists would still rather get the data from a Spark cluster than do CQL queries. It's just more natural for them.

Then again, rather than Spark, you might have an Elastic cluster to do free searches. What *sort* of queries are you trying to do? Set, repeatable batch analytics, streaming analytics, or free text ad hoc queries?

Lacking methods for consolidation, you need to fall back on lambda architectures where you have a speed layer for OLTP (#1 above) and a serving layer for OLAP.

And also, as a reminder, there's no way your standard ERP system is keeping up with the raw rate of ingestion and analytics of IIoT. And no way the CFO is going to let quarter close be impacted because someone's trying to run an ad hoc data query on the ERP system. It is going to be a second and maybe a third or even fourth system system, though information can flow between them all if and as needed.

There can and will be multiple other systems out there. Real-time adtech/martech systems. These might have their own RDBMS, separate from your central corporate ERP, with a subset of user data.

Another part might be a graph data model, totally orthogonal to your RDBMS, to track 360º total touch of users across multiple social networks with your brand. The size of that graph alone would make your typical ERP DBAs grimace if you suggested trying to store those nodes and edges in your ERP. ("How?!?" they'll ask.) Hint: SQL ≠ Gremlin/Tinkerpop.

Yet another system might be a shopping cart system for your website. Sure, it might tie in closely to your ERP system, but for the purposes of keeping timeouts and abandonment low, it's built on a separate system to remain closer to your website than your back office. It will, of course, need to be able to traverse the firewall to get those orders into your ERP system for fulfillment. It may be SQL or NoSQL, depending on your use case.

Another database, or even a cluster of databases, may be used specifically for your R&D processes that are *not* the same as your corporate ERP system. There is NO WAY engineering wants to have to rely on a "please-mother-may-I" game with IT. They're going to do their own. You cannot stop them. The rest of the corporation won't even know what the acronyms and code words for these servers stand for. ("Good," the engineers will say.)

Other database(s) will track your own internal infrastructure. Every server and network device on your net will consolidate its logs for your CIO's needs. (Think Kafka feeding a NoSQL key-value store for log consolidation. For real-time uptime, latencies and throughput, Prometheus or Datadog for metrics, etc.)

And so on.

All of this is just your *on-prem* infrastructure. But what percentage of your infrastructure is now on the cloud? How much do you really know about what the back-end you are running on when you are getting the equivalent of a monthly time-share in a virtual server, or you are running APIs that are completely serverless anyway?

Every time someone says that there will be "one database to rule them all," I keep thinking, "Okay Sauron. That's just not how corporations work. Not even in 2000. Certainly not 2020." If anything, there's a reason for the explosion of database types and classes, and they all run on their own hardware because trying to make them all "one thing" is utterly absurd. Each has its own design patterns for consistency vs. availability, data distribution vs. consolidation, and so on.

The good news is that each of the database systems your corporation does run on will get faster and, as hardware generations roll out, cheaper overall to run the same workload. But don't be fooled entirely. You may save more money if you are being licensed "per server," but frankly many licenses are calculated "per core," in which case your higher densities will actual cost *more.*

Though, of course, if your database software is open source, you will not have to care one whit about your database cost as you scale.

Lastly, churning your on-prem servers just to get to higher densities has a capital cost expenditure which is never free. You might need to wait until you've reached the depreciation schedule on your current servers before funding opens up for new on-prem hardware. And then your boss might ask, with a hairy eyeball and a big spreadsheet, "So how does this compare to a public cloud, in terms of monthly spend?"

Disclosure: I work for ScyllaDB (scylladb.com). We make a monstrously fast NoSQL database, but we have to play well with an infinite number of other great systems.

If I had a formula for bypassing trouble, I would not pass it around.
Trouble creates a capacity to handle it. I don't say embrace trouble; that's
as bad as treating it as an enemy. But I do say meet it as a friend, for
you'll see a lot of it and you had better be on speaking terms with it.
-- Oliver Wendell Holmes, Jr.