IT systems capacity planning. This is hard ... but how hard? Inquiring minds wish to know
- Reference: 1637152205
- News link: https://www.theregister.co.uk/2021/11/17/survey_it_systems_capacity_planning/
- Source link:
Right-click -> Add storage. Right-click -> Add RAM .
Job done.
[1]
Which is fine, but it leads us into temptation – we don’t do capacity planning because the need to do so feels like it has gone away.
[2]
[3]
This is the case all through IT, of course. We get away with designing algorithms poorly because today’s ultra-fast CPU cores save our bacon through sheer speed. We don’t index our databases properly because solid-state storage rescues us when our queries do full table scans. The thing is, though, we get away with this approach most of the time, but definitely not all of the time.
In this survey – see below – we’re keen to find out the extent to which our readers have had to cope with changes in demand for capacity in their systems and, more importantly, how they have managed the capacity planning process. Many of us have had to scale up systems – particularly things like virtual desktop and VPN services – due to users being sent home to work during the COVID-19 lockdowns.
[4]
But some organizations will have kept capacity at roughly the same levels, and it’s likely that some have scaled down – perhaps through exploiting opportunities to finally get around to decommissioning resource-hungry legacy systems.
Systems perform great in the test environment but then tank when put live – often because the production database was ten times the size of the test one
We’re also interested in the science of performance and capacity planning. Most of us have come across systems that performed great in the test environment but then tanked when put live – often because the production database was ten times the size of the test one – but did we do anything to predict that?
Did we ask the users whether the app felt snappy enough during testing? Did we, for that matter, run up any electronic measures of performance and resource usage, or perhaps simulate the actions of hundreds of users with automation tools?
This correspondent was a performance tester in a previous life, and I can confirm how good it feels to know that the app will scale to 250 users thanks to the stats gathered by the test harness that simulated 250 users hammering it at once. And after go-live, did we keep asking the users and/or carry on with our electronic monitoring to gauge behavior against expected performance?
And, finally, what do we do in the long term? If you’ve devised a regime of user feedback or software-based monitoring during development testing, have you continued to use these tools – or something similar – in the medium and long term? Proactive evaluation has clear benefits, particularly if the systems are at a point where further scaling up would need new hardware or a step-up in cost.
[5]
Do please let us know your approach, warts and all, by taking part in our short survey below. There are three questions to answer. We'll run the poll for a few days and then publish a summary on The Register thereafter.
Don’t feel bad if you tick all the “we don’t do that” boxes, because there could be many reasons (not least time and cost) for not having a humongous capacity monitoring and planning regime. And if you tick all the “we do that in spades” boxes, try not to be too smug... ®
JavaScript Disabled Please Enable JavaScript to use this feature.
Get our [6]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YZU1QDV2JxdjgzD24treWQAAAFY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZU1QDV2JxdjgzD24treWQAAAFY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZU1QDV2JxdjgzD24treWQAAAFY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZU1QDV2JxdjgzD24treWQAAAFY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZU1QDV2JxdjgzD24treWQAAAFY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://whitepapers.theregister.com/
Our planning ahead - until now.
The computers I am asked to quote for are used as general workstations and are expected to have at least a 5yr working life. (Applications range from basic coding to Autodesk CAD suites, and all stops in between)
As such, we go for systems that are not-quite bleeding-edge to avoid the massive premiums that attracts, but they will 'fly' on pretty much anything thrown at them.
Last purchasing cycle was 2016. Minimum of 64GB RAM, latest mobo support chipset, and minimum of 12 cores (6+6). A 'fast' enterprise 1GB spinning rust was the standard boot, with slower secondary storage. Graphics was whatever the fastest non-mental-priced card at the time.
The only upgrades we've done since 2016 is change out the boot drives for SSDs and give our users new video cards where required - intensive CAD users.
Win11 has put a bit of a downer on this long-term strategy, but as hardware scarcity and stupid prices have skewed the market, our poor users will just have to soldier on for now.
choke points
I used to do a lot of capacity planning. The actual hardest part was persuading t' management that:
a) I knew what I was talking about
b) that they would have to spend money
However, the traditional process as described in ITIL and ISO9000 has largely become obsolete. What seems to happen more is that systems become vulnerable to unpredicted and often unknown behaviours due to crappy software design and architectures. Oh .... and networks.
Things that cannot be solved simply by running down to PC World for all the SIMMs they have in stock, or invoking the magic of capacity on demand.