GitLab versus The Zombie Repos: An old plot needs a new twist
- Reference: 1659952868
- News link: https://www.theregister.co.uk/2022/08/08/gitlab_versus_the_zombie_repos/
- Source link:
$1 million is certainly a lot to be wasting on a fossil collection, and is a full quarter of the company's total hosting costs. Who wouldn't want to spend it on something more fun? One answer is to cut 'em loose, which was what GitLab was [1]expected to do from September . In an attempt to forestall the inevitable tsunami of techiness, the GitLabbers set very generous rules – a project has to be untouched for a year, there'll be plenty of warning, and the merest brush of a code fairy's gossamer wings will reset the clock.
But that was never going to quell the outrage. Some of this is entitlement bias, but a lot of it is because of the harm done to open source when stuff just vanishes from places where it was once assured a safe harbor. Last week, just hours after The Reg [2]exclusively broke the story , the org made a [3]quick U-turn .
The problem of freemium
On the face of it, GitLab has fallen into the Freemium Brand Trap. The logic behind free tiers on a paid-for service is sound enough. If you have a system where the incremental cost of adding a new user is low enough, you can give restricted user access away for free. People can get their feet wet and build useful things; once they see your service as beneficial for larger projects, they can hand over the cash for the full-fat experience. For a service that's aimed at open source communities, there's the added bonus of Good Egg status – your brand becomes beloved. It's win-win.
The trap comes when the cumulative load of free tier users starts to cost more than you planned. If you're a crass commercial outfit, then you can put the screws on and trust that the paying base will keep your profile high and your brand aloft. With open source, you're seen as damaging the community. Because you are.
[4]
This is the first problem with the assumption that untouched code is dead code. Open source is built on stable components. Stable means not being fiddled with. But when you do need to revisit long-established code, say because a big yet ancient security flaw has just surfaced, you really need to revisit it. Open source should mean that this is always possible, no matter who used to own the code and what its life story has been since it was looked at.
[5]
[6]
Those are immediate concerns about vanishing code. Open source also means having code available for research, for education, for whatever unforeseen reasons. Nobody forced GitLab to offer a free service to open source, but it did – and that means a responsibility. It's part of the world's communal memory now.
Noble ideals aren't any good if you can't afford them, though, and a megabuck off the bottom line won't stop bleeding away by itself. Let's look at that figure. Is it legit? You can store a petabyte for up to five years in Fujifilm's Object Archive service for around $45,000, or less than $10k a year, one of the best deals around. GitLab has around 29 million non-active users at 5GB free repo space. That's 145 petabytes, or $1.3 million a year. The tyranny of the Freemium Brand Trap is the tyranny of numbers.
[7]GitLab U-turns on deleting dormant projects after backlash
[8]GitLab plans to delete dormant projects in free accounts
[9]Gtk 5 might drop X11 support, says GNOME dev
[10]Open source body quits GitHub, urges you to do the same
How much does the average zombie user actually use? It won't be 5GB. Only GitLab knows for sure, but there are clues. When the Windows source was moved to a Git repo [11]for the first time it was 3.5 million files supported by 4,000 engineers, and that came to a 300GB repo. That's under 100MB per user. Repos grow rapidly as old data is kept, up to a point, but that's amenable to cleaning. At Windows levels of repo use, that's $26,000 annual zombie tax.
Nobody can deny that GitLab has a zombie problem. It's not clear how big it really is, nor whether GitLab's proposed solution is optimal, especially given the nature of open source. But if the tyranny of numbers can work against you, it can work for you. What if the free tier was contingent on offering 10GB of your local storage to the community, with the resultant aggregated free tier storage managed by GitLab as the hosting system?
[12]
Would that even work? That turns out to be a set of really [13]interesting questions which could turn a lot of freemium models on their head.
It's not GitLab's job to develop startlingly innovative solutions to basic problems. But it is GitLab's job to be a good open source citizen, and that means not brooding in secret about problems and concocting The Answer. Go to the community. Lay out the facts. Say what solutions look plausible to you, then ask for input. Better believe you'll be having that conversation whether you like it or not, so own it from day one.
It goes against the grain for commercial companies to present problems rather than solutions, but the whole idea of open source is community problem solving and a rather better willingness to accept the truth than is the commercial norm. Sometimes, this means accepting that there's no such thing as a free lunch. Sometimes, this means compromise. Sometimes, this means a brilliant new idea.
[14]
We won't know, and GitLab won't know, if we don't try and find out. ®
Get our [15]Tech Resources
[1] https://www.theregister.com/2022/08/04/gitlab_data_retention_policy/
[2] https://www.theregister.com/2022/08/04/gitlab_data_retention_policy/
[3] https://www.theregister.com/2022/08/05/gitlab_reverses_deletion_policy/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YvEzI839oenbQ5QKIjsmgwAAAA0&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YvEzI839oenbQ5QKIjsmgwAAAA0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YvEzI839oenbQ5QKIjsmgwAAAA0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2022/08/05/gitlab_reverses_deletion_policy/
[8] https://www.theregister.com/2022/08/04/gitlab_data_retention_policy/
[9] https://www.theregister.com/2022/07/05/gtk_5_might_drop_x11/
[10] https://www.theregister.com/2022/06/30/software_freedom_conservancy_quits_github/
[11] https://devblogs.microsoft.com/bharry/the-largest-git-repo-on-the-planet/
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YvEzI839oenbQ5QKIjsmgwAAAA0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[13] https://www.cncf.io/blog/2019/11/04/building-a-large-scale-distributed-storage-system-based-on-raft/
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YvEzI839oenbQ5QKIjsmgwAAAA0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://whitepapers.theregister.com/
I have 25 or so repositories, totalling about 1GB. Much of that hasn't been active recently.
> the merest brush of a code fairy's gossamer wings will reset the clock.
I'm not sure GitLab thought this through. It's not hard to write a bit of code that scans all local repositories, updates a small managed file and push the change. I've done it, purely for research purposes of course! So they would get involved in an "arms race" with devious members trying to keep their repositories live and GitLab constantly changing what they consider "meaningful changes" to a repository.
I hope these repositories are set to not let some random El Reg reader update them on a whim.
Even raising an issue was supposed to be enough to keep it alive. And even then, it would have culled the very very dead wood that truly nobody cares about, not even enough to set up a cron job, let alone move to (gasp) a 5$/mo tier.
When is a project a Zombie?
Would there be merit in saying something that hasn't been downloaded or cloned is a better indicator, rather than looking for updates?
Unfortunate timing
The timing of this has been disastrous for GitLab. This happens just when groups such as the Software Freedom Conservancy were making inroads with their [1]Give up GitHub campaign on the back of GitHub giving away other people’s open source code as part of its CoPilot feature.
That GiveUpGitHub move seems to have caught on, judging from noises around the #GiveUpGitHub hashtag [2]on Twitter and [3]Mastodon and alternative services such as [4]Hostea and [5]Codeberg are reporting a lot of interest.
Of course the most pure open source advocates would still have been suspicious of GitLab but if they had played their cards right GitLab could have been a refuge for many people leaving GitHub but slightly hesitant of moving to smaller forges such as Hostea or Codeberg.
That doesn’t solve GitLab’s Freemium economics problem, of course, though perhaps an influx of paying projects on the back of the GitHub exodus could have changed it for the better.
[1] https://sfconservancy.org/GiveUpGitHub/
[2] https://twitter.com/hashtag/GiveUpGitHub
[3] https://mastodon.social/tags/GiveUpGitHub
[4] https://hostea.org/
[5] https://codeberg.org/
Off the cloud
Maybe they should move off the cloud? All these "zombie" project could probably be hosted on one beefy server for like $500 a month.
I am not a programmer, (so not a user of Gitlab, Github, or any other repo service), but one way, that I would have considered reasonable, is that if you are no longer contactable after a year of inactivity then your account gets shut down. So after a year of no activity, Gitlab sends you an email to your associated email address. If you dont respond, then your account goes away. Hell for safety make it multiple emails over the space of 6 months, but if your not contactable by then, I dont really see a problem.
Perhaps for safety, after that 6 months, Gitlab removes the Repo from being viewed or used or whatever. Keep the data but with noone able to contact, it view it, use it , clone it, whatever. And keep it in that limbo for 3 months. If there are no massive howls of anger from the community, then it goes to big the Data graveyard in the sky. If someone suddenly finds themselves needing that repo, then it can be restored, and Gitlab can probably charge for its restoration at that point.
I cant really see a problem with this idea, beyond that Gitlab will have to put up with zombie projects for another 18 months...
This could be disastrous where a small project is completely stable and doesn't need any changes, but in regular use by large numbers of other code. The original author might not be available for a number of reasons.
Define "disastrous".
If "large numbers of other code" was making regular use, one would certainly hope that those responsible for that other code had measures in place to guard against exactly this potential scenario. Or the scenario that the original author takes down their own repository. Or various other scenarios.
Because if they don't take those measures then all bets are off and frankly GitLab is only one of a number of problems you now have,
Turtles All the Way Down Scenario?
Not having had to personally deal with this, I'll ask: how easy/difficult is it for a programmer to know ALL his/her project's dependencies? If a project depends on, say, X11 can the project-depending-on-X11's programmer reasonably discover all X11's dependencies and sub-dependencies ?
Point being, there may be an inactive project which has many things at higher levels depending on it, but due to the depth of the dependency stack, people running the higher-level projects might not realize how important that inactive project is when GitLab, or whomever, asks, "Is anybody using this? No? Okay, we'll pull the plug."
I did suggest that when it gets taken down after 6 months of no replies, if there is no outcry from the community, then it gets deleted. Are you suggesting that people wouldnt notice for 3 months it being down, and would not complain about it disappearing within that 3 months?
If it's vital, and suddenly disappears people would absolutely complain. Then Gitlab could bring it back, but perhaps by assigning it to someone else, as the original Author has not repsond to 6 months of messages, and so cannot be considered to be enaged with the community.
OK. For those of you who have downvoted, propose your own solutions...
The problem is
The project owner may have disappeared / died, but the code is still being actively used - a bit like [1]this well know XKCD
[1] https://xkcd.com/2347/
I don't understand the suggestion...
"What if the free tier was contingent on offering 10GB of your local storage to the community, with the resultant aggregated free tier storage managed by GitLab as the hosting system" I don't follow!
Re: I don't understand the suggestion...
I think the person who wrote about "offering 10 GB of storage" meant, "dedicating 10GB of storage on their home PC, and running a distributed filesystem package on that home PC to make it available to GitLab."
Re: I don't understand the suggestion...
I guess it's alluding to a sort of BitTorrent-ish thing, where you would use some Web3.0 distributed decentralized filesystem. Even the up/down ratio thing finds itself again as storage donated vs used.
29 million non-active users, $1.3 million a year. (Should that be "accounts?")
That works out to 4 cents a year per user.
If they charged 1 dollar a year for a 5GB account, they could make a profit,
and there would be less chance of Gitlab going belly up.
If Gitlab goes belly up, all the code gets lost.
There is no way to escape a free tier with 5GB being targeted for use as backup storage.
If GitLab goes belly up, only the hosted-repo is lost.
There will still be local clones of the repo with the associated users.
> Some of this is entitlement bias
No Rupert it isn't. And since you're going to be that kind of dick I'm not reading the rest of your article.
TL;DR
seems the real problem is using a development platform as an archive - two different jobs. A little like trying to run a publishers inside a library.
And that is not a new, or even a tech problem. It's a little worrying no one at Gitlab spotted that.
I wonder how quickly Google etc would pony up if a vanished archive broke Chrome(OS) ?
Read-only content
How easy would it be for Gitlab to spend a few $10k for one of their engineers to cook up a read-only version of the repo interface that lets it be a dependency of other things, but does not implement all of the expensive parts. In other words, be as near as a few files behind and nginx service as possible?
Re: Read-only content
Not very.
But I missed the answer to the preceding question of "Why ?"
One thing I've noticed on GitHub - I'm guessing it extends to GitLab too - is that people tend to use forks almost like bookmarks, so you'll often find a project has dozens of forks which have never actually diverged from their parent commit.
If they're not already doing something smart storage-wise to avoid duplication then removing inert forks (and ideally replacing them with some kind of redirect to the upstream project) could save quite a bit.
The article goes on about the responsibilities of GitLab, but I must have missed the responsibilities of the code creators to not just shoot their code into the universe and expect somebody else to host it perpetually for free.
Remember, please, that GitLab the software is Open Source!
There's one critical thing that's missing from this article: GitLab, the software, is an open-source software!
However much GitLab might try to lean on the fact that GitLab dot com offers some Enterprise Edition features -- not fully open-source -- to free users, the GitLab product stems from an open-source background and the core functionality certainly is still open-source. Many of the supposed freeloaders contributed patches and debugging time and feedback and well researched issue reports and other input into that product!
It is quite dishonest for GitLab dot com, the commercial entity, to simply sum up the cost of keeping some hard-drives spinning! They also should perform the impossible calculation of how much of their income from actual paying customers should rightly be attributed to work from the community they're now spurning.
I don't think anyone on the open-source side of this equation was or is complaining that GitLab dot com brings in income from exploiting the open-source portion of their code base -- it's within the terms of the license. But, to appreciate exactly *why* this feels like a massive rug-pull to many of us, ask this: would anybody have ever contributed to GitLab open-source, had they know they were just free labour for a corporation that chooses to optimise its bottom-line at the expense of this very community -- pretty much just like any other capitalist corporation?
Prolly not, yeah? Capitalism and community don't mix!
I mean, I'm bitter because I've just had to spend a tonne of my time migrating from self-hosted GitLab to self-hosted Gitea. This, it turns out, was a very good decision but I rather liked GitLab, back in the day, and do somewhat resent the way that they've been treating GitLab CE users as second-class citizens for a while -- pretty much making from-source builds too onerous to bother with, forcing the use of Omnibus or official, bloated Docker images, and pushing U.I. junk that can't be disabled, readily, in CE, that nobody asked for, but does nothing but plug an EE-only feature.
The writing has rather been on the wall for at least some years!
I did wonder about the volume of data held on free projects. 5GB is a *lot* of code, so unless folks are stuffing it with backup ZIP files, etc, it is hard to see that being used up.
There was talk of moving stuff to slower storage, you would think that was already automatic (i.e. only files that are frequently/recently requested stay on even the HDD-tier of the back-end storage) and certainty if they are suffering from the $/GB for SSD use than why not have tiring where paid users get the fast/expensive stuff, and the free user's projects get punted down a layer on to less expensive and slower storage?
Edit: Just checked, I have 4 projects on GitLab, one public using 2.4MB and 3 private, all totalling 4MB and last updated 2 years ago. How typical am I?