Boffins rate npm and PyPI package security and it's not good
- Reference: 1660179265
- News link: https://www.theregister.co.uk/2022/08/11/npm_pypi_security/
- Source link:
Computer scientists at North Carolina State University have put one of its tools to the test by evaluating software package registries npm and PyPI using [1]OpenSSF Scorecards .
As security expert Bruce Schneier [2]observed two decades ago, "Security is a process, not a product."
[3]
Yet the processes used appear not to have worked very well - certainly not well enough to prevent the supply chain attack on SolarWinds' Orion software, the ransomware attack on the Colonial Pipeline network, or the exploitation of the Apache Log4j bug. Nor have the blizzard of responses - executive orders, government initiatives, academic engagement, and private sector efforts - made malicious hacking much more of a challenge.
Packages could achieve a score of 10 even if it had multiple unpinned dependencies
The open source ecosystem, upon which so much public and private sector code depends, has been a remedial focal point. The Linux Foundation's August 3 2020 [4]establishment of the OpenSSF , backed by IBM, GitHub, Google, JPMorgan Chase, Microsoft, NCC Group, and Red Hat, among others, represents one recent attempt at amelioration.
The OpenSSF Scorecard project was launched in November 2020, to provide an automated tool to determine whether particular security practices are followed. Scorecards rate 18 different heuristics, or checks, with a score ranging from 0 to 10. These include things like [5]Binary Artifacts , [6]Branch Protection , and [7]Dangerous Workflow , to name a few.
[8]
[9]
In [10]a preprint paper distributed via ArXiv, NCSU researchers Nusrat Zahan, Parth Kanakiya, Brian Hambleton, Shohanuzzaman Shohan, and Laurie Williams applied the OpenSSF Scorecard to software packages within npm and PyPI in order to see what security practices could be identified among the developers using those registries.
"Our study shows a gap in security practices for both ecosystems," said Nusrat Zahan, corresponding author of the study and a doctoral student at North Carolina State University, in an email to The Register . "Code-Review, Maintained, Binary Artifacts, License, and Branch Protection are practices that assess a repository's security posture. Aside from Binary Artifacts, you will notice that both ecosystems failed to implement these practices at scale."
[11]
"On the contrary, practices like Dangerous Workflows and Token Permission scan GitHub workflows to verify the presence of good practices. But what would happen if a repository did not contain GitHub workflows? The tool would still give a high score to that package because it could not detect any bad practices. Hence, even if these metrics had a high percentage of packages with good practices, it also opens up the debate about whether Scorecard should check for the existence of GitHub workflows before verifying the good or bad practices for accurate results."
[12]Email domain for NPM lib with 6m downloads a week grabbed by expert to make a point
[13]Homeland Security warns: Expect Log4j risks for 'a decade or longer'
[14]Open source isn't the security problem – misusing it is
[15]How do you fix a problem like open-source security? Google has an idea, though constraints may not go down well
The researchers' results – which they expect to update later this week in a revised draft – show both the value and limitations of automated security testing.
Both npm and PyPI scored well in the " [16]Dangerous Workflow " check, the only metric rated "Critical" in terms of importance. This check looks for untrusted code checkout and for script injection with untrusted context variables in packages' GitHub workflows as a result of misconfigured GitHub Actions (automation scripts).
"More than 99 percent of packages passed the check," the researchers' paper says. "However, we found 1,938 npm packages and 508 PyPI packages where Scorecard found vulnerable code patterns."
An attacker could abuse a vulnerable package, for example, by crafting a malicious [17]GitHub issue title that injects code and opens a reverse shell connection. The fact that 99 percent of packages dealt with this risk is heartening but as security types [18]often observe , "Defenders have to be right 100 percent of the time and attackers have to be right once…"
[19]
"It shows that we can use the Scorecard tool to detect open vulnerabilities in a potential dependency," explained Zahan. "The statistic might not accurately reflect the number of packages with open vulnerabilities. OSV [Open Source Vulnerabilities database] only contains a list of vulnerabilities that have been reported."
"Packages may contain more vulnerabilities than are listed. For example, [20]Elder et al. showed in a study that they found 95 times more vulnerabilities than reported. Hence, if we do more in-depth studies to detect vulnerabilities, we might find more than we know, and in that case, our finding shows evidence that we need to focus on secure coding. Note that the scorecard tool gives us a way to measure these security practices, but it is up to the practitioners to determine how they can improve package security."
The " [21]Maintained " check underscores just how much open source software is not attentively maintained. The researchers found "more than 85 percent of packages in npm and 75 percent of PyPI packages were unmaintained in GitHub."
The " [22]Code Review " check also revealed a useful finding: Only 30 percent of npm packages and 34 percent of PyPI packages declared code review practices in their repositories.
The researchers say that's to be expected given that, particularly in npm, packages often have only a single maintainer. They point to [23]a study published last year , conducted by some of the same computer scientists, that found 1.5 million npm packages had an average of 1.7 maintainers.
But given the solo nature of so many of these software libraries, those involved in securing the open source ecosystem may want to explore whether cost-effective code reviews can be made available for popular single-person projects.
Both npm and PyPI scored poorly on checks like "Security-Policy," "Packaging," "Signed Releases," and "Fuzzing." While none of these gaps represent urgent problems, they show how these package ecosystems and participating developers could take security more seriously.
Another heuristic, " [24]Pinned Dependencies ," seems to show npm and PyPI in a good light, with more than 99 percent of packages having at least one pinned dependency. Of these, 81 percent of npm packages and 66 percent of PyPI packages scored 10 – they had no unpinned dependencies, which is generally considered safer.
But those high scores masked frailties. "We found packages could achieve a score of 10 even if the package had multiple unpinned dependencies in the JavaScript package’s package.json file, indicating Scorecard findings do not indicate the accurate status of pinned dependencies in an ecosystem," the paper explains. "We also observed that Scorecard does not verify the presence of Dockerfiles, shell scripts, and GitHub workflows files in a repository."
What this suggests is that automation alone isn't enough. Automated tools have to be able to make accurate measurements and those tools remain works-in-progress.
"Scorecard provides a head start for practitioners to measure package security practices," said Zahan. "The Scorecard project is evolving based on the findings and recommendations from practitioners. The software industry seeks to standardize supply chain security procedures through initiatives including Scorecard, [25]Alpha-Omega , [26]OSV , and [27]OSI ."
"Research like ours helps to understand how a package performs against other OSS packages in an ecosystem and how the scorecard can improve automated testing. The Scorecard team welcomed our research and agreed to work on these findings to enable automated testing to run more effectively. But it will require community efforts to standardize and implement these tests." ®
Get our [28]Tech Resources
[1] https://github.com/ossf/scorecard
[2] https://www.schneier.com/essays/archives/2000/04/the_process_of_secur.html
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YvR@5n1didhn56Vudx02owAAAME&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://www.theregister.com/2020/08/03/linux_foundation_forms_openssf/
[5] https://github.com/ossf/scorecard/blob/main/docs/checks.md#binary-artifacts
[6] https://github.com/ossf/scorecard/blob/main/docs/checks.md#branch-protection
[7] https://github.com/ossf/scorecard/blob/main/docs/checks.md#dangerous-workflow
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YvR@5n1didhn56Vudx02owAAAME&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YvR@5n1didhn56Vudx02owAAAME&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] https://arxiv.org/abs/2208.03412
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YvR@5n1didhn56Vudx02owAAAME&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[12] https://www.theregister.com/2022/05/10/security_npm_email/
[13] https://www.theregister.com/2022/07/14/dhs_warns_expect_log4j_risks/
[14] https://www.theregister.com/2022/01/12/open_source_isnt_the_problem/
[15] https://www.theregister.com/2021/02/04/google_open_source_security/
[16] https://github.com/ossf/scorecard/blob/main/docs/checks.md#dangerous-workflow
[17] https://docs.github.com/en/issues/tracking-your-work-with-issues/about-issues
[18] https://blog.apnic.net/2019/07/04/forcing-attackers-to-be-100-right/
[19] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YvR@5n1didhn56Vudx02owAAAME&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[20] https://trebuchet.public.springernature.app/get_content/fabb4f89-51cc-4e82-95f5-b8f6bd3716ae
[21] https://github.com/ossf/scorecard/blob/main/docs/checks.md#maintained
[22] https://github.com/ossf/scorecard/blob/main/docs/checks.md#code-review
[23] https://arxiv.org/abs/2112.10165
[24] https://github.com/ossf/scorecard/blob/main/docs/checks.md#pinned-dependencies
[25] https://github.com/ossf/alpha-omega
[26] https://osv.dev/
[27] https://deps.dev/
[28] https://whitepapers.theregister.com/
Not Impressed.
I can't comment on NPM, because I haven't used it, but I do have projects on PyPI. I had a look at the paper, and it's pretty clear that the "problems" listed are not in PyPI, but rather in Github.
To start with, they don't actually look at PyPI except to get a list of projects which they then look for on Github. There is no link between PyPI and Github. You can have packages in PyPI without having a Github account or any code in Github. They are two completely independent things.
Their scorecard is entirely based on the assumption that you do everything through Github and use all of it's workflow features. If you use Github just as a place to publish code for the public, then you will get a low score. If you use all the Github bells and whistles and use them the right way, then you get a high score.
In other words, part of the score is based on result, and part of it is based on "process". And by "process" they only mean is your process conducted in Github rather than somewhere else.
A good example is "maintained". If a project doesn't get at least one commit per week to Github, then it is is marked down. There's no reason why that should be a valid criteria. The project may not be unmaintained. It may simply be stable and isn't getting updates because there isn't anything wrong which needs fixing. Or you could be working away on new features, but Github is just where you publish the source code as opposed to the place where you actually work from.
This is why there are so many projects which score highly in terms of not having anything wrong, but most seem to have low scores in terms of making use of Github's automated work processes.
I have a Github account and I have packages in PyPI. Part of my work process is to push code to Github for source publishing and to upload packages to PyPI for users. I have my own testing and QA processes which I run on my own hardware as I have no intention of locking myself into Github. It's just a convenient place to host the source code for anyone who wants it. I have been planning to also push source to another Git repo aside from Github to reduce my dependency on them for some time, but I simply haven't got around to it yet.
Overall, I'm not impressed with the report.
P.S. "Standard" security mode (the most relaxed standard setting) in Firefox seems to give The Register fits and result in a page not found error. I can only post on this site by fiddling with the security settings and manually turning off tracking protection. I've no problems anywhere else. El Reg should get a "fail" on the testing and maintaining score card.