Academics tell Brit MPs to check the software used when considering reproducibility in science and tech research
- Reference: 1637843588
- News link: https://www.theregister.co.uk/2021/11/25/research_software_inquiry/
- Source link:
According to joint academics group the Software Sustainability Institute, about 69 per cent of research is produced with specialist software, which could be anything from short scripts to solve a specific problem, to complex spreadsheets analysing collected data, to the millions of lines of code behind the Large Hadron Collider and the Square Kilometre Array.
"With many studies, research published without the underlying software used to produce the results is unverifiable," the institute said in its submission to the Parliamentary Science and Technology Committee's Reproducibility and Research Integrity Inquiry.
[2]
The institute said a 2014 workshop had revealed that research software "is infrequently subjected to the same scientific rigour as is applied to more conventional experimental apparatus."
[3]
[4]
The committee has just [5]published 86 similar pieces of evidence addressing this broad and difficult problem.
The reproducibility crisis came to the fore in 2005 when Stanford School of Medicine professor John Ioannidis [6]published a paper titled: "Why most published research findings are false."
[7]
Since then, the issue has arisen in a number of studies demonstrating the prevalence of irreproducible data, the committee said. UK Research and Innovation (UKRI), the non-departmental public body funded through the Department for Business, Energy and Industrial Strategy, is in the process of establishing a national research integrity committee.
However, the specific issue of reproducible research has thus far been overlooked, the Science and Technology Committee said.
Others question whether "crisis" is the right word. King's College London said in its submission: "We believe all stakeholders have a responsibility to use sober and accurate language to discuss these challenges and question whether 'crisis' risks overgeneralising a complex set of issues and may allow for misrepresentation in the media which has the potential to needlessly damage public trust.
[8]
"Referring to irreproducible research findings as a 'crisis' may also imply that this is an acute issue that can be swiftly resolved, rather than a characteristic embedded in our current research culture."
Crisis or not, the problem has many different strands, from medicine to neuropsychology, and even to research into battery performance, according to the submissions.
[9]Theranos' Holmes admits she slapped Big Pharma logos on lab reports to boost her biz
[10]NASA boffins seem to think we're worth saving from fiery asteroid death so they're shooting a spaceship at one
[11]Genetically modified E coli bacteria produce ink for 3D printing programmable objects
[12]Swiss lab's rooftop demo shows sunlight and air can make fuel
None of this should prevent the underlying role of software from being ignored, the Software Sustainability Institute argued.
"To enable systemic change to improve reproducibility and research integrity, the quality and transparency of software must be improved," it said in its submission.
"Research software engineers have a key role to play by making the software used in research more robust and reusable, and helping train researchers in the fundamentals of publishing code so others can review and inspect it."
According to its website, the Software Sustainability Institute has facilitated the advancement of software in research by cultivating better, more sustainable research software to enable world-class research since 2010. It was previously awarded funding from all seven research councils and its mission is to become the world-leading hub for research software practice.
The Science and Technology Committee will hold oral sessions to take evidence from December and a report will follow. Then we'll see if software is given due consideration. ®
Get our [13]Tech Resources
[1] https://www.oxfordeconomics.com/recent-releases/the-impact-of-the-innovation-research-and-technology-sector-on-the-uk-economy
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YZ-Bc4d0MQYZPtP8asL9kQAAAEk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZ-Bc4d0MQYZPtP8asL9kQAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZ-Bc4d0MQYZPtP8asL9kQAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://committees.parliament.uk/work/1433/reproducibility-and-research-integrity/publications/written-evidence/
[6] https://en.wikipedia.org/wiki/Why_Most_Published_Research_Findings_Are_False
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZ-Bc4d0MQYZPtP8asL9kQAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZ-Bc4d0MQYZPtP8asL9kQAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2021/11/24/theranos_fraud_trial/
[10] https://www.theregister.com/2021/11/24/double_asteroid_redirection_test/
[11] https://www.theregister.com/2021/11/23/genetically_modified_e_coli_bacteria/
[12] https://www.theregister.com/2021/11/10/swiss_labs_rooftop_demo/
[13] https://whitepapers.theregister.com/
Re: Career development disguised as science
The other issue is that many of those writing the software are new to programming.
This post generated a lot of discussion of the research software mailing list:
https://shape-of-code.com/2021/02/21/research-software-code-is-likely-to-remain-a-tangled-mess/
"an acute issue that can be swiftly resolved"
So a crisis is something that can be quickly resolved.
Then climate change is endemic to our civilization.
Yup, sounds about right.
Re: "an acute issue that can be swiftly resolved"
Strictly, a crisis is what we now call a 'tipping point' due to the misapplication of 'crisis' to mean anything serious that's occurring.
'Crisis' is derived from the same root as 'critic' and 'critical', all relating to the ancient Greek for 'to decide'.
The good old OED [vol 2, 1933ed.] defines it as" A vitally important decisive stage in the development of anything; a turning point " . Hence, when a nuclear reactor goes critical, it's the point where the reaction becomes self-sustaining, but a crisis can by definition not persist , so an 'ongoing crisis' is a nonsense.
What's the betting that...
...in most cases the software used will be Excel (probably an old unsupported version) that's running lots of macros that were programmed by someone who no longer works for the institution, and no one there has any clue how it works, so they will leave it to do its "magic" until the day it dies*.
* which will be 2 weeks before it needs to be used for something super-critical whose deadline can't be moved.
Re: What's the betting that...
No idea why you're getting downvotes. I do know for certain that the government usually insists that computational deliverables from their research contracts are done as Excel spreadsheets rather than languages like Python, simply because Excel is regarded as universally available.
Re: What's the betting that...
>I do know for certain that the government usually insists that computational deliverables from their research contracts are done as Excel spreadsheets
Your certainty is very much misplaced.
Re: What's the betting that...
...in most cases the software used will be Excel (probably an old unsupported version) that's running lots of macros that were programmed by someone who no longer works for the institution, and no one there has any clue how it works, so they will leave it to do its "magic" until the day it dies*.
* which will be 2 weeks before it needs to be used for something super-critical whose deadline can't be moved.
Hmm, looking at the latest version of our software:
$ ls -R $BASE_DIR > foo
$ find $BASE_DIR -type f | wc
87781 87781 11047154
$ grep xls foo
No data
$ grep '[.]cc$' foo | wc
15382 15382 356232
$ grep '[.]h$' foo|wc
13890 13890 284324
grep '[.]cpp$' foo | wc
785 785 16326
$ grep '[.]hpp$' foo | wc
5 5 104
$ grep '[.]py$' foo | wc
17493 17493 471904
$ grep '[.]sh$' foo | wc
831 831 15764
$ grep '[.]xml$' foo | wc
5599 5599 95946
I think I can confidently say there's not much Excel in our analyses. ROOT, of course, is another matter.
Re: What's the betting that...
I think you need to know about the -l flag to wc
Re: What's the betting that...
> I think you need to know about the -l flag to wc
It's three extra characters to type...
Re: What's the betting that...
Unfortunately Excel does tend to be heavily used. It shouldn't be. It's totally unsuited to most tasks that end up in publications. I've spent many hours correcting errors introduced by the kinds of slips and trips that are far too easy to make in Excel.
The whole point about reproducibility is that it eliminates the possibility that the software (or the instruments, or the power supply, or the orientation of the lab bench or ...) is responsible for the results. If they can't be reproduced with a different set up, great. If they can't it might be of interest to investigate why, at which point all these details can be considered.
None of this applies to psychology, which is almost all irreproducible because it's a load of bollocks. Technical scientific term there.
And let us not forget that in many papers published no detailed raw data in the paper / linked online, maybe you will get a few graphs & brief description of methodology and results (the cynic in me notes that in many cases you look for details on statistical significance of the results and its not present) ... if you want the actual raw data its normally a matter of contacting the paper authors & hoping they can oblige.
Disclosure - Before IT I worked in biosciences research & still "follow" various research areas that interest
me, subjectively it does seem that there's more pressure on researchers to churn out papers quickly and often & a fair amount of paper number inflation now (i.e. work "fragments" dribbled out over several papers that not so long ago would have just been published in a single paper, turning 1 chunky paper into 4 smaller ones dealing with different facets of the same work).
Not a new phenomenon
" work "fragments" dribbled out over several papers that not so long ago would have just been published in a single paper, "
Even in the late '80s to early '90s when I was working in science (both physics and ecology) this was recognised and referred to as ' ravioli publishing ' (lots of little bite size pieces).
The lack of raw data is understandable because journals simply could not print it. That in turn reinforces what most of us think; that the whole idea of printed journals is ludicrously outdated. In many areas of science they are completely irrelevant, because by the time a paper is printed everybody who cares has read it from a preprint server.
The only purpose journals serve is one for which they were never intended; as proxy measures of the worth of a particular piece of research. The whole concept of a "high-impact journal" is a piece of self-perpetuating nonsense.
The whole system of dissemination needs urgent and radical overhaul. Open access journals aren't the answer, because they "publish" any old crap. On the other hand a revamped version of the reviewer system isn't the answer either, because it's inherently unreliable, rewards personal connection over actual worth and demonstrably discriminates against non-white and women researchers.
A preprint server with a star rating system? Weighted by the star weight of the reviewers, maybe?
And yes, there are far, far, far too many papers, most of which are wrong and almost all of which are irrelevant. It's an artefact of the stupidly gladiatorial funding system. I have a friend who is a very senior tenured humanities professor in the US. He writes a book every five years.
and then ignore email requests for the data
And when you email asking for the data, the researchers don't reply, or at least 30% of them don't:
https://shape-of-code.com/2017/02/02/i-have-been-reading-your-interesting-paper/
However, the situation with regard to data being made available on sites such as Zenodo is getting better.
The whole point about reproducibility is that it eliminates the possibility that the software is responsible for the results.
I don't know whether it's still the case, but back in the 70s/80s there was a semi-joke amongst particle physicists that half the particles being found by the various accelerators round the world(*) were actually bugs in the Fortran code that was used for analysis everywhere.
(*) Those were the days when accelerators were small and cheap enough that rich countries could afford their own rather than having to share one world wide one.
Although one can get discrepancies when data are analysed by different versions of the same software even.
But will the Central IT allow people to freeze the software versions on their machines until the study is completed? Not a chance.
They will in research-heavy organisations, yes. It's a key requirement when working with the likes of pharma or life sciences to be able to set in stone and verify a combination of software and data, exactly to enable this kind of working pattern.
Someone should let them know then. I've* been banging my head on a brick wall about this for over 5 years.
*I work in life-sciences in a research-heavy organisation supported by a central IT service that's totally focussed on "The Student Experience". And when they say "student" they mean "undergraduate student".
Documented software and a scripted demonstration that it builds/runs on a clean OS system should be published with any peer-reviewed paper?
Would send shock waves through the industry but might just force some sound practice and the ability to analyse the process.
The above would rule out GNU radio being used...
How far do you go with that? Should any paper with a plot produced by Gnuplot have to include the full Gnuplot source (at the version used) in case apparently interesting features are actually artefacts? Can I analyse data using Matlab or do the hard maths with Maple, both of which are closed source?
+1 etc to this. This is a genuinely hard problem, especially when you get into some of the more large-scale stuff happening with ML models. Sure, I could give you my code, and I can even tell you exactly which commit to use in the underlying library and what toolchain to build it with.
But unless you've got a couple of hundred thousand dollars kicking around to download, store and process the training data and re-run the validations you've not got a snowball's hope of reproducing my work.
We need to be a lot more clever than "just publish the source" or "do more testing"
Papers should include the data and methodology used in the paper, or as supplementary information. This doesn't always happen. One of my favorite examples came from climate science where a novel technique called 'Rhamstorf smoothing' was used. And it eventually turned out to have been a standard triangle filter. Other times, data aren't included at all.
Not sure what the solution should be, ie including documented source for a climate model might be excessive, but it should be open to peer review. And I guess how much software engineerin we should expect.
For generally available software then yes, the version numbers should be included just in case.
Must it be open source? No, but it needs to be reproducible so your data run on my machine using XYZ's software gives the same result. And the method(s) used also possible with something that is open-source in broadest sense (i.e. can be inspected, need not be GPL or any other specific license).
>"is infrequently subjected to the same scientific rigour as is applied to more conventional experimental apparatus."
Part of me very much wants to argue that it probably shouldn't be subject to the same rigours, because fixing a broken/invalid piece of software is almost always a damned sight cheaper, faster and easier than replacing a possibly-very-expensive bit of kit, as it is generally cheaper/easier to validate the correctness of a piece of software. There is value to agility in software, and serious risk in saying we need to put the same levels of effort and cost into assuring the intangible as we do the tangible.
We should absolutely be talking about software assurance and how to properly review software assets, but starting from "well we do this with the physical kit" is probably the wrong place to start.
Ultimately it's a question of costs and funds. Academics are incentivised to publish their results. They are not necessarily incentivised to spend an extra few weeks tidying up their code data and getting them published on an equal footing.
We also need to be careful conflating the "reproducibility crisis" as published by Ioannidis with issues of software quality, version control and provenance as discussed in the rest of the article. Ioannidis identified serious, structural and fundamental issues in research practices producing studies across medicine, and none of them had anything to do with technical issues like kit or code. Higher quality or peer-reviewed research software wouldn't have made the issues identified by Ioannidis any better.
"scientists"
The scientific community does nothing to root out hacks from its ranks.
The litmus test I am using as a benchmark for rubbishness of the community is cannabis research coming from its bowels.
Inadequate control groups or lack thereof, massaging the data, fitting results to match the politically correct outcome and so on.
Not heard of any "scientist" being fired, department closed or defunded. Not a word of critique from the community.
Then that "research" gets widely published and is being used to propel the governments' propaganda.
It used to be that "scientists" could only receive grants if their research demonstrated bad effects of cannabis use.
There is hope though [1]Covid-19: Researcher blows the whistle on data integrity issues in Pfizer’s vaccine trial
[1] https://www.bmj.com/content/375/bmj.n2635
Re: "scientists"
Political considerations affect all sorts of research. I have a pal in the Geography Dept of one of the UK's best known universities who tells me that even to propose research looking at the likely extent of global warming (not denying it, you understand, just looking for some numbers) is seriously career limiting. Likewise it is politically impossible to research whether parental behaviour leads to ASD/ADHD and so on in children, or whether children who go to nurseries do better than those raised at home.
And so we end up with the nonsense of "neurodiversity" when in the overwhelming majority of cases there isn't a shred of evidence that there is - in computing terms - a hardware error rather than a software one.
Great principle
It's undeniably an important principle but it all comes down to implementation. I thought the CovidSim code check was a nice example of what it would look like in practice. I remember there was a spurt of complaints about the code quality (I thought the complaint that using a single letter variable names, like "r", was bad practice aged particularly well) and then an independent review found it worked fine.
I didn't really get the sense we got a lot of scientific value out of the exercise
Nature article on it: https://www.nature.com/articles/d41586-020-01685-y
Tell that the to the consultants that insist on using ANSYS, SolidWorks etc.
I have lost count of how many recent PhD's involve running 50-year old finite-element analysis methods on some modern problem; inside the sandbox of one of the commercial modelling tools.
Not saying it's not valuable work; quite the opposite, but call me weird they aren't quite the same as the PhD's of old where you had to actually write the model! (Perhaps just as well, given the effort involved in creating and validating such a thing).
Nowadays we use software tools to do research. Thirty years ago we wrote the software tools to do research. Thirty years before that we wrote the operating system as well. Jolly good. Life moves on.
Career development disguised as science
The rewards are given for publication counts and quality tokens (publication in top rated journals). Long-term work on thorough methods development is not well rewarded. And it should not be forgotten that the art of good bias design - difficult to detect and easy to defend - is highly profitable when judiciously applied.