revisit: a Workflow Tool for Data Science
Abstract
In recent years there has been widespread concern in the scientific community over a reproducibility crisis. Among the major causes that have been identified is statistical: In many scientific research the statistical analysis (including data preparation) suffers from a lack of transparency and methodological problems, major obstructions to reproducibility. The revisit package aims toward remedying this problem, by generating a "software paper trail" of the statistical operations applied to a dataset. This record can be "replayed" for verification purposes, as well as be modified to enable alternative analyses. The software also issues warnings of certain kinds of potential errors in statistical methodology, again related to the reproducibility issue.
- Publication:
-
arXiv e-prints
- Pub Date:
- August 2017
- DOI:
- 10.48550/arXiv.1708.04789
- arXiv:
- arXiv:1708.04789
- Bibcode:
- 2017arXiv170804789M
- Keywords:
-
- Statistics - Applications;
- Computer Science - Computers and Society