Article ID Journal Published Year Pages File Type
424636 Future Generation Computer Systems 2013 13 Pages PDF
Abstract

•We created an automated, in situ approach for capturing provenance in spreadsheets.•We implemented an Excel Provenance Add-in with functions accessible via a ribbon menu.•Our tool generates understandable provenance logs and visualizations.•Our case studies suggest that our approach is efficient and beneficial to researchers.

One of the most important tasks in eScience is capturing the provenance of data. While scientists frequently use off-the-shelf analysis tools to process and manipulate data, current provenance techniques such as those based on scientific workflows are typically not able to trace internal data manipulations that occur within these tools. In this paper, we focus on one such off-the-shelf tool, MS Excel, which is used by many scientists; specifically, we propose InSituTrac, an automated in situ provenance approach for spreadsheet data in Excel. Our framework captures data provenance unobtrusively in the background, allows for user annotations, provides undo/redo functionality at various levels of granularity, presents the captured provenance in an accessible format, and visualizes captured provenance to support analysis of the provenance log. We highlight several motivating use case scenarios which show how provenance queries can be answered by our approach. Finally, case studies with an atmospheric science research group and a fisheries research group suggest that the automated provenance approach is both efficient and useful to scientists.

Related Topics
Physical Sciences and Engineering Computer Science Computational Theory and Mathematics
Authors
,