Skip to main content
University of Wisconsin Madison crest Carbon Model modeling carbon budgets in large terrestrial ecosystems

Policy Issues in Accessibility and Interoperability of Scientific Data: Experiences from the Carbon Modeling Field

Puneet Kishor presented at AGU 2010 in San Francisco on why data access and interoperability decide the pace of large scale ecosystem modeling.

Puneet Kishor presented at AGU 2010 in San Francisco, California.

Abstract. Large scale terrestrial ecosystem modeling is highly parameterized, and requires a great deal of historical data. Routine model runs can easily use hundreds of gigabytes, even terabytes, of data on tens and perhaps hundreds of parameters. It is a given that no single modeler can or does collect all the required data. All modelers depend upon other scientists, and upon governmental and research agencies, for their data needs.

This is where data access and interoperability become crucial for the success of a project. Having well documented and quality data available in a timely fashion can greatly assist a project, while the converse can bring it to a standstill, leading to a large amount of wasted staff time and resources.

Data accessibility is a complex issue. At best it is an unscientific composite of a variety of factors: technological, legal, cultural, semantic, and economic. In reality it is a concept that most scientists worry about only when they need some data, and mostly never after their project is complete. The exigencies of the vetting, review, and publishing processes overtake the long term view of making one's own data available to others with the same ease and openness that was desired when seeking data from others.

This presentation describes our experience with acquiring data for our carbon modeling efforts, dealing with federal, state, and local agencies, a variety of data formats, some published and some not so easy to find, and documentation that ranges from excellent to non existent. A set of indicators is proposed to place and determine the accessibility of the data we are seeking and the data we are producing, in order to bring some transparency and clarity that can make data acquisition and sharing easier. The paper concludes with a proposal to use free, open, and well recognized data marks such as the Creative Commons CC0 public domain dedication and the CC BY attribution license, which would advertise the openness of scientific data to everyone.

Back to the Carbon Blog