The document discusses data publication and outlines its importance in ensuring scientific integrity and reproducibility. It notes that data publication, with persistent identifiers like DOIs, allows data to be cited similarly to scholarly articles. This facilitates data sharing, access, and credit to data creators. The document reviews why data publication is increasingly important given growing data volumes and open access policies. It provides examples of data citation and outlines the benefits of publishing data, such as increasing research impact. Finally, it discusses requirements for trusted data repositories and the services provided by the IEDA repository to support the full data publication process.
2. December 11, 2013
2
Data Publication: Definition
⢠Data publication with a small ď¨pď§
⢠Sharing data via web sites or submission to databases
⢠Data Publication with the big ď¨Pď§
⢠Publication of data as part of scholarly communication
⢠Citable
This is what we are talking about!
⢠Persistent access
⢠Quality assurance (repository review or peer-review)
3. 3
December 11, 2013
Data Publication: Why?
⢠Data access is key to science
⢠Data are the nucleus of scientific collaboration
re-use
⢠Scientific progress requires community that competes and collaborates
in pursuit of common goals
⢠Without access to the same materials no community exists
⢠Data are needed for scientific replication
reproduce
⢠The value of an article that canât be replicated:?
⢠Scholarly articles are summaries, not the actual research results
⢠Experimental data expensive to verify, observational data impossible
⢠Replication projects show: many published articles cannot be replicated
accessible at
http://www.slideshare.net/
4. 4
December 11, 2013
Data Publication: Why now?
⢠Exponentially increasing data volumes
⢠Rapidly expanding cyberinfrastructure capabilities to mine
and analyze data
⢠New paradigms in publishing
⢠Growing enforcement of policies for open access to
research data
âIn the last decades, rapid technical developments, such
as digital data and high-throughput techniques,
dramatically changed the scholarly publishing paradigm.
This requires new approaches in order to ensure
availability and usability of science data.â
http://www.icsu-wds.org/working-groups/data-publication
5. 5
December 11, 2013
Data Policies
⢠Agencies
⢠Societies
⢠Journals
February 22, 2013
May 9, 2013
8. December 11, 2013
Data Publication: Why?
⢠Ensure proper citation of the data and credit to their
creator(s)
â˘
8
9. December 11, 2013
9
Data Citation Principles
http://www.force11.org/datacitation
âSound, reproducible scholarship rests upon a foundation of
robust, accessible data. For this to be so in practice as well
as theory, data must be accorded due importance in the
practice of scholarship and in the enduring scholarly record.
In other words, data should be considered legitimate, citable
products of research. Data citation, like the citation of other
evidence and sources, is good research practice.â
10. 10
December 11, 2013
Data Citation Principles (Force 11)
⢠Importance: Data should be considered legitimate, citable products of research. Data
â˘
â˘
â˘
â˘
â˘
â˘
â˘
citations should be accorded the same importance in the scholarly record as citations of
other research objects, such as publications.
Credit and attribution: Data citations should facilitate giving scholarly credit and normative
and legal attribution to all contributors to the data, recognizing that a single style or
mechanism of attribution may not be applicable to all data.
Evidence: Where a specific claim rests upon data, the corresponding data citation should be
provided.
Unique Identification: A data citation should include a persistent method for identification
that is machine actionable, globally unique, and widely used by a community.
Access: Data citations should facilitate access to the data themselves and to such
associated metadata, documentation, and other materials, as are necessary for both humans
and machines to make informed use of the referenced data.
Persistence: Metadata describing the data, and unique identifiers should persist, even
beyond the lifespan of the data they describe.
Versioning and granularity: Data citations should facilitate identification and access to
different versions and/or subsets of data. Citations should include sufficient detail to verifiably
link the citing work to the portion and version of data cited.
Interoperability and flexibility: Data citation methods should be sufficiently flexible to
accommodate the variant practices among communities but should not differ so much that
they compromise interoperability of data citation practices across communities.
http://www.force11.org/datacitation
13. December 11, 2013
13
Data Citation Example
⢠Global Multi-Resolution Topography (GMRT) synthesis described by
â˘
â˘
â˘
â˘
Ryan et al., 2009.
Its data DOI is 10.1594/IEDA/100001.
Its citation would be: Ryan, William B.F. (2009): Global MultiResolution Topography (GMRT) synthesis. Integrated Earth Data
Applications (IEDA). http://dx.doi.org/10.1594/IEDA/100001
It can be accessed at the
URL: http://dx.doi.org/10.1594/IEDA/100001.
The data set DOI is different from the DOI for the publication that cites
the data set, doi: 10.1029/2008GC002332.
14. December 11, 2013
14
Data Publication: Benefits
⢠Scientific integrity
⢠publishing your data and citing its location in published research
papers can allow others to replicate, validate, or correct your
results, thereby improving the scientific record.
⢠Increase the impact of your research
⢠those who make use of your data and cite it in their own research
will help to increase your impact within your field and beyond it.
Users of your data may include those in other disciplines, sectors,
and countries.
⢠Preserve your data for your own future use
⢠by preparing your data for sharing with others, you will benefit by
being able to identify, retrieve, and understand the data yourself
after you have lost familiarity with it, perhaps several years hence.
15. 15
December 11, 2013
Data Publication: Benefits
from: Heather A. Piwowar, âSharing Detailed Research Data Is Associated with Increased Citation Rateâ
http://precedings.nature.com/documents/361/version/1
16. December 11, 2013
RDA/WDS Publishing Data IG
⢠address practical aspects in publishing research data
16
17. 17
December 11, 2013
Data Publication: Options
Conventional
publication
Institutional
Repositories
Disciplinary
Repositories
Data Article
18. 18
December 11, 2013
Data Publication âBest Practiceâ
Trusted Data
Repository
Journal
Data
File
Data
Description
Reciprocal citation by DOI
19. December 11, 2013
19
Journal Guidelines: Examples
âElsevier encourages authors to deposit raw experimental
data at relevant data repositories.â
http://www.elsevier.com/about/content-innovation/database-linking
âAGU encourages authors to identify and archive
their data in approved data centers.â
AGU Data Policy, approved by AGU Council Dec 8, 2013
20. 20
Repository Requirements
âThe use of published digital data, like the use of digitally published literature, depends
upon the ability to identify, authenticate, locate, access, and interpret them.â
Report of the CODATA - ICSTI Task Group on Data Citation Standards and Practices: âOUT OF CITE, OUT OF MINDâ
⢠Open access
⢠Long-term preservation
⢠Persistent & unique identification
⢠Data quality assurance (peer-review?)
âĽDeposition of data in âTrusted
Repositoriesâ
22. 22
Domain-Specific Repositories
⢠Are best poised to ensure âFitness for Re-useâ through
domain-specific data stewardship
⢠Must ensure professional data curation services
⢠Long-term archiving & access
⢠Persistent, unique identification
⢠Discoverability (metadata registration)
⢠Must integrate with the âscholarly communication
ecosystemâ
23. 23
Adding Value: Domain-Specific Data
Stewardship
⢠Development, maintenance, and promotion of domain-
specific, community-based standards for data and metadata
⢠Provenance documentation, uncertainties, semantics (vocabularies,
taxonomy), formats
⢠Domain-specific guidelines, software tools, and user
support/training that facilitate data submission
⢠Harmonization & integration of data for advanced analysis
⢠Mapping of data to standards-based interfaces for
interoperability
26. December 11, 2013
26
IEDA Infrastructure
⢠Cooperative Agreement with NSF
⢠Sustainable funding
⢠Formal community governance & guidance
⢠Disciplinary expertise
⢠Professional data management policies & procedures
⢠Persistent identification of data & samples (DOI, IGSN)
⢠Standards-compliant metadata catalog
⢠Long-term archiving agreements with National Geophysical Data
Center & Columbia University Libraries
⢠Risk management
⢠âAccreditationâ as member of the World Data System
27. December 11, 2013
27
IEDA Repository Services
⢠Rich and standards-compliant metadata catalog
⢠Data publication with DataCite (Registration with DOI)
⢠Online data submission tools
⢠QA/QC of datasets and metadata
⢠Storage & risk management of submitted data
⢠Long-term preservation (via partners)
⢠Cross-referencing with journals, data citation index, etc.
30. December 11, 2013
IEDA Repository Services
⢠Investigator support
⢠Data Management Plan tool
⢠Data Compliance Report tool
⢠User Support
⢠Online submission tools
⢠Data templates
⢠Tutorials & Help pages
⢠YouTube videos
⢠Personal assistance (info@iedadata.org)
⢠Workshops, webinars, etc.
http://www.iedadata.org/help
30
31. 31
December 11, 2013
IEDA Data Publication Process
User Submission
Data
Review
IEDA Data
Managers
Publication
IEDA Repository
Integration
IEDA Data
Managers
Synthesis
databases
DOI linking
Journal
Manuscript
Editors
Portal
DOI linking