Where should I share my Data, Materials, Code now?
This Article Is Licensed Under CCO For Maximum Reuse.
The following is a Table of Contents that links directly to specific sections within the guide:
Selecting an appropriate repository
After November 16, 2026, the OSF will no longer function as a generalist repository with the disabling of the OSF Projects workflow.
Your existing public projects will remain discoverable and DOIs will continue to resolve. COS recommends that you do not move public content if the content in the project is complete/static; instead, you may allow COS to host the content. For projects that you choose to move, this guide describes several options for you to consider. You should consult with any additional contributors on your projects to ensure that you are not creating unintentional copies of public content.
Some options for other repositories are below, but you might want to first check with your institution for their recommendations. Many universities and other research facilities own or support specific infrastructure for hosting papers, preprints, datasets, and other materials. Using your institution's repository may be the best option. If you are at a university, first check with the university library or office of research. More guidelines for how to select an appropriate repository are also available further down in this guide.
If an institutional repository is not available or appropriate, the following table lists a few repositories that support sharing of data, materials, code, and other research outputs. Each offers a basic description, key features, and other information that may be important to determining the best place for your content.
Table of repositories:
| Repository | Type | Description | Access Options |
Cost (as of Aug 2026) |
Provides DOI | Open Source? | Non-Profit? |
|---|---|---|---|---|---|---|---|
| Dataverse | Data | A robust repository for research data across all disciplines, hosted across multiple institutions globally. | Embargo at file level, repository managed access, preview links | Free, institutional membership | YES | YES | YES |
| Dryad | Data | Focused on research data underlying published findings. Best suited for datasets that accompany a journal article. | Embargo with journal approval, preview links | Data Publishing charge (DPC), institutional membership | YES | YES | YES |
| Figshare | Data, code, images, | A repository that accepts datasets, code, images, presentations, and more from any discipline, and supports sharing multiple output types from a single research project. | Embargo at file and dataset level, preview links | Free for researchers, institutional memberships, organizational memberships | YES | YES | NO |
| Mendeley (Data) | Data | A cloud-based collaborative repository for data storage, linked to Mendeley Reference Manager and select Elsevier journals. | Embargo at file and dataset level, user managed access, preview links during embargo | Free for researchers, institutional memberships | YES | NO | NO |
| Vivli | Data | Specialized for clinical and biomedical research, particularly individual participant-level data from clinical trials. | Embargo at file and dataset level, repository managed access, access using secure research environment | Managed access, institutional memberships | YES | NO | YES |
| Zenodo | Data, code, images, research outputs | A free, general-purpose repository hosted by CERN for researchers. Accepts any research output type and assigns a permanent DOI to every object. | Embargo at dataset level, preview links | Free for researchers, communities and Institutional login integration | YES | YES | YES |
| ICPSR | Data | ICPSR is self-publishing repository for social, behavioral, and health sciences research data | Embargo at dataset level, repository managed access | ICPSR is partially free, depending on your institutional affiliation and the specific dataset you want to access. It is free to search the site, and many public-use datasets are open to everyone, but others require a paid membership or a per-study fee. | YES | NO | YES |
| Qualitative Data Repository (QDR) | Data | The Qualitative Data Repository (QDR) is a dedicated archive for storing and sharing digital data (and accompanying documentation) generated or collected through qualitative and multi-method research in the social sciences and related disciplines. | Embargo at dataset level, repository managed access | Typical deposit fees range from $750-2500; deposit-fee assistance is possible. Institutional members can often deposit for free. | YES | YES | YES |
| Protocols.io | Methods | A secure platform to develop, share, and discover reproducible research methods, protocols, and workflows across teams and the global scientific community. | User managed access | Public, user-deposited, content published via protocols.io is open access and always free to read and publish and a user can have up to 2 private protocols. They do charge for additional private use, collaboration and some additional premium functionalities, as detailed on the 'plans' page. | YES | NO | NO |
| Codeberg | Code, software | Codeberg is a non-profit community-led organization that aims to help free and open source projects prosper by giving them a safe and friendly home. | User managed access | Free, though you can donate | No; they have a guide for creating citable code | YES | YES |
| Forejo | Code, software | Forgejo is a self-hosted lightweight software forge. | User managed access | Free, though you can donate | No; users can follow Codeberg’s guide for creating citable code | YES | YES |
| GitHub | Code, software | GitHub is a proprietary developer platform that allows developers to create, store, manage, and share their code. | User managed access | Free with paid options | No; they have a guide for creating citable code | NO | NO |
| GitLab | Code, software | GitLab is a proprietary software forge. | User managed access | Free with paid options | No; users can follow GitHub’s guide for creating citable code | NO | NO |
How to select a repository
Repositories provide a place to store, share, and discover research outputs. Different repositories specialize in different types of content, research disciplines, and workflows.
As you decide where to store your research content, consider what you are sharing, who needs access to it, and what requirements apply to your project. This guide will help you identify the type of repository that best fits your needs.
Consider the following questions:
| Does your funder, institution, or journal have requirements or offer services? |
Some funders, institutions, and journals require researchers to deposit certain outputs in specific repositories or repositories that meet defined standards. Check whether your project has requirements related to:
Your institution may also host infrastructure for hosting many types of materials. Depositing in specific repositories may be required for some types of materials at some institutions. Check with your institution's library or research office to find out if there are options for you. |
| What type of research content are you sharing? |
Different repositories are designed to support different types of research outputs, including:
|
| What kind of repository is most appropriate? |
Talk to your institutional librarian about who should host your data and other research outputs. Generally, your options fall into one of four categories:
|
| Who should be able to find and use your content? |
Consider whether your content should be:
|
Managing Private Data with Controlled Access
If you have private data that requires managed or controlled access, the main choice is: Do you want the materials to be publicly discoverable and available by request, or known only to people who have the link?
| If you have… | Start with… | Why |
| Materials that are ready to be publicly discoverable and available by request |
A research repository with restricted access services |
A public record supports discovery, citation, preservation, and an access-request process. |
| Materials that should not be publicly discoverable and should be accessed only through a shared link |
A non-repository secure content-management platform such as Box, Dropbox, Google Drive, One Drive, etc. |
Link-based access without a public record and public requests. These platforms can also support collaborator uploads and editing. |
| Sensitive or regulated datasets that should be publicly discoverable but available only through a restricted-use process |
A research repository with managed, restricted access services |
A public record supports discovery while the repository manages curation and secure access. |
Additional Resources:
- NIH-supported Data Repositories for Sharing Scientific Data, generalist repositories and domain/discipline specific repositories
- A comparison chart and comparison flow chart for generalist repositories
- Re3data
- Desirable characteristics of data repositories for federally funded research by the National Science and Technology Council
- Selecting a Data Repository by the National Institutes of Health
- Data Repository Guidance by Nature Journals
Preserve and Share Data with a Repository
A Repository is a tool to share, preserve, and discover research outputs, including but not limited to data or datasets.
Types of repositories:
- Data Repository: A central database in a business used to aggregate and manage information.
- Institutional Repository: A digital archive maintained by a university or library to store and share research papers and datasets.
- Generalist Repository: Cross-disciplinary or generalist repositories allow data from all areas of study. They are also often used to house conference proceedings or other products such as handouts and educational materials. They may offer options for collaboration, and public vs. private settings for your data. Examples include Dataverse, Figshare, Zenodo, and Dryad. A comparison chart and comparison flow chart for generalist repositories can help you make a selection.
- Discipline-specific Repository: is a digital archive tailored to a specific academic field (i.e. Clinical trials- clinicaltrials.gov, Vivli; Genomics data repositories- Gene Expression Omnibus (GEO), GenBank ) Discipline-specific repositories collect research outputs from a particular field of study and may be broad or granular. A specific organism or type of research may have its own repository. Best practice for data sharing is to put "like with like" to facilitate discoverability and data reproducibility. Some funders, like NIH, may mandate a particular disciplinary repository based on the type of research, such as genomic data.
- Software/Code Repository: A digital storage space where developers keep codebases, documentation, and a project's revision history. It allows teams to track changes, collaborate, and merge code. (i.e Github, Gitlab, or Zenodo).
Contents
1. What is a Repository?
- Purpose: A repository is a tool to share, preserve, and discover research outputs, including but not limited to data or datasets. While workflows and processes will vary across repositories, generally speaking, researchers submit and describe their own data which is then ingested into the repository for storage. Other researchers can then download - or request to download - the data directly from the repository.
- Why put data in a repository?
- Comply with funder requirements
- Long-term, safe, managed storage and backup
- Make your data findable and accessible to others, contributing to a culture of open scholarship, sharing, and reproducibility. Share "like with like" when possible.
- Ease of collaboration with others
- Receive a persistent identifier (DOI) for your data
2. Preparing Your Data for Deposit
Preparing your data for future use and long-term storage at the beginning of your project can save you time and frustration at the end. Using stable file formats, researching if your repository requires a fee or has requirements for deposit, and sharing well-documented datasets can all prepare you for successful long-term storage – and future use – of your data.
Furthermore, not everything needs to be shared or preserved! Iterative versions of your data or code – while valuable to your process – don’t necessarily need to be preserved. If a final, clean version of your data or code accurately represents information in your publication, that is what you should focus on sharing. Your final versions should be shared in stable file formats with an associated readme file.
- Data Prepared – Verify that your data files are ready, documented, and meet any ethical or privacy requirements. This includes de-identifying any participant data. If you want your data to undergo blind peer review, you will also want to remove contributor information from any data files.
- Select an appropriate repository. Carefully consider your needs when you are looking for a repository, and make sure that it meets your criteria. You also want to check that it meets or complies with any funder requirements for your grant, if applicable.
- What disciplines of data do they accept? Which formats/file types?
- Is there a size limit on overall storage or on individual files?
- Is there a cost? What is the cost structure?
- Does the repository provide data curation if you need it, and if so, is there a cost?
- Are there settings for collaboration?
- What data repository is standard in your lab, your field, or your discipline?
- Do you work with datasets that need access restrictions? This includes datasets with personally identifiable information (PII), confidential information, sensitive information, or protected information.
- Other questions to think about.
- What file formats will you use to store and share your data?
- How long will your data be kept in your selected repository? How long will it be kept in your lab?
- Are there any additional resources or requirements to prepare your data for deposit?
- Determine what data needs to be shared:
- Is it observational data that cannot be reproduced? Observational data that cannot be reproduced may need to be stored into perpetuity.
- Is it difficult to reproduce? Did your analyses require a lot of supercomputing time, or are very specialized instruments needed for their creation? Simulations may only require the source code, initial conditions, and verification data. However, if simulations are time- or resource-intensive to produce, the models or results may need to be stored into perpetuity.
- Do they underlie a publication?
- Are they null results that tell you something about a process or a procedure?
- Five steps to decide what data to keep A checklist from the Digital Curation Centre to help you appraise your data and determine what to keep.
- README files README files contain detailed explanations of your data to enable future use.
-
Giving context to your data is of immense benefit to any future users of your data – including yourself! Memories are fallible, and documenting important aspects of your data will help prevent you from second-guessing your work.
Context can be as simple as a README file with a few relevant details, or as complex as a robust metadata schema with granular details to greatly enhance reuse. Data dictionaries go hand-in-hand with giving context to your data. Plan to have some or all of this documentation as part of your data package for long-term storage.
Choosing a Repository
When sharing your research data code, and documentation, there are many repositories to choose from. Some repositories are domain specific, and focus on a very narrow type of research. Other repositories, known as generalists, accept data from multiple disciplines and in various file formats. Another aspect of repositories is the ease with which data is accessible (controlled-access or open).
Qualities to Look For in a Repository
- NIH guidance on selecting a repository: List of desirable characteristics to look for when choosing a repository to manage and share data resulting from Federally funded research.
- Nature Data Repository guidance: Because Nature does not host data, this is guide for researchers who are publishing with them on qualities to look for when determining where to share their data.
Desirable Characteristics for All Data Repositories
When choosing a repository to manage and share data resulting from Federally funded research, here are some desirable characteristics to look for:
- Unique Persistent Identifiers: Assigns datasets a citable, unique persistent identifier, such as a digital object identifier (DOI) or accession number, to support data discovery, reporting, and research assessment. The identifier points to a persistent landing page that remains accessible even if the dataset is de-accessioned or no longer available.
- Long-Term Sustainability: Has a plan for long-term management of data, including maintaining integrity, authenticity, and availability of datasets; building on a stable technical infrastructure and funding plans; and having contingency plans to ensure data are available and maintained during and after unforeseen events.
- Metadata: Ensures datasets are accompanied by metadata to enable discovery, reuse, and citation of datasets, using schema that are appropriate to, and ideally widely used across, the community(ies) the repository serves. Domain-specific repositories would generally have more detailed metadata than generalist repositories.
- Curation and Quality Assurance: Provides, or has a mechanism for others to provide, expert curation and quality assurance to improve the accuracy and integrity of datasets and metadata.
- Free and Easy Access: Provides broad, equitable, and maximally open access to datasets and their metadata free of charge in a timely manner after submission, consistent with legal and ethical limits required to maintain privacy and confidentiality, Tribal sovereignty, and protection of other sensitive data.
- Broad and Measured Reuse: Makes datasets and their metadata available with broadest possible terms of reuse; and provides the ability to measure attribution, citation, and reuse of data (i.e., through assignment of adequate metadata and unique PIDs).
- Clear Use Guidance: Provides accompanying documentation describing terms of dataset access and use (e.g., particular licenses, need for approval by a data use committee).
- Security and Integrity: Has documented measures in place to meet generally accepted criteria for preventing unauthorized access to, modification of, or release of data, with levels of security that are appropriate to the sensitivity of data.
- Confidentiality: Has documented capabilities for ensuring that administrative, technical, and physical safeguards are employed to comply with applicable confidentiality, risk management, and continuous monitoring requirements for sensitive data.
- Common Format: Allows datasets and metadata downloaded, accessed, or exported from the repository to be in widely used, preferably non-proprietary, formats consistent with those used in the community(ies) the repository serves.
- Provenance: Has mechanisms in place to record the origin, chain of custody, and any modifications to submitted datasets and metadata.
- Retention Policy: Provides documentation on policies for data retention within the repository.
For a list of NIH-supported data repositories NIH-supported data repositories see Repositories for Sharing Scientific Data, generalist repositories and domain/discipline specific. A comparison chart and comparison flow chart for generalist repositories are also provided.
This Article Is Licensed Under CCO For Maximum Reuse.