Chapter 7 Keeping Data Safe

Or, how to ensure your data isn’t altered or removed accidentally or intentionally.

7.1 Lesson Objective

Understand and, when possible, implement safeguards to prevent unauthorized or accidental data modification.

An important caveat: sensitive data (including that derived from human participants, about sensitive species or locations, or otherwise protected information) necessitates additional data safety and security measures that will not be covered in this guide.

7.2 Key Terms

  • Cloud-based Storage: A system that lets you save, access, and share your research files securely over the internet from any device, rather than storing them solely on your local computer or external drive. Google Drive and Box are two examples of cloud-based storage (but these storage providers still have physical locations and servers(.
  • On-premises Storage: A system where your research files are saved on physical servers and hardware maintained locally by your institution, accessible only through your campus network or VPN rather than directly over the internet.
  • Authoritative Version: The official copy of a file or dataset that serves as the definitive reference point for your research, ensuring everyone works from the same version.

7.3 Lesson

We have all been there before: you go to review your manuscript draft, and your most recent save didn’t take–or worst case, the document is completely gone. Or, you saved your important presentation to a flash drive, and now you’ve lost that flash drive. Now imagine this is your research data, and you can’t access the recent and time-bound data you painstakingly captured. Not only is this a frustrating setback to the project, it can also feel embarrassing, especially in a team setting. To avoid this, practice robust data security and preservation standards.

If you are entering a lab, you may not have much control over the approach to data storage and backup. But that doesn’t mean you shouldn’t know about it! If you’re entering a project, there are (or should be) structures in place to ensure that data are backed up regularly, version controls to prevent accidental editing, and processes to document your data collection workflows. If there are not, this is a good chance to start the conversation, if that is something you feel comfortable doing.

To the extent you have control over the project data–or at least your data within the project–there are a few rules of thumb you can adopt to help prevent data loss.

7.3.1 Practice 1: Use group storage

Each research lab, center, department, etc. will have access to different types of storage, which will depend on your institution. It’s common for there to be institutional support for cloud-based storage. All storage solutions have options for group accounts–meaning multiple users will have access to the storage. In Google, this looks like a Shared Google Drive (NOT a shared Folder); in Microsoft, this is storage managed through a Teams Shared Drive. While the exact terminology may vary, it is imperative that you contribute your data to the shared or group storage instead of your personal Drive, OneDrive, or computer storage. In particular, when you leave the institution, this will ensure your data remain managed by and accessible to the lab, center, department, etc.

7.3.2 Practice 2: Remember 3-2-1

This best practice suggests that, to avoid data loss or corruption, you should save 3 copies of important files in 2 locations with 1 copy stored remotely (e.g., a cloud storage system, such as Google Drive). This can be challenging, especially when working on a team that is actively updating data files. But don’t be daunted! A relatively straightforward way to implement this is to ensure that you have periodic (e.g., weekly, monthly) backups in place. These can be automated so you don’t have to worry about it.

Regardless of the means you use to implement backups, it is important these are documented and understood so you always know where the authoritative version of the working files lives. This can be accomplished through a robust folder organization and file naming system. Getting into the habit of adding the current date to the beginning of file names can help you both sort files and keep track of the most recent and working version, the raw copies, and everything in between.

7.3.3 Practice 3: Only edit raw data when necessary

After the data collection phase, consider carefully when, if at all, you need to be using the raw and original data (i.e., that which you created or received). Sometimes data cannot be recreated: a weather condition on a certain day a week ago, voter perceptions before an election, health data from a wearable device immediately after a major surgery. You will never be able to get that data back–if you accidentally delete something or change a value, you could have just irrevocably altered your data.

So, when you are conducting analyses, running scripts, using R to mutate, using Excel formulas, etc, make a derivative (copy) of the data for analysis. As with above, though, be sure that you understand and have clearly labeled (such as through file names) which version(s) of the files you are working with.

7.3.4 Practice 4: Save data from key steps

As your project progresses, and you create derivatives to run analyses, be sure to keep data files from key steps of your research, without necessarily keeping every file. Consider your research project as a pipeline: data are collected, created, or derived from a secondary source; then the data will undergo various transformations, as you clean it, analyze it, and visualize it. You don’t need to keep every piece of data created, but be sure that you retain enough data (and documentation!) to provide enough information on the process. This is useful not only for you and your collaborators, but also for ensuring your work is reusable and reproducible.

7.3.5 Practice 5: Keep your guest list up to date

As people transition on and off the project, ensure that access to files, documentation, and servers is added and removed. This is especially important if your data have any sort of sensitivity, but is good practice regardless. When someone leaves the project, even if you think their account will be deactivated by your central IT, you should still plan to remove their account’s access.

7.4 Exercises

Spend a few minutes reviewing the storage for your files. If you / your department / your lab use cloud storage (e.g., Google Drive, One Drive, Box), look over who has access to the files.

  • Do you know everyone on the access list?
  • If you are using local storage (managed by your institution), do you know how often your data are backed up?
  • Do you know if the storage you are using has retention limits that might result in your data being deleted?

Regardless of the storage you are using, consider the documentation for your project. Is it clear what the authoritative version of the files is?

This is also a good time to get to know your IT personnel, if you don’t already! They are there to support you - even if it might not always feel like that.

7.5 Further Readings

  • Data Management Security Tips from James Madison University Office of Research Integrity and Compliance. This resources provides best practices on keeping physical and digital data safe.
  • If you have specific questions, reach out to your IT support.

Get credit for your work >>>