{"id":51,"date":"2020-01-22T19:45:16","date_gmt":"2020-01-22T19:45:16","guid":{"rendered":"https:\/\/ni.cmu.edu\/computing\/?post_type=ht_kb&#038;p=51"},"modified":"2021-03-31T20:35:10","modified_gmt":"2021-03-31T20:35:10","slug":"data-management","status":"publish","type":"ht_kb","link":"https:\/\/ni.cmu.edu\/computing\/knowledge-base\/data-management\/","title":{"rendered":"Data Management"},"content":{"rendered":"<p>As a user, you should be familiar with various volumes on the cluster.\u00a0 The first three (3) volumes below are extreamly important for you to understand their purpose. This will help you manage your files and data effectively.<\/p>\n<ol>\n<li><strong>\/home\/&lt;username&gt;<\/strong> is the user&#8217;s home directory and should only be used to store small files such as configuration files, documents, and source code <strong>[keep this to a maximum of 300GB]. <\/strong> (ZFS Snapshots, ZFS replication, and nightly backup to tape archive using <a href=\"https:\/\/www.teradactyl.com\/products\/teradactyl-software\/atli-options\/\">TiBS<\/a>.)<\/li>\n<li><strong>\/user_data\/&lt;username&gt;<\/strong> space is a 1TB partition for users to save their work. (ZFS Snapshots, ZFS replication, and nightly backed up by School of Computer Science Facilities to tape archive using <a href=\"https:\/\/www.teradactyl.com\/products\/teradactyl-software\/atli-options\/\">TiBS<\/a>.)<\/li>\n<li><strong>\/lab_data\/&lt;labgroupname&gt;<\/strong> is a 10TB partition for lab members to share and save their work.(ZFZ Snapshots and ZFS replication)<\/li>\n<li><strong>\/containers<\/strong> &#8211; some standard singularity images that we provide to our users can be found here.<\/li>\n<\/ol>\n<div class=\"hkb-article__content\">\n<blockquote><p>The storage on the cluster was put in place so users could utilized enterprise grade storage for their computational needs. <strong>The cluster is not intended to be used for a file and\/or backup server.<\/strong> Please do not store your processed data on the cluster.\u00a0 Only data currently being used for processing and current results of cluster processing should\u00a0 reside on the cluster.<\/p><\/blockquote>\n<h2>Data protection &#8211; ZFS snapshots, ZFS replication and Tape backup<\/h2>\n<blockquote><p><strong>DATA LOSS<\/strong>: We care about your data, and we do everything we can to retain and save any and all data whenever possible.\u00a0 We will not be held legally liable for any data loss. We do this as a courtesy to our users, but we offer no guarantees.<\/p><\/blockquote>\n<p>The data protection in place on the cluster is designed to minimize downtime and data loss.\u00a0 Most of this is done without the user really noticing.<\/p>\n<h4>If you need your files restored, you should send David an email\u00a0 with the following information:<\/h4>\n<ol>\n<li>Details on what files\/directories need recovered and include full paths if possible.<\/li>\n<li>Date and which you would like the files to be recovered from.<\/li>\n<li>The location you would like the files to be recover to<\/li>\n<\/ol>\n<h4>In summary, besides for the redundancy we have configured in the RAID on our volumes, we also have the following data protection:<\/h4>\n<ol>\n<li>Tape backup is done nightly. This servers as an archival history of the files, but could be slow to restore.<\/li>\n<li>ZFS snapshot is a feature in the ZFS file system which a point-in-time copy of the file system.\u00a0 This servers as a way to immediately recover accidentally deleted or corrupt files without having to go to tape to recover files.<\/li>\n<li>ZFS replication duplicates the snapshots on the primary storage to a secondary storage device.\u00a0 This servers as a disaster recovery platform for catastrophic events.<\/li>\n<\/ol>\n<h4>ZFS snapshots and your quota.<\/h4>\n<p>Have you received a &#8220;Disk Quota Exceed&#8221; message, delete some files and that failed to resolve the quota problem?<\/p>\n<p>When a snapshot is created, the space is initially shared between the snapshot and the file system, and could also be shared with previous snapshots.\u00a0 As the files in the file system changes due to files being updated or deleted, some of the space in the snapshots becomes unique to the snapshot.<\/p>\n<p>Space deleted by users isn&#8217;t immediately freed because the files deleted are still taking up space in the previous snapshot(s).\u00a0 What is confusing to users is that du, ls and other standard UNIX commands show the files were deleted and they are below their quota, when actually their space is still being used due to the snapshots.\u00a0 To resolve this issue, you will have to send an email request to David asking him to delete the snapshots.\u00a0 Keep in mind that this will eliminate the ZFS snapshot option for file recovery for any files contained in the deleted snapshots.<\/p>\n<\/div>\n<p><!-----------------\nDetails of the various volumes\n\n\/home is a 4.6TB space. This is the volume where your home directory is located (\/home\/&lt;username&gt;). This volume is incrementally backed up, nightly, to the Carnegie Mellon's School of Computer Science tape backup system.The size of this volume limits us to what can be stored on it. In addition, if the volume becomes 100% full, it could bring the complete cluster to a standstill. Therefore, we have enabled quotes on this volume. This means that each user will only be able to use, at maximum, predetermined amount of space. Please limit the files to smaller sized files. (e.g. configuration files, notes and documents, software development programs). I believe the quota will be in the range of 100-200GBs of space.\n\nFor data files and logs which tend to be considerably larger files, we have a large RAID 6 volume. The file system used for this space is a \"zfs\" file system. The <strong>ZFS file system<\/strong> is a relatively new file system that contains features and benefits not found in more traditional UNIX file systems. This volume will NOT be backed by the CMU SCS tape backup system (it is too large of this). The volume has been configured to be reliable, but we will also have a method in place to protect the data in an unlikely event of a complete RAID failure. We will be using a second RAID volume and a the snapshot feature included in the zfs file system. A <strong>snapshot<\/strong> is a read-only copy of a file system or volume. Snapshots can be created, and they initially consume no additional disk space within the pool. However, as data within the active dataset changes, the snapshot consumes disk space by continuing to reference the old data, thus preventing the disk space from being freed. This will give users some protection of the accidental deletion of files. The ZFS filesystems are built on top of virtual storage pools called zpools. We plan to create a zpool for each lab. In addition, we will create a UNIX group for each individual lab and add users in whatever groups they need access to that lab's zpool. Once the zpools are created, they will be automounted through \/data2 (e.g. \/data2\/tarrlab, \/data2\/coaxlab, etc.) Within the zpool, labs can organize their data in whichever way they would like. If files need recovered, we can do so using the snapshot for that zpool.\n\n\n<h4><a name=\"Automounting\"><\/a> Automounting<\/h4>\n\n\nAutomount: As aforementioned, we will automount the zpools in addition to other mount points on the cluster. Automounting is when the system will automatically mount the volume in response to access operations by the user and\/or user programs. When there is a period of time that the mount point is not accessed, the system will unmount the volume.\n\nFor instance, if you were to do a listing (UNIX command 'ls') on \/data2 and the zpool on \/data2, \/data2\/tarrlab had not been accessed for a while, it would NOT show up in the listing. But if you would directly perform a listing on 'ls \/data2\/tarrlab' you would see its contents. When the tarrlab zpool had been directly access ( using the 'ls \/data2\/tarrlab command), it was automatically mounted. It will now be seen using the command 'ls \/data2'\n\n---><\/p>\n<h3>Options for sharing files:<\/h3>\n<p><a href=\"http:\/\/www.cmu.edu\/computing\/accounts\/storage\/\">http:\/\/www.cmu.edu\/computing\/accounts\/storage\/<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As a user, you should be familiar with various volumes on the cluster.\u00a0 The first three (3) volumes below are extreamly important for you to understand their purpose. This will help you manage your files and data effectively. \/home\/&lt;username&gt; is the user&#8217;s home directory and should only be used to&#8230;<\/p>\n","protected":false},"author":1,"comment_status":"closed","ping_status":"closed","template":"","format":"standard","meta":{"footnotes":""},"ht-kb-category":[7],"ht-kb-tag":[],"class_list":["post-51","ht_kb","type-ht_kb","status-publish","format-standard","hentry","ht_kb_category-cluster"],"jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb\/51","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb"}],"about":[{"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/types\/ht_kb"}],"author":[{"embeddable":true,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/comments?post=51"}],"version-history":[{"count":17,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb\/51\/revisions"}],"predecessor-version":[{"id":346,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb\/51\/revisions\/346"}],"wp:attachment":[{"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/media?parent=51"}],"wp:term":[{"taxonomy":"ht_kb_category","embeddable":true,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb-category?post=51"},{"taxonomy":"ht_kb_tag","embeddable":true,"href":"https:\/\/ni.cmu.edu\/computing\/wp-json\/wp\/v2\/ht-kb-tag?post=51"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}