24 ms·
Amazon Glacier
http://aws.amazon.com/glacier/
- klodolph 14y agoI get a 404, and when I go to http://aws.amazon.com/ http://aws.amazon.com/ I see nothing about glacier in the news feed.
- jeffchuber 14y agoTry this link http://aws.amazon.com/glacier/?utm_source=AWS&utm_medium=website&utm_campaign=BA_glacier_launch http://aws.amazon.com/glacier/?utm_source=AWS&utm_medium...
- benguild 14y agoGreat. Now if someone would simply create a badass client for this for Mac (like Backblaze or Mozy) … we'd be in business. :)
- sa1f 14y agoArq's support is enough. http://www.haystacksoftware.com/arq/ http://www.haystacksoftware.com/arq/
- benguild 14y agoIt already supports it?
- jordibunster 14y agoKinda: "In the coming months, Amazon S3 will introduce an option that will allow customers to seamlessly move data between Amazon S3 and Amazon Glacier based on data lifecycle policies."
- ojilles 14y agoURL? Can't find this on the website, blog or twitter account.
- icebraining 14y agoIt's in the Glacier FAQ: http://aws.amazon.com/glacier/faqs/#How_should_I_choose_between_Amazon_Glacier_and_Amazon_S3 http://aws.amazon.com/glacier/faqs/#How_should_I_choose_betw...
- benguild 14y agoIt already supports it?
- terhechte 14y agoThis is fantastic. I've long searched for a solution like that. This is really suitable for a remote backup that only needs to be accessed if something really bad happens (i.e. a fire breaking out, etc). I'm a lone entrepreneur, so I do have backup hard disks here, but being able to additionally save this data in the cloud is great. I'm often creating pretty big media assets, so Dropbox doesn't necessarily offer enough space or is - for me - too expensive in the 500gb version (i.e. $50 a month). Glacier would be $10 a month for 1 terabyte. Fantastic.
- nemesisj 14y agoAgreed - I'm really happy about this. I have a home NAS solution that is a few TB and it's too expensive to store on S3. This is perfect to prevent the "house burned down" scenario on very large storage devices!
- Dysanovic 14y agoI'm trying to work out the best way to back-up my NAS to the cloud just in case my house burns down. It's really annoying how lots of cloud storage solutions dont cater for home users with a NAS. I agree that this could be a good, cheap way of backing stuff up.
- brunnsbe 14y agoThinking exactly the same, Glacier would be a good place for a backup off family photos and videos.
- pontifier 14y agoI recently heard about a startup (spacemonkey) that will be offering 1TB with redundancy and no access delay for $10/month. The way they do it is really clever as well. Cloud storage has always seemed way too expensive to me, but these lower prices have me re-evaluating that.
- yungchin 14y agoThe only issue I see is that verifying archive integrity (you don't want to find out the archive was bad after you lost the local backup...) would be somewhat complicated, given their retrieval policies. Also, the billing for data-transfer out plus peak retrievals sounds so convoluted, I can't begin to work out what a regular test-restore procedure would cost me. Nevertheless, it's some exciting progress in remote storage!
- nemesisj 14y agoI had a quick skim through the marketing stuff and the FAQs and didn't see anywhere that actually details what the backend of this is. I'd be curious if they're actually using tape, older machines, Backblaze pods, etc. I guess if it's the latter, the time to recover could be an artificial barrier to prevent people from getting cute.
- dabeeeenster 14y agoI think Amazon have their own "backblaze pods"...
- tommoor 14y agoagreed, it would be great to know how this is running from a hardware point of view - just out of personal interest :-)
- freehunter 14y agoSomeone further up mentioned a very plausible (in my experience) answer. Magnetic tape, using hard drive arrays as RAM. The wait time in this situation would be the time needed to complete all the current tasks waiting to be written/read in the queue before your data is written from tape to hard drive so you can access it.
- flyt 14y agoAmazon specifically said it does not use tapes.
- flyt 14y agoIt appears to use S3 as its basic backend. My guess is that S3 has been modified to have "zones" of data storage that can be allocated for Glacier. Once these zones have been filled with data (and of course that data is replicated to another region) the hard drives are spun down and essentially turned off. This is why the cost of retrieval is so high: every time they need to pull data the drives need to be spun back up (including drives holding data for people other than you), accessed, pulled from, then spun back down and put to sleep. Doing this frequently will put more wear and tear on the components and cost Amazon money in power utilization. As is Glacier should be extremely cheap for AWS to operate, regardless of the total amount of data stored with it. Beyond the initial cost of purchasing hard drives, installing, and configuring them the usual ongoing maintenance and power requirements go away.
- sharth 14y agoSince this is built on top of S3, I'd love to be able to store EBS Snapshots in this.
- kdsudac 14y agoWow, order of magnitude cheaper than S3 (1 cent per month vs. 12-14 cents for S3). Data transfer priced the same. Drawback is "jobs typically complete in 3.5 to 4.5 hours." Seeing as how people tend to be pack rats, I can see this being huge.
- martey 14y agoI think one of the most interesting parts of this is how they plan to ensure that people do not use it for transient backup: https://aws.amazon.com/glacier/faqs/#How_am_I_charged_for_deleting_data_that_is_less_than_3_months_old https://aws.amazon.com/glacier/faqs/#How_am_I_charged_for_de... Deleting data from Amazon Glacier is free if the archive being deleted has been stored for three months or longer. If an archive is deleted within three months of being uploaded, you will be charged an early deletion fee. In the US East (Northern Virginia) Region, you would be charged a prorated early deletion fee of $0.03 per gigabyte deleted within three months.
- kdsudac 14y agoNice catch, that's an important piece of fine print.
- cayblood 14y agoSo I guess that means it would still work well for a scheme like time machine uses, where incremental changes are added but deletions are simply made note of. At least I think that's how it works.
- rektide 14y agoAnd yet they support a max of 100 vaults per account, so some roll behind recompaction of incrementals is still necessary.
- Ecio78 14y agoIt will be interesting to know if they will upgrade AWS Storage Gateway to use this kind of backend instead of S3 http://aws.amazon.com/storagegateway/ http://aws.amazon.com/storagegateway/
- regularfry 14y agoWhat does "99.999999999% durability" mean? Does it mean they allow themselves to lose one byte per terabyte on average?
- Swizec 14y agoI think it means there is a small chance enough hard drives might fail at the same time that there happen to be no backups of those drives. They make so many backups so quickly that there is only a 0.00000000001% (I didn't count the zeros) chance of this occuring.
- gjm11 14y agoWhich of course means that (if they're telling the truth) the probability of losing your data mostly comes from really big events: collapse of civilization, global thermonuclear war, Amazon being bought by some entity that just wants to melt its servers down for scrap, etc. (Whose probability is clearly a lot more than 10^-11 per year; the big bang was only on the order of 10^10 years ago.)
- justinsb 14y agoThere's some clever wordplay/marketing here... "designed to provide 99.99..99%" means that the theoretical model of the system tells you that you lose 1 in X files per year when everything is working as modeled (e.g. "disks fail at expected rate as independent random variables"). If something not in the model goes wrong (e.g. power goes out, a bug in S3 code), data can be lost above and beyond this "designed" percentage. The actual probability of data loss is therefore much, much higher than this theoretical percentage. A more comical way to look at it: The percentage is actually AWS saying "to keep costs low, we plan to lose this many files per year"; when we screw up and things don't go quite to plan, we lose a _lot_ more.
- Jabbles 14y agoper object. So although the chance of losing any particular object is tiny, the chance of you losing something is proportional† to the number of objects. Still extremely small. †roughly proportional if you have << 1e11 objects
- Swizec 14y agoThis is amazing! Just ~3 weeks ago I finally broke down and started using S3 to store my digital photos and other such crap. Sure hope transferring 10 or 20 gigabytes of data from S3 to Glacier is easy.
- ukd1 14y ago11 9's. Impressive.
- phil 14y agoStorage experts: I'd love to know more about what might be backing this service. What kind of system has Amazon most likely built that takes 3-4 hours to perform retrieval? What are some examples of similar systems, and where are they installed?
- jpalomaki 14y agoCould be for example some tape robot where you can have huge amounts of tapes in the storage but only have few devices for reading/writing them. With tape you can't really stream the data to the web. Instead you would probably first copy it somewhere. If there is lots of data, say few terabytes, even this process takes some time. Or in case they are using regular hard drives, you might want to have this kind of time limits in order to pool requests going to a specific set of drives. This would enable them to power down the drives for longer periods of time. The 3-4 hour estimate may also be artificial. Even if you can in most cases retrieve the data faster, it would be good to give an estimate you can always meet. They might also want to differentiate this more clearly from standard S3. And we should not forget that it does take time to transfer for example one terabyte of data over network.
- tezza 14y agoTypically they are tiered. There'll be a near-line HDD array. This is for the recent content and content they profile as being common-access. Then there'll be a robotic tape library. Any restore request will go in a queue annd when an arm-tapedrive becomes free they'll seek to the data and read it into the HDD array. Waiting for a slot with the robot arm - tape drive is what will take 4 hours. EMC(kinda), Fujitsu etc make these. http://en.wikipedia.org/wiki/Tape_library http://en.wikipedia.org/wiki/Tape_library http://www.theregister.co.uk/2012/06/26/emc_tape_sucks_no_more/ http://www.theregister.co.uk/2012/06/26/emc_tape_sucks_no_mo...
- nodata 14y agoWouldn't there also need to be a lot of logic to prevent fragmentation? You'd probably want data from one user near other data from that user, i.e. on the same tape.
- ghshephard 14y agoWhat's particularly awesome, is that this likely represents an upper bound on cost. It will only go down as time goes on.
- zach 14y agoAmazon Glacier is an extremely low-cost, pay-as-you-go storage service that can cost as little as $0.01 per gigabyte per month. What would be absolutely fascinating is a pay-before-you-go storage service — data cryonics. Paying $12 to store a gigabyte of data for 100 years seems like a pretty intriguing deal as we emerge from an era of bit rot.
- simonw 14y agoI really want that service.
- natep 14y ago> Paying $12 to store a gigabyte of data for 100 years seems like a pretty intriguing deal as we emerge from an era of bit rot. As long as that data is decode-able and more importantly, find-able (out of all the GBs frozen for 100 years, why would you want to look at any particular one of them?).
- jaggederest 14y agoTo be fair, if improvements in hardware and software continue at the rate they have been, or some moderate percentage thereof, in 100 years it will be no problem to trawl a few exabytes of data for anything interesting.
- sp332 14y agoThere's a blog that's analyzing Geocities, that's about 1 terabyte of 1 KB files. http://contemporary-home-computing.org/1tb/ http://contemporary-home-computing.org/1tb/ The analysis tracks changes in template design, follows modifications to logos and gifs, and unearths collections of shrines to dead children etc. But that's from when it was harder to make and upload data, so people only put meaningful (to them) stuff online. These days we'd have a hundred thousand copies of a few popular MP3's and everyone's crappy digital photos. The percentage of meaningful stuff would be a lot lower.
- bradfa 14y agoAs long as that data is decode-able and more importantly, find-able (out of all the GBs frozen for 100 years, why would you want to look at any particular one of them?) I'd store my pictures there. Finding old pictures of grandparents when they were little, or even older stuff, is amazing. Wouldn't it be cool if my descendants could still look at pictures of my family in 100 years? Provided that downloading from this 100 year store is something I could do X times per year, and so long as I could append more data to it over time, it's an interesting business model.
- kristofferR 14y agoI hope they can get the access time down from 3-5 hours to about 1 hour - that's the difference for me between it being a viable alternative for storing backups of my client's web sites or not. I might create a script that uploads everything to Glacier and just keeps a couple of the latest backups on S3 though.
- bbgm 14y agoPer Werner's blog post [1] "in the coming months, Amazon S3 will introduce an option that will allow customers to seamlessly move data between Amazon S3 and Amazon Glacier based on data lifecycle policies." 1. http://www.allthingsdistributed.com/2012/08/amazon-glacier.html http://www.allthingsdistributed.com/2012/08/amazon-glacier.h...
- ibotty 14y agothat's certainly interesting. as there will be migration from s3 to glacier, it would be nice if tarsnap had an option to store only the (say) last week in s3 (with .3$/gb/month) and the rest in glacier (with, say, .03$/gb/month). that would certainly be very nice. cperciva, what do you think?
- cperciva 14y agoI can't see any way for Tarsnap to use this right now. When you create a new archive, you're only uploading new blocks of data; the server has no way of knowing which old blocks of data are being re-used. As a result, storing any significant portion of a user's data in Amazon Glacier would mean that all archive extracts would need to go out to Glacier for data... Also, with Tarsnap's average block size (~ 64 kB uncompressed, typically ~ 32 kB compressed) the 50 microdollar cost per Glacier RETRIEVAL request means that I'd need to bump the pricing for tarsnap downloads up to about $1.75 / GB just to cover the AWS costs. I may find a use for Glacier at some point, but it's not something Tarsnap is going to be using in the near future.
- PanMan 14y agoWhile I have no idea how you would fit it in your current infrastructure, I certainly see a (BIG) use-case for: I have this 100 GB, store it somewhere safe (in glacier), I won't need it for the next year (unless my house burns down). I agree that is a bit different from ongoing daily backups with changes, but its also not THAT different from a customer perspective. That it doesn't fit with how you store blocks on the backend won't matter to a lot of customers.
- cperciva 14y agoOh, I absolutely agree that Glacier has lots of great use cases. I wish Tarsnap was able to make good use of it.
- ibotty 14y agoi understand. i hoped tarsnap knew what blocks (do not) get reused. it's unfortunate, because some backups happen to just lie around for very long. it would be nice to take advantage of (the low cost of) glacier for that. that said, if it's not possible with tarsnap now, it's not possible now. :D. if you find a satisfying possibility to incorporate it in the new backend(s) design (if that's fixable in the backend(s) alone), i'd surely be pleased.
- WalterBright 14y agoI'm curious about data security. It says it is encrypted with AES. But is it encrypted locally and the encrypted files are transferred? I.e. does Amazon ever see the encryption keys? Or is the only way to encrypt it yourself, and then transfer it?
- zhoutong 14y agoAFAIK the keys are managed by Amazon, just like S3. It's more for compliance reasons rather than real security. Encryption still has to be done yourself to protect the data.
- guan 14y agoIt does protect you against situations where Amazon loses disks or tapes, or disposes of them improperly, or they are stolen.
- ghshephard 14y agoI'm a long time user of backblaze, and I'm a big fan of the product - it does a great job of always making sure my working documents are backed up, particularly when I'm traveling overseas, and my laptop is more vulnerable to theft or damage. With that said - Backblaze is optimized for working documents - and the default "exclusion" list makes it clear they don't want to be backing up your "wab~,vmc,vhd,vo1,vo2,vsv,vud,vmdk,vmsn,vmsd,hdd,vdi,vmwarevm,nvram,vmx,vmem,iso,dmg,sparseimage,sys,cab,exe,msi,dll,dl_,wim,ost,o,log,m4v" files. They also don't want to backup your /applications, /library, /etc, and so on locations. They also make it clear that backing up a NAS is not the target case for their service. I can live with that - because, honestly, it's $4/month, and my goal is to keep my working files backed up. System Image backups, I've been using Super Duper to a $50 external hard drive. Glacier + a product like http://www.haystacksoftware.com/arq/ http://www.haystacksoftware.com/arq/ means I get the benefit of both worlds - Amazon will be fine with me dropping my entire 256 Gigabyte Drive onto Glacier (total cost - $2.56/month) and I get the benefit of off site backup. The world is about to get a whole lot simpler (and inexpensive) for backups.
- jwr 14y agoWhat's even more important, you will be able to encrypt your backups without having to disclose the encryption key in case you ever need to restore (client-side encryption and decryption). This is not the case with Backblaze, which is why I switched to CrashPlan — but I'm still looking for other solutions.
- OoTheNigerian 14y agoWhat I find fascinating with Amazon's infrastructure push is the the successful 'homonymization' of their brand name Amazon. Amazon simultaneously stands for ecommerce and web infrastructure depending on the context. e.g "Hey I want to host my server".. "Why don't you try Amazon". "Do you know where I can get a fair priced laptop?" "Check Amazon". Is there any other brand that has done this successfully? Edit: I should have specified internet brand.
- masterzora 14y agoThis actually confused me when I saw the headline. I saw 'Amazon' and, despite having been working with AWS all night, thought ecommerce first and couldn't figure out what they could name 'Glacier'. I was kind-of-not-really hoping it was going to be a new shipping option guaranteed to take a long time. That said, the actual service looks solid.
- waterlesscloud 14y agoI'm completely amused that there's an internet service named "Glacier" in such a way that it is a positive connotation!
- vidarh 14y agoVirgin (200+ businesses operating or having operated, under the Virgin brand, ranging from infrastructures - trains, airlines - to records, banking, bridal saloons under the name Virgin Bride...). Mitsubishi and Samsung springs to mind as two of the best known ones internationally where their brands are known in multiple markets internationally, though many of their businesses are less known outside Asia (e.g. Mitsubishi's bank is Japans largest). Any number of other Asian conglomerates. ITT used to fall in that category back in the day: Fridges, PC's, hotels, insurance, schools,telecoms and lots more. I remember we at one point had both an ITT PC and fridge. The name was well known in many of its markets. The large, sprawling, unfocused conglomerate have fallen a bit out of favor in Europe and the US. ITT was often criticized for their lack of focus even back in the 80's, and have since broken itself into more and more pieces and renamed and/or sold off many of them (e.g. the hotel group is now owned by Starwood).
- sharingancoder 14y agoLooks like the perfect solution to backup all my photos. Considering 100 GB of photos, that's still just 100 * 0.01 * 12 = $12 a year! I'm sold on this!
- jordanthoms 14y agoWow, the cadence of releases from aws recently has been amazing. Can't wait to see what else they have in store.
- buro9 14y agoThis is a really good offering for media that you typically will keep locally for instant access, yet you want to have an off-site backup in a way that lives for a very long time. Dropbox should work here, but it's simply too expensive. My photo library is 175GB. That isn't excessive when considering I store the digital negatives and this represents over a decade. I don't mind not being able to access it for a few hours, I'm thinking disaster recovery of highly sentimental digital memories here. If my flat burns down destroying my local copy, and my personal off-site backup (HDD at an old friends' house) is also destroyed... then nothing would be lost if Amazon have a copy. In fact I very much doubt anyone I know that isn't a tech actually keeps all of their data backed-up even in that manner. I find myself already wondering: My 12TB NAS, how much is used (4TB)... could I backup all of it remotely? It's crazy that this approaches being feasible. It's under 30GBP per month for Ireland storage for my all of the data on my NAS. To be able to say, "All of my photos are safe for years and it's easy to append to the backup.". That would be something. A service offering a simple consumer interface for this could really do well.
- tommoor 14y agoPerhaps DropBox et al will be able to make use of this Infrastructure to lower the costs of their service...
- lubos 14y agoor finally to break even?
- ch0wn 14y agoHow do you think they could use this service? I can't think of a use case for them.
- biot 14y agoThey store the last 30 days worth of versions of every file you modify. Dropbox could keep versions from the last 10 days in S3 but move the rest to Glacier. Restores for older versions wouldn't be instant, but if S3 storage is a nontrivial expense for them moving the bulk of previous versions to Glacier would cut down on costs.
- newhouseb 14y agoIt's a bit strange how S3 has a 'US Standard' region option, while Glacier has the usual set of regions (US East, US West, etc). I wonder if this means that unlike S3, Glacier isn't replicated across regions?
- miles932 14y agoUS Standard isn't replicated across regions.
- newhouseb 14y agoHm, you could be right, AWS says: > The US Standard Region automatically routes requests to facilities in Northern Virginia or the Pacific Northwest using network maps. I guess this means that the data is automatically geographically sharded?
- jordanthoms 14y agoKeep in mind that even if your data is in one AWS region, it'll still be stored in multiple different datacenters some distance apart. Just not on the other side of the US.
- mvanveen 14y agoI just started a project where I'm keeping a raspberry pi in my backpack and am archiving a constant stream of jpgs to the cloud. I've been looking at all the available cloud archival options over the last few days and have been horrified at the pricing models. This is a blessing!
- lucaspiller 14y agoGreat idea, have you got any more details?
- bambax 14y agoThis is the best thing ever! I've been dreaming about such a service for a very long time -- posted about how I needed this here 6 months ago: http://news.ycombinator.com/item?id=3560952 http://news.ycombinator.com/item?id=3560952 Rotating hard drives on the NAS in my attic is going to get a LOT simpler...
- nicolas314 14y agoOne recurrent issue with Amazon services is that they charge in US$ and currently do not accept euros. European banks charge an arm and a leg for micropayment conversions: last time I got a bill from AWS for 0.02$ it ended up costing me 20 euros or so. Pretty much kills the deal. One solution could be to pre-pay 100$ on an account and let them debit as needed.
- EwanToo 14y agoThat's really odd, I've never heard of fees that high - in the UK I can use my credit cards with Amazon and only pay a few pennies in transaction fees. Have a look at the CaxtonFX Dollar Traveller, you need to load it with $200 but then there's no transaction fees I think, except a slightly worse exchange rate..
- nicolas314 14y agoSure there are workarounds. But why should I have to jump through hoops when my account at Amazon is already entrusted with a European credit card and merrily charging euros for all other goods?
- cbg0 14y agoThis sounds like it's an issue related with your bank trying to scrounge more money off you. I have a European debit card and I get charged the equivalent in EUR, nothing more.
- dasil003 14y agoYou'll have to ask your bank since they're the ones screwing you. I've been using my US Wells Fargo account in Europe, and my British Lloyds TSB account in Europe and the US, and I've never gotten screwed anywhere near that bad. Maybe 1-5% premium I had to pay in some cases. But in any case, you should look into getting a USD account. I have a USD and EUR account from Lloyds TSB International and it is a good way to guarantee I never get screwed even a little bit when I'm traveling and doing a lot of small transactions.
- 14y ago
- Gussy 14y agoFrom the retrieval times they are giving, it seems plausible that they could be only booting the servers 5 or 6 times a day, to run the upload and retrieval jobs stored in a buffer system of sorts. Having the servers turned off for the majority of the time would same an immense amount of power, although I wonder about the wear on drives spinning up and down compared to being always on. Any other theories on how this works on the backend while still being profitable?
- Baughnie 14y agoTape drives. Lots and lots of tape drives.
- jhack 14y agoThis sounds really appealing as a NAS backup solution, but I'm a bit concerned about security and privacy. Let's say I want to backup and upload my CDs and movies, would Amazon be monitoring what I upload and assume I'm doing something illegal?
- catastrophe 14y agoAny easy to use Windows clients for Glacier (for backing up an OS and/or other files)? If not, to any developers reading this: there's money in them thar Glacier.
- josteink 14y agoThis looks like exactly what I need. I'm currently using S3 for backup and archiving by regularly running s3cmd to sync my data on my NAS. And while not super-duper expensive, s3 provides much more than I really need, and hence a more limited (but cheaper) service would definitely be appreciated. If there is anything with the easy of use like s3cmd to accompany this service I will be switching in a heartbeat.
- akurilin 14y agoAny desktop apps out there that will let you add folders for backup to Glacier and have them be automatically synced up to the cloud as they change? That would be quite useful.
- d0ugal 14y agohttp://www.haystacksoftware.com/arq/ http://www.haystacksoftware.com/arq/ I assume they will support Glacier soonish.
- gadders 14y agoAnyone know what the TOS are for this? I couldn't find them on a scan of the announcement. A lot of the consumer-level services refuse any liability for any data loss. Does Amazon do the same for this?
- ghshephard 14y agoPresumably the standard S3 SLA will apply: http://aws.amazon.com/s3-sla/ http://aws.amazon.com/s3-sla/ Realistically, you'd want to have at least two diverse cloud backup systems - I doubt you'd be happy with Service Credits if your data went missing.
- gadders 14y agoIt always seem funny though that these companies say "We will keep your data safe! * * T&C's apply, if we lose it, you're on your own." I shouldn't imagine Bank Vaults deal the same way with physical property. Can you get insurance for digital assets the same way you can for physical ones?
- omh 14y agoInterestingly they penalise you for short-term storage: Amazon Glacier is designed for use cases where data is retained for months, years, or decades. Deleting data from Amazon Glacier is free if the archive being deleted has been stored for three months or longer. If an archive is deleted within three months of being uploaded, you will be charged an early deletion fee. In the US East (Northern Virginia) Region, you would be charged a prorated early deletion fee of $0.03 per gigabyte deleted within three months
- MattSayar 14y agoAfter reading tezza's explanation [1] of how they're probably using tape storage, this makes sense; Amazon wants the mechanical robot arm to spend the majority of its time writing to the tapes. If you're constantly tying it up with writes/deletes, you're taking time away from its primary mission: to archive your data. Charging you for early deletes discourages that practice. [1] http://news.ycombinator.com/item?id=4411697 http://news.ycombinator.com/item?id=4411697
- ratzkewatzke 14y agoThey're not using tape storage. The ZDNet story confirms this: http://www.zdnet.com/amazon-launches-glacier-cloud-storage-hopes-enterprise-will-go-cold-on-tape-use-7000002926 http://www.zdnet.com/amazon-launches-glacier-cloud-storage-h.... There are any number of reasons why deletes would be discouraged. One is packing: if your objects are "tarred" together in a compiled object, discouraging early deletes makes it more cost-effective to optimistically pack early.
- brlewis 14y agoLikely they are losing money on the cost of initially writing data, then profiting from storage. With free early deletion they would lose money.
- jalada 14y agoCan anyone make sense of the retrieval fees? Seems like the most confusing thing ever. If I'm storing 4TB and one day I want to restore all 4TB, how much is it going to cost me?
- kondro 14y ago1 x retrieval request per archive (it's really designed to store a small number of large files/tars) plus $0.12/GB... therefore, 4TB = $480.
- jalada 14y agoNo, that's not right. That's only the data transfer rate. If you check the FAQ, they bill you based on your peak hourly retrieval rate. If I download 4TB at say...10MB/s, not only do I need to pay $480, but I also have to pay ~$257 as a retrieval fee? Their wording is confusing, but ignoring the free retrieval amount (negligible difference on a 4TB transfer): Fee = Peak hourly retrieval * number of hours in month * $0.01 https://aws.amazon.com/glacier/faqs/#How_much_data_can_I_retrieve_for_free https://aws.amazon.com/glacier/faqs/#How_much_data_can_I_ret... Admittedly ~$737 isn't the end of the world if your house has burned down and you need all your data back, but it's still important to know the details. I think in that situation, it would be cheaper to use their bulk import/export, which would be roughly $300 for 4TB
- kondro 14y agoYou're right it is complicated. But it also depends on how much you have archived because you get to access 5% of your data on a prorated basis too and that's based on your peak hourly rate. But ultimately, this product isn't designed for backup purposes. It's designed for archive purposes. If you have 4TB of customer data from 3+ years ago that you never access, but need to keep in case the IRS does an audit, then this is the place to put it.
- Dylan16807 14y agoThis is so confusing. So apparently if you spend the entire month retrieving the data at 1.6MB/s it only costs $40 plus transfer fees? And more importantly, how do you throttle your retrieval? Edit: So I'm working through a scenario in my head and trying to figure out how charging based on the peak hour isn't completely ridiculous. I have 8GB stored to try out the system. This costs a whopping dollar per year. One day I decide to test out the restore feature. So I go tell Amazon to get my files and wait a few hours. When Amazon is ready, I hit download. I'm on a relatively fast cable connection so the download finishes in an hour. I look at the data transfer prices and expect to be charged one dollar. But I didn't take into account this 'peak hour' method. I just used roughly 8GB/hour over the minimal free retrieval. This gets multiplied out times 24 hours and 30 days to cost 8 * 720 * $0.01 = $57. Fifty-seven times my annual budget because I downloaded my data too quickly after waiting hours for Amazon to get ready.
- Sami_Lehtinen 14y agoExcellent! Is there any good open source client / backup application for this? I would start using it immediately. I'm currently using ridiculously expensive backup solution.
- nodata 14y agoDuplciity supports S3, so I'd watch it for Glacier support: http://duplicity.nongnu.org/ http://duplicity.nongnu.org/
- wladimir 14y agoI wonder if Glacier support in Duplicity will be possible without large changes. AFAIK, duplicity also reads some state from the remote end to determine what to backup (Although it also keeps a local cache of this?). To use glacier, the protocol would have to completely write-only.
- nodata 14y agoTurns out there is less of a need for direct support: https://news.ycombinator.com/item?id=4411649 https://news.ycombinator.com/item?id=4411649
- takluyver 14y agoI'd guess it would use a hybrid approach, with recent backups on S3 (which duplicity already does) being shifted to glacier after a period of time. The FAQ indicates that Amazon plans to make this easy.
- gvalkov 14y agoHere's to hoping that duplicity and git-annex could somehow make use of this service. I'm far more optimistic about duplicity support though, as incremental archives seem to fit the glacier storage model much better. A git-annex special remote [1] might turn out to be much more challenging, if at all possible. [1] http://git-annex.branchable.com/special_remotes/ http://git-annex.branchable.com/special_remotes/
- squidsoup 14y agoOnce there's support for Glacier in boto (https://github.com/boto/boto https://github.com/boto/boto) I would imagine a duplicity backend would be easy to implement.
- kondro 14y agoI know a lot of people seem to have jumped on the backup options of Glacier here and, whilst there is some potential for home users to make use of this product for back, that is not what Glacier is intended for. Glacier is an archive product. It's for data you don't really see yourself ever needing to access in the general course of business ever again. If you're a company and you have lots of invoice/purchase transactional information that's 2+ years old that you never use for anything, but you still have to keep it for 5 - 10 years for compliance reasons, Glacier is the perfect product for you. Even its pricing is designed to take into account that the average use case is to only access small portions of the total archive store at a reasonable price (5% prorated for free in the pricing page).
- ghshephard 14y agoFor many users, though - they will never use the restore capability. And for those who do, with Backblaze, they'll usually get a FedEx of a Hard Drive - so Recovery time is measured in about a day or so. I wouldn't downplay the consumer backup/restore angle so quickly - for many (most?) consumers, the ability to restore rapidly is balanced by their desire to have low monthly payments. I think we're going to see a lot of consumer backup applications built on top of Glacier in the next several months that will be competing with (the already excellent) backblaze and friends. (Note - Backblaze has excellent real-time restore, with date versioning, for those of us who use it as an online data recovery tool as well)
- kloc 14y agoAmazon should complement this service with data contact centers which are connected to their data centers network. Then people could go to these centers in person and hand over their hard drives full of data for back up. It will be like bank lockers but only digital. At this low price people would want to upload terabytes of data which will be pain to upload/download.
- akh 14y agoWe've just added support for Amazon Glacier to http://www.PlanForCloud.com http://www.PlanForCloud.com so you quickly forecast your costs and compare it with other options.
- akh 14y agoWe just ran a quick cost forecast and it's interesting: If you start with 100GB then add 10GB/month, it would cost $102.60 after 3 years on AWS Glacier vs $1,282.50 on AWS S3!
- moontear 14y agoCrashplan is still cheaper for storage larger than 400GB. Crashplan+ Unlimited is USD 2.92/month if you take the 4 year package. When I upload 300GB to Amazon and pay 0.01 * 300 = USD 3/month. Amazon would be even more expensive for larger amounts of data. Is there some fine print I'm missing with Crashplan unlimited?
- josephagoss 14y agoAlso remember its free to recover your entire Crashplan archive (I have 500GB with them). If you wanted to recover 500GB with Glacier it would cost $200 @10MB/s (according to someones calculation further down) You have to pay for retrieval
- moontear 14y agoNot totally true, you have to pay for retrieval for 1GB and upwards per month: http://aws.amazon.com/glacier/#pricing http://aws.amazon.com/glacier/#pricing as well as 5cent per upload/retrieval request.
- ghshephard 14y agoWhenever I read about "Unlimited" plans (Backblaze has it as well, for $3.96 if you get a 1 year plan) I always think of things like Joyent and their "Lifetime" hosting, or AT&T and their "Unlimited" data plans. Their business plan is usually structured around people not actually using the service, and those who do use the "Unlimited" option usually end up either (A) being rate limited, or (B) having a conversation with the hosting/data provider to encourage them to transition elsewhere. What's exciting about this, is that Amazon doesn't care _how_ much data you send them - presumably they've priced this so it's profitable at any level you wish to use. It's a sustainable model. Services like TarSnap/Arq will likely adopt this new service (Possibly offering tiered backup/archival services?). I have (close to) zero doubt that Amazon's Glacier Archival Storage will be available 5 years from now at (probably less than) $0.01/Gigabyte/Month. They are a (reasonably) safe archival choice. Now that light users (<300 Gigabytes) have a financial incentive to move off of CrashPlan onto Amazon - it further exacerbates the challenges that "Unlimited" backup providers will face. All their least costly/most profitable may leave (or, at the very least, the new ones may chose Amazon first) With that said - I love Backblaze (Been a user since 2008) for working data backups, rapid-online (free) restores - and I will continue to use them, but I wouldn't plan on archiving a Terabyte of Data to them for the next 20 years.
- monkeypizza 14y agoThey should provide a "time capsule" option - pay X dollars, and after a set number of years, your data archive will be opened to the public for a given amount of time. There'd be no better way to ensure that information would eventually be made public.
- spindritf 14y agoYou can build it on top of Glacier.
- gregtour 14y agoWake up.
- Keyframe 14y agoI'm not sure if cost is right. Each project I work on is approx. 50-60 TB in size (video). Recent one got backed up on 20 LTO 5 tapes times three. That's $600 for tapes per project. Each tape set went to a separate location - two secure ones for about $20/year and one at studio archive for immediate access, if needed. I find this method extremely reliable and it cost ~$700 initially to back all up and virtually non existent further fees. With Glacier it would cost $600 per month.
- ghshephard 14y agoYou have a pretty niche (but interesting!) use case. An LTO-5 Tape can store 1.5 Terabytes Raw [1] - Call it 2 Terabytes with a bit of compression (your Video probably doesn't losslessly compress at 2:1). 60 Terabytes requires 30 Tapes - around $15/month to store at Iron Mountain. [2] . The Glacier Charge for 60 Terabytes is $600/month vs $15/month for Tape Storage. Also - upload/recovery times are problematic when you are talking 10s of terabytes. Right now, the equation is in favor of archiving tapes at that level (Even presuming you store multiple copies for redundancy/safety). Glacier is for the people wanting to archive in the sub-ten terabyte range - they can avoid the hassle/cost of purchasing tape drives, tapes, software - and just have their online archive. The needle will move - in 10 years Glacier might make sense for people wanting to store sub 100 Terabytes, and tapes will be for the multi-petabyte people. [1] http://en.wikipedia.org/wiki/Linear_Tape-Open http://en.wikipedia.org/wiki/Linear_Tape-Open [2] http://www.ironmountain.com/Solutions/~/media/9F17511FA1A74131A0ECF1C132C642E0.pdf http://www.ironmountain.com/Solutions/~/media/9F17511FA1A741...
- urza 14y agoI am a bit confused with the pricing of retrieval. Could some good soul tell me how much would cost to: Store 150 GB as one big file for 5 years. To this I will add 10 GB (also as one file) every year. And lets say I will need to retrieve the whole archive (original file + additions) at the end of year 2 and 5. How much will it cost?
- Yrlec 14y agoThis looks awesome! We are currently developing a P2P-based backup solution (http://degoo.com http://degoo.com) where we are using S3 as fall-back. This will allow us to be much cheaper and I am sure it will enable many other backup providers to lower their prices to.
- deleted 14y ago[deleted]
- xedarius 14y agoWould be nice if I could just type my Amazon login details into TimeMachine and magic off-site backups just happened.
- ck2 14y agoNow if Tim Kay just adds support for glacier and I can leave all the backup scripts as is...
- djbender 14y agoIn case anyone is wondering it appears you can only upload from their APIs right now. I wonder if they intend to make it accessible through their web interface at any point?
- rglover 14y agoI'm currently using an app called Arq that backs everything up to S3. If I had to guess, I'd say there's about 50-60 gigs or more on there. Last months bill was something like .60 cents. How does glacier compare or contrast to this setup (the app does something similar with the archive concept)?
- dont_believe_u 14y agoWould love to know where you're getting that $0.60 number (I assume you meant 60 cents and not 0.60 cents). Even with the first-year free tier, it costs $7.00/mo ($6.00/mo on RRS) to store 60 GB of data on S3.
- sreitshamer 14y agoI'd also like to know how the bill was so low (I'm the developer behind Arq). Is it perhaps because it's the first month's bill and you haven't had the 60GB on S3 for very long (not a full month)?
- tc 14y agoBeware that retrieval fee! The retrieval fee for 3TB could be as high as $22,082 based on my reading of their FAQ [1]. It's not clear to me how they calculate the hourly retrieval rate. Is it based on how fast you download the data once it's available, how much data you request divided by how long it takes them to retrieve it (3.5-4.5 hours), or the size of the archives you request for retrieval in a given hour? This last case seems most plausible to me [6] -- that the retrieval rate is based solely on the rate of your requests. In that case, the math would work as follows: After uploading 3TB (3 * 2^40 bytes) as a single archive, your retrieval allowance would be 153.6 GB/mo (3TB * 5%), or 5.12 GB/day (3TB * 5% / 30). Assuming this one retrieval was the only retrieval of the day, and as it's a single archive you can't break it into smaller pieces, your billable peak hourly retrieval would be 3072 GB - 5.12 GB = 3066.88 GB. Thus your retrieval fee would be 3066.88 * 720 * .01 = $22081.535 (719x your monthly storage fee). That would be a wake-up call for someone just doing some testing. -- [1] http://aws.amazon.com/glacier/faqs/#How_will_I_be_charged_when_retrieving_large_amounts_of_data_from_Amazon_Glacier http://aws.amazon.com/glacier/faqs/#How_will_I_be_charged_wh... [2] After paying that fee, you might be reminded of S4: http://www.supersimplestorageservice.com/ http://www.supersimplestorageservice.com/ [3] How do you think this interacts with AWS Export? It seems that AWS Export would maximize your financial pain by making retrieval requests at an extraordinarily fast rate. [(edit) 4] Once you make a retrieval request the data is only available for 24 hours. So even in the best case, that they charge you based on how long it takes you to download it (and you're careful to throttle accurately), the charge would be $920 ($0.2995/GB) -- that's the lower bound here. Which is better, of course, but I wouldn't rely on it until they clarify how they calculate. My calculations above represent an upper bound ("as high as"). Also note that they charge separately for bandwidth out of AWS ($368.52 in this case). [(edit) 5] Answering an objection below, I looked at the docs and it doesn't appear that you can make a ranged retrieval request. It appears you have to grab an entire archive at once. You can make a ranged GET request, but that only helps if they charge based on the download rate and not based on the request rate. [(edit) 6] I think charging this way is more plausible because they incur their cost during the retrieval regardless of whether or how fast you download the result during the 24 hour period it's available to you (retrieval is the dominant expense, not internal network bandwidth). As for the other alternative, charging based on how long it takes them to retrieve it would seem odd as you have no control over that.
- nl 14y agoInteresting to see that in some cases it probably makes sense to just stop paying the bills, rather than pay the early deletion fee[1]. [1]https://aws.amazon.com/glacier/faqs/#How_am_I_charged_for_deleting_data_that_is_less_than_3_months_old https://aws.amazon.com/glacier/faqs/#How_am_I_charged_for_de...
- mp99e99 14y agoI'm with Atlantic.net cloud [AWS competitor, full disclosure]; the price point for storage is great but retrieval seems expensive -- perhaps retrieval is very rare its offset by the savings on the storage. I know you can mail in drives for storage, can you have them mail you drives for retrieval? (for Glacier specifically) Also, prior comments made mention they were using some sort of robotic tape devices, but according to this blog: http://www.zdnet.com/amazon-launches-glacier-cloud-storage-hopes-enterprise-will-go-cold-on-tape-use-7000002926/ http://www.zdnet.com/amazon-launches-glacier-cloud-storage-h... Its using "commodity hardware components". So, thats why I thought maybe they are loss-leadering on the storage and making up on the retrieval prices? Its definately a interesting product and I love how there's a reason they called it Glacier. AMZN is a wild boar going after everyone!
- sintaks 14y agoI don't think they're loss-leadering on storage, but if they are, they don't think they will be for long. AWS (EC2 and S3 in particular) does very well when it comes to profit margins. I suspect they'd like to keep it that way, and that whatever they're charging gives them some slice of profit, however small.
- portentint 14y agoWorst. name. ever.
- milesokeefe 14y agoHow so? Doesn't it represent that your data is safe, as it is "frozen", and slow to retrieve like a glacier?
- talaketu 14y agoI like the name too. An accreting mountain of preserved data.
- res0nat0r 14y agoDoes anyone know if there is a CLI tool to interface with this yet? I see SDK's mentioned on the product homepage but I dont see any simple CLI tools for this yet to upload/download data and query etc. http://aws.amazon.com/developertools/ http://aws.amazon.com/developertools/
- mslot 14y agoI think this is more or less the formula for calculating monthly costs (corrections welcome): 0.01S+max(0,7.20*(R-0.0017S)/4) S is number of GB stored R is biggest retrieval in the month 4 is the average number of hours a retrieval For an example with 10TB storage (replace 10000 to change): http://fooplot.com/plot/4pu7u2gpox http://fooplot.com/plot/4pu7u2gpox x is biggest retrieval in GB, y is $/month
- mslot 14y agoCorrection: See my other post http://news.ycombinator.com/item?id=4416684 http://news.ycombinator.com/item?id=4416684
- JakeSc 14y ago> Amazon Glacier is designed to provide average annual durability of 99.999999999% for an archive. That's pretty impressive. I wonder how many bytes they lost to lose that .0000000001% of data.
- csears 14y agoI wonder if homeowner's insurance would cover the retrieval fee if your computer/hard drives were protected assets on the policy.
- jebblue 14y agoI just saw this and thought wow finally cheap mass storage. After reading the comments (and the Amazon Glacier web page) it's clear it's cheap archiving but not cheap retrieval.