Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think none of these issues really matter.

Sure, it isn't perfect. But all of those issues are liveable-with. Most of them involve handling corrupt data - but any serious archive will hash the whole file as part of the cataloging process.

The commonly-used technology is by far the best choice for long term archiving, because something that has a billion users will go obsolete/unreadable a long time after a fancy compressor written by a phd student in your lab.

If I were running an archive today, I would be keeping everything in .zip files, because there is a ~35 year window of computers that can open them, and I wouldn't be surprised if the format remains in common use for a further 35 years. That means that in the year 3000, someone wanting to access the data only needs to find/emulate a system from anytime between 1990 and 2060 to have a good chance of reading the data.



> The commonly-used technology is by far the best choice for long term archiving, because something that has a billion users will go obsolete/unreadable a long time after a fancy compressor written by a phd student in your lab.

Did you read the article? `xz` is far more complex than the alternatives. Your analogy doesn't make sense in this scenario.

> If I were running an archive today, I would be keeping everything in .zip files

So...not xz. Okay.


complexity doesn't matter really... You aren't going to be reading these with a hex editor. What matters is that decompression software is widely available today, and will be long into the future. While thats kinda true for .xz, it's far far more true for .zip.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: