this post was submitted on 21 Aug 2023
358 points (98.4% liked)

Piracy: ꜱᴀɪʟ ᴛʜᴇ ʜɪɢʜ ꜱᴇᴀꜱ

54462 readers
283 users here now

⚓ Dedicated to the discussion of digital piracy, including ethical problems and legal advancements.

Rules • Full Version

1. Posts must be related to the discussion of digital piracy

2. Don't request invites, trade, sell, or self-promote

3. Don't request or link to specific pirated titles, including DMs

4. Don't submit low-quality posts, be entitled, or harass others



Loot, Pillage, & Plunder

📜 c/Piracy Wiki (Community Edition):


💰 Please help cover server costs.

Ko-Fi Liberapay
Ko-fi Liberapay

founded 1 year ago
MODERATORS
 

See linked posting. I've commented there with a link to a CLI tool in Python that allows downloading of IA collections. I've submitted a patch to enable specifying start and end points so that it's easier to resume downloading a huge collection, or to allow multiple people to split up the work.

https://archive.org/details/georgeblood

https://archive.org/details/78rpm_bowling_green

F*ck the RIAA and absurdly long copyright.


EDIT: There is more than one collection of 78s on IA, so I updated the title.


The issue with these collections are that they're absolutely HUGE. And yes, IA offers torrents for them, but as a separate torrent for every. single. album. And the torrents have all data in them -- FLAC, fixed-rate MP3, VBR MP3, PDF liner notes, etc. etc... there may be some extremely hardcore data-hoarders out there who want everything, but IMHO as these are scratchy old 78 records, FLAC is overkill to just save the audio in a listenable format. The George Blood collection, just the VBR MP3s, is looking to be about 6TB. With ALL data it might be over 40TB! I can't afford that many hard drives :)


So, my approach at the moment is to save just the VBR MP3s (they seem to be done at up to 320kbps VBR) and the JPEG album cover. If I have a chance and any storage left afterwards, I can make a separate pass to get the album liner PDFs...


Tool used: https://github.com/jjjake/internetarchive


Patch to allow setting start and end item indices for downloads: https://github.com/jjjake/internetarchive/pull/605


Example usage to grab just the VBR MP3 and record label JPG for each (note the --start-idx and --end-idx arguments):

#ia download --start-idx=4001 --end-idx=8000 -a -i --format="VBR MP3" --format="JPEG" --search collection:georgeblood

I'm going to concentrate on the George Blood collection for now.. I'm starting at item 1. It would be great if others started at index 50,000, 100,000, 150,000, ... and others started at the end and worked backwards in similarly-sized chunks, so that it's assured someone gets each of them.

all 35 comments
sorted by: hot top controversial new old
[–] Haui@discuss.tchncs.de 33 points 1 year ago (3 children)

Probably stating the obvious but „are in no threat of being deleted“ is an absolute joke.

A company holding the IP can just make it unavailable tormorrow. A big chunk of us is here because reddit somehow is allowed to delete our posts because the law is idiotic. At least european people are allowed to get their data but the cooperative works of thousands of people is threatened due to those laws.

The concept of IP needs to be reformed.

[–] sxan@midwest.social 15 points 1 year ago (2 children)

As concrete examples, try to get a copy of Disney's 1946 movie, "Song of the South." It's been removed from circulation because of its whitewashed presentation of "happy slaves." Similarly, 6 of Dr. Seuss' books, including "And to Think That I Saw It on Mulberry Street" were withdrawn because of racial imagery (the mentioned book had a "Chinaman" drawn with a WWII stereotype style - rice hat, sloping eyes, buck teeth).

There's media you simply can't get anymore.

[–] Haui@discuss.tchncs.de 8 points 1 year ago (1 children)

Our culture has been copyrighted.

[–] sxan@midwest.social 2 points 1 year ago (1 children)

In this case, the media was withdrawn for (arguably) good reasons: the representations were deemed hurtful or harmful.

Good reasons or bad, they still stand as stark examples of how media can disappear at the whims of a single organization.

[–] Haui@discuss.tchncs.de 3 points 1 year ago

Yes and it’s horrific

[–] wizardbeard@lemmy.dbzer0.com 2 points 1 year ago (1 children)

Fun fact, there is a fan made blu-ray quality remaster of Song of the South available on IA.

[–] sxan@midwest.social 2 points 1 year ago (1 children)

Say what? Now I'm curious how they handled the slavery topic, and found actors for it.

Thanks for the heads-up!

[–] GnuLinuxDude@lemmy.ml 2 points 1 year ago* (last edited 1 year ago)

Song of the South does whitewash being black in the USA, but it is set in post-civil war America, so superficially it does not need to handle the slavery topic, which can be dismissed as having been dealt with already.

[–] Arghblarg@lemmy.ca 14 points 1 year ago (1 children)

Yeah. And whenever anyone says "Oh the music companies would never let these old recordings die, it's their bread and butter!" I give them this story.

We cannot trust our cultural heritage to any one entity.

[–] Haui@discuss.tchncs.de 3 points 1 year ago

Oof. I just read this. It’s pretty brutal.

[–] maudefi@lemm.ee 22 points 1 year ago (1 children)

Cool tool! Please consider leaving GitHub for any of the numerous FOSS options.

[–] Arghblarg@lemmy.ca 8 points 1 year ago (1 children)

Oh, it's not my project -- I already have moved my own projects off there, yeah.

[–] maudefi@lemm.ee 7 points 1 year ago (1 children)

That's awesome! Really encouraging seeing projects and devs migrate away from closed-source and proprietary systems and features. 💪

[–] Arghblarg@lemmy.ca 4 points 1 year ago* (last edited 1 year ago)

sourcehut, self-hosted Gogs or Forgejo are some good candidates. Gitea is popular, but there's apparently been some drama about them going commercial without proper buy-in from their contributors. (The code lineage is AFAIK Gogs → Gitea → Forgejo).


All the above solutions make it super-easy to mirror a github project as well, just in case it goes away :) Doing so has saved my arse more than a few times when github takes a repo down for stupid reasons.


Mandatory plug for !selfhosted@lemmy.world :)


Gitlab seems too heavyweight to me. I use Gogs myself on my home server. No code review tools via PR ala github/gitlab, but I don't need those in my web frontend.

[–] the_lone_wolf@lemmy.ml 9 points 1 year ago

pirates sail your ship

[–] Cyb3rManiak@kbin.social 7 points 1 year ago

Instructions unclear. Linked posting explains nothing. Will assume this is about 78 missing dragonballs and move on.

Jokes aside, we must preserve the 78 collection. What if in the future an alien signal will reach earth and no one can understand it because all 78s are extinct? We don't have Starfleet to go back in time and get a 78 from San Francisco in the past to save the future!