| 00:48:04 | <justcool393> | Quick question: are there any published ratelimits for the Internet Archive? I'm trying to get a large amount of files from the Wayback Machine (potentially all captured at different points), but I don't want to get blocked or completely hammer their service |
| 01:30:55 | <DopefishJustin> | justcool393: I haven't seen any, and they have a number of posts straight up telling people to download it all https://blog.archive.org/2020/10/21/want-some-terabytes-from-the-internet-archive-to-play-with/ |
| 01:31:28 | <DopefishJustin> | I would advise using the "ia" python tool they provide rather than trying to scrape the pages |
| 01:33:24 | <DopefishJustin> | oh I see, you're talking about the wayback machine and not the file libraries |
| 01:34:22 | <DopefishJustin> | even then I think the philosphy is "just go for it and we'll worry about making it work" |
| 01:58:26 | | Stiletto quits [Ping timeout: 252 seconds] |
| 02:08:59 | | Stiletto joins |
| 04:08:22 | <justcool393> | ah ok |
| 04:08:23 | <justcool393> | thanks |
| 14:23:10 | | Arcorann quits [Read error: Connection reset by peer] |
| 16:48:10 | | @arkiver quits [Excess Flood] |
| 16:48:42 | | arkiver (arkiver) joins |
| 16:48:42 | | @ChanServ sets mode: +o arkiver |
| 16:51:28 | <@arkiver> | justcool393: how many requests would this be/ |
| 16:51:31 | <@arkiver> | ? |
| 18:08:44 | | HP_Archivist quits [Ping timeout: 240 seconds] |
| 18:15:38 | | OrIdow6 quits [Read error: Connection reset by peer] |
| 18:15:56 | | OrIdow6 (OrIdow6) joins |
| 18:55:57 | <justcool393> | arkiver: I was going to say in the ballpark of 24,000 |
| 18:56:05 | <justcool393> | but it's seems to be way less so far |
| 18:56:23 | <justcool393> | given i'm filtering by MIME type and such |
| 18:56:45 | <justcool393> | right now any time my script makes a request it'll wait 5 seconds |
| 19:29:52 | <@arkiver> | justcool393: ok, that number is not a problem |
| 22:35:19 | | Arcorann (Arcorann) joins |