00:07:20<@arkiver>yeah they are adding resources every now and then
00:07:24<@arkiver>but demand is high
00:22:01<Ryz>Is there extra reserves for people who have Internet Archive accounts?
00:30:07Wayward (wayward) joins
00:43:53<@arkiver>as far as I know now
00:43:54<@arkiver>no
00:43:55<@arkiver>*
00:57:20DogsRNice (Webuser299) joins
00:57:26AK quits [Client Quit]
00:57:35nothere quits [Client Quit]
00:57:46AK (AK) joins
00:57:47nothere joins
01:04:52<Ryz>...I wonder if there's a way to see what's directly being archived on WBM via SPN
01:05:12<Ryz>Akin to https://archive.is/rss
01:07:18<Ryz>Hmm, I'm trying to find the RSS feed of Archive.is - it's been really some time since seeing that...
01:21:41<OrIdow6>I don't think that exists for SPN (though IIRC IA does have internal information about it)
01:21:56<OrIdow6>Could be a sort of privacy issue as well
01:22:16<OrIdow6>Though I suppose it's technically possible to figure it out after it happens (except for blacklisted sites)
01:58:18<@JAA>Oh yeah, another SPN issue that's currently very pronounced is that it fails to replay the snapshot shortly afterwards, claiming that it hasn't been archived. I've also seen infinite loops on images for the same reason.
01:59:04<@JAA>I imagine that makes things worse for the system load as well as many casual users won't understand that the snapshots will be visible later, so they try to save it again, etc.
04:06:56balrog quits [Ping timeout: 250 seconds]
04:25:48balrog joins
04:31:33qw3rty__ joins
04:35:31qw3rty_ quits [Ping timeout: 264 seconds]
04:47:39HP_Archivist quits [Quit: Leaving]
05:44:35HP_Archivist (HP_Archivist) joins
06:22:12DogsRNice quits [Read error: Connection reset by peer]
07:05:34Ajay quits [Ping timeout: 244 seconds]
07:19:56Ajay joins
08:08:31Wayward quits [Ping timeout: 264 seconds]
08:12:49flashfire42 quits [Quit: Ping timeout (120 seconds)]
08:13:18flashfire42 (flashfire42) joins
08:27:49Wayward (wayward) joins
09:00:19britmob2 quits [Quit: britmob2]
13:46:17britmob2 joins
15:18:40<HP_Archivist>Can IA not accept PDF/A-1a standard documents?
15:19:00<HP_Archivist>I received a 400 bad data error when trying to upload with the following:
15:19:12<HP_Archivist><?xml version='1.0' encoding='UTF-8'?>
15:19:12<HP_Archivist><Error><Code>BadContent</Code><Message>Uploaded content is unacceptable.</Message><Resource>Syntax error detected in pdf data. You may be able to repair the pdf file with a repair tool, pdftk is one such tool.
15:19:12<HP_Archivist></Resource><RequestId>355202e6-c5db-423a-8bd9-e864c49846f3</RequestId></Error>
15:26:09<HP_Archivist>Weird. Uploaded a regular PDF to the same item and it came back with the same error message. Only thing about this PDF is that it contains redactions
15:26:11<HP_Archivist>Any ideas?
15:27:32<@JAA>https://github.com/jjjake/internetarchive/issues/388
15:28:21<HP_Archivist>Ugh
15:28:22<HP_Archivist>Alright
15:28:25<HP_Archivist>Thanks JAA
15:29:26<@JAA>Either the PDF isn't actually standard-conformant, or there's a bug in whatever PDF validation stuff IA uses. In the latter case, good luck...
15:30:59<HP_Archivist>I've uploaded a redacted PDF before without issue. But it's been a while. I uploaded a regular PDF just a few days ago to a separate item without issue. But this new PDF contains redactions. Guessing it's that maybe
15:33:13<@JAA>Nah, if the file is valid, it should work fine.
15:33:25<HP_Archivist>Opens locally without any problems.
15:33:33<HP_Archivist>Alternatively, I could try uploading in a zip?
15:33:35<@JAA>But if the software that was used to redact the document produces invalid PDF files...
15:33:59<HP_Archivist>I'm using Nuance PowerPDF and the redactions are someone's personal email
15:34:04<@JAA>'Works fine with $PDFVIEWER' just means that your viewer might work around any inconsistencies.
15:34:30<@JAA>Just like browsers will happily render invalid HTML and hence fuck up the language, the same is true for PDF or really any file format, sadly.
15:34:43<HP_Archivist>Hm
15:34:54<@JAA>Instead of forcing the software devs or people who produce those invalid files to fix their crap.
15:35:15<HP_Archivist>Yeah but strikes me as odd is I've done this before and they've uploaded to IA just fine
15:35:25<HP_Archivist>Using the same version of the software to create the PDF
15:35:32<HP_Archivist>So I think it's an issue with IA
15:35:34<@JAA>¯\_(ツ)_/¯
15:35:51<HP_Archivist>I'm going to upload in a zip, see if that works
15:35:52<@JAA>Or IA updated their PDF checker to verify PDFs more rigidly.
15:36:03<HP_Archivist>^ fair point, hadn't thought of that
15:36:14<@JAA>In any case, try validating the PDF with a software that checks it for standard compliance.
15:37:38<HP_Archivist>Will the free version of pdftk suffice?
15:37:48<HP_Archivist>Never used it
15:38:25<@JAA>Neither have I.
15:39:16<HP_Archivist>Idk if makes a difference, but a previous upload of a redacted PDF using the same version of the same software still opens just fine but no idea if that's to be expected if IA updated their checker
15:39:31<HP_Archivist>Still opens in the browser from IA *
15:39:43<HP_Archivist>I'm just going to upload in a zip
15:41:44<@JAA>That would be expected since the files are only checked on upload.
15:41:56<@JAA>Uploading in a ZIP makes it less accessible.
15:42:30<HP_Archivist>I understand that. However, the user can still 'view contents' and it appears to load just fine in the browser
15:42:31<HP_Archivist>https://ia601507.us.archive.org/view_archive.php?archive=/1/items/harry-potter-games-talking-with-senior-producer-stuart-whyte/Harry%20Potter%20Archive%20Project%20Speaks%20with%20Senior%20Game%20Producer%20Stuart%20Whyte.zip&file=Harry%20Potter%20Archive%20Project%20Speaks%20with%20Senior%20Game%20Producer%20Stuart%20Whyte.pdf
15:45:00<HP_Archivist>*shrugs* Alternatively, I could also upload the screenshots of this email without that person's email address shown.
15:45:24<HP_Archivist>6 or one half a dozen or another heh
15:46:00<@JAA>Screenshots of text are always great for searchability.
15:46:37<HP_Archivist>Yeah I think I might do that. Also, I can revisit this at a later time. Try and upload just the PDF and maybe the issue would be fixed
15:48:02<HP_Archivist>Thank you for help JAA
15:48:06<HP_Archivist>for the help*
15:49:16<@JAA>Always happy to give advice that gets ignored. :-)
15:50:49<HP_Archivist>Heh :)
16:04:42Larsenv_ is now known as Larsenv
16:05:42Doranwen quits [Remote host closed the connection]
16:11:13Doranwen (Doranwen) joins
16:31:56DogsRNice (Webuser299) joins
21:18:42qw3rty__ quits [Ping timeout: 250 seconds]
21:21:55<mgrandi>Why does IA have a pdf validator? Does it do that for other formats?
21:22:48<mgrandi>And also, since the WBM/SPN is having such trouble lately, do stuff like archivebot retry if the job fails?
21:23:26<@arkiver>mgrandi: afaik AB does a few tries, JAA ^?
21:23:32<@arkiver>mgrandi: also yes they check other formats
21:25:30<@JAA>mgrandi: AB makes three attempts, but that's entirely unrelated to the SPN issues, which have (usually) nothing to do with the origin server and everything to do with IA's servers.
21:27:29<mgrandi>Yeah that's what I meant, would it retry if it gets a "job failed" error or similar
21:28:38<@JAA>It won't because it doesn't use SPN in the first place.
21:28:46<@JAA>Never did, never will.
21:30:05<mgrandi>Ah, today I learned, I thought it used the super not so secret SPN api
21:31:21<@JAA>lol no, we wouldn't even manage 1 ‰ of what we're doing now.
21:47:23qw3rty joins
21:48:12AK quits [Client Quit]
21:49:20AK (AK) joins
23:11:09atphoenix quits [Remote host closed the connection]
23:11:47atphoenix (atphoenix) joins
23:24:15Larsenv quits [Client Quit]
23:25:02Larsenv (Larsenv) joins