03:41:09qw3rty__ joins
03:45:02qw3rty_ quits [Ping timeout: 258 seconds]
04:30:57Stiletto joins
04:32:57Stilett0 quits [Ping timeout: 258 seconds]
06:36:15Ryz quits [Remote host closed the connection]
06:39:37Ryz (Ryz) joins
08:18:09rsn quits [Quit: Leaving]
09:27:52Stiletto quits [Ping timeout: 250 seconds]
09:30:03Stiletto joins
10:16:03<HP_Archivist>Hey JAA remember a while back we talked about SingleFileZ extension and how saving HTML files to Archive.org would load the live web pages right on IA? You had mentioned that Chrome doesn't support local viewing of saved HTML pages from SingleFileZ.
10:16:26<HP_Archivist>So routine checking on my uploads from time to time, I discovered that it appears that function of loading live web pages is broken.
10:16:28<HP_Archivist>https://archive.org/details/harry-potter-games-auxiliary
10:16:52<HP_Archivist>Click HTML on the side, the first HTML in the list - https://ia803204.us.archive.org/31/items/harry-potter-games-auxiliary/Adam%20on%20Twitter_bgolus%20Wow%2C%20that%27s%20crazy%20the%20whole%20game%20was%20made%20in%206%20months.html
10:17:02<HP_Archivist>Something broke on Archive-side
10:17:17<HP_Archivist>But when saved locally and opened in FF it loads just fine (still won't load in Chrome.)
10:18:12<HP_Archivist>Every single one I tried to open *on* IA loads a jumbled mess like that one above... So something must've broken over the past 6 months...
10:21:26<HP_Archivist>Part of me wishes I had saved every single one as a PDF/A and uploaded that way. But now with the new issue of PDFs, too, seems like nothing is perfect on Archive.org heh
10:22:01<OrIdow6>Type is sent as text/plain rather than text/html
10:27:40<OrIdow6>The data is still there, it's just not displaying right
10:28:04<OrIdow6>You would lose much more saving as PDF, than by requiring the user to download it and reopen
10:29:20<HP_Archivist>OrIdow6 yeah I know it's still there otherwise it wouldn't have loaded locally
10:29:31<OrIdow6>I'm guessing that whatever is serving it uses the non-ASCII to infer that it's not HTML
10:29:33<HP_Archivist>The whole point is to have it accessible ON Archive as much as possible
10:31:23<OrIdow6>I'd say you should just user singlefile instead of singlefilez
10:31:42<HP_Archivist>Yeah maybe it'll be resolved in time, but a bummer because it's alright full of files for people to go through. Kind of embarrassing to offer it all up and then have caveats like 'Archive.org won't load it, you're gonna have to download and view locally'
10:31:44<OrIdow6>I don't blame them for assuming that HTML files won't have huge chunks of binary data in the middle
10:32:30<HP_Archivist>Fair point and yeah noted for SingleFile
10:34:10<HP_Archivist>Oh and yeah OrIdow6, I agree about losing far more if saved to PDF. Though I would argue the avg user doesn't care about the elements so much as long as most of the text/visuals are still there akin to the live version unless it's heavy on multimedia or other script-type stuff
10:35:41<HP_Archivist>Saving every aspect or element of a web page is the goal, seems like there's never a clear way of archiving pages though heh
10:41:57<HP_Archivist>OrIdow6 worth trying to contact staff or is this more complicated than an email request
10:42:57<OrIdow6>I'd guess that reversing this would require a lot of technical changes
10:43:35<OrIdow6>SingleFileZ is a hack that depends on browsers being extremely lenient
10:43:42<OrIdow6>In how they parse HTML
10:44:18<OrIdow6>And the fact that they're doing it like this doesn't mean they're wrong
10:45:54<HP_Archivist>Well, it's not like any of this pages saved aren't in the WBM already. They're all there. My goal, here, was to try and make these sets of pages more accessible because the average person, non-technical, wouldn't even know to go to the WBM machine let alone how to use it to find an old web page
10:46:10<HP_Archivist>So this was my attempt at making these older pages more easily accessible
10:46:41<HP_Archivist>Perhaps I created more problems than I solved heh
10:47:18<OrIdow6>The Wayback Machine isn't that badly-known
10:47:36<OrIdow6>I've seen it come up in disucssions unrelated to "archiving"
10:47:58<OrIdow6>But dynamic pages obviously cause a problems, and for now singlefile and the like are good stopgaps
10:48:30<HP_Archivist>I suppose. But when I think about the avg person, think about polling 10 random people in a major city. How many of them will know what it is, and how many of them will have used it, how many of the will know how to search in a savvy way
10:48:33<OrIdow6>(I think the solution will be to do something like what Purplebot does, and serialize the rendered page back into HTML and put it in a warc record, but that's for later)
10:49:45<HP_Archivist>Perhaps it's worth, in time, re-saving these pages with SingleFile. Idk. There are over 200 pages there, took several weeks to compile
10:51:09<HP_Archivist>Also, I'm not familiar with Purplebot
10:51:37<OrIdow6>That's not something that's relevant to your case anyhow
10:51:46<OrIdow6>(And I meant Chromebot, not Purplebot, obviously)
10:52:18<HP_Archivist>I was under the assumption that was another bot for different html-type jobs heh no worries
12:32:16AK quits [Quit: AK]
12:33:22AK (AK) joins
13:46:38OrIdow6 quits [Quit: Quitting.]
13:47:58OrIdow6 (OrIdow6) joins
13:48:54AK quits [Client Quit]
13:51:55AK (AK) joins
13:53:10OrIdow6 quits [Remote host closed the connection]
13:54:39OrIdow6 (OrIdow6) joins
13:58:43OrIdow6 quits [Remote host closed the connection]
13:59:24OrIdow6 (OrIdow6) joins
14:03:19OrIdow6 quits [Remote host closed the connection]
14:03:48OrIdow6 (OrIdow6) joins
14:07:33OrIdow6 quits [Remote host closed the connection]
14:08:15OrIdow6 (OrIdow6) joins
15:00:05offog quits [Ping timeout: 258 seconds]
15:29:20offog joins
15:41:27OrIdow6 quits [Changing host]
15:41:27OrIdow6 (OrIdow6) joins
16:30:16Fusl quits [Excess Flood]
16:30:38Fusl (Fusl) joins
16:37:08@ChanServ sets mode: +o OrIdow6
16:56:29@ChanServ quits [shutting down]
16:56:31ChanServ joins
16:56:31ing.hackint.org sets mode: +o ChanServ
18:35:42<@arkiver>OrIdow6: JAA: on why those records are not in wayback
18:36:16<@arkiver>we have more indexes than i was aware of, that item missed the large 'all' index
18:36:40<@arkiver>it should be added when the next 'all' index for wayback is created, which is likely up to a month from now
18:36:52<@arkiver>i'm keeping an eye on this
18:37:12<@JAA>Two hard things in computer science etc. etc.
18:39:08<@arkiver>JAA: what do you mean?
18:39:30<@JAA>arkiver: There are two hard things in computer science: caching, naming things, and off-by-one errors.
18:39:40<@JAA>This is basically a caching bug.
18:39:43<@arkiver>right yeah
18:40:02<@arkiver>well, the reason is missed the 'all' index is because the data was being shuffled to another drive when the index was being built
18:41:02<@JAA>So there is an index that gets rebuilt frequently (daily?) based on the item CDXs, and then there's that 'all' index that gets rebuilt less frequently for the long term I guess?
18:41:20<@arkiver>exactly yes
18:41:24<@JAA>Plus a separate one for liveweb/SPN so that gets reflected immediately.
18:41:33<@arkiver>yes
18:41:43<@JAA>Good to know, thanks.
18:42:02<@arkiver>eventually records are removed from the index, under the assumption they are then in the 'all' index
18:42:13<@JAA>Makes sense.
21:37:36Wayward quits [Ping timeout: 258 seconds]