00:16:31godane1 joins
00:19:00godane quits [Ping timeout: 252 seconds]
00:51:29VADemon quits [Client Quit]
00:57:39<Ryz>Oh, so it isn't just me, YouTube's broken: https://twitter.com/disclosetv/status/1326682623595442176
00:57:52<Ryz>I was about to play a song for myself to listen before I have to do other things
00:58:22<Ryz>Videos won't play~
00:59:10<Ryz>Alternatively, videos take a real long while to load
01:00:34<Craigle>Definitely seeing the same thing here. I had children running in panic to inform me YouTube stopped working, lol
01:15:56<EggplantN><N​orthern#2000> It's 1am tf am I gonna do till 3am. God damnit Google
01:16:31<@JAA>Set up a system that automatically downloads videos you're interested to local storage. :-P
01:16:54<EggplantN><N​orthern#2000> That is on a to-do list somewhere with one trillion other things.
01:17:23<@JAA>I hope fixing your client is #1 on that list.
01:17:49<EggplantN><N​orthern#2000> It's around about 5th - 8th depending how soon this RHEL5 box falls over
01:34:57<@arkiver>i hope so too
01:46:09vitzli (vitzli) joins
01:47:25Stiletto joins
01:49:35Stilett0 quits [Ping timeout: 252 seconds]
01:50:33godane2 joins
01:52:43godane1 quits [Read error: Connection reset by peer]
02:06:34vitzli quits [Client Quit]
03:04:25on joins
03:04:36<on>yo, I return
03:04:47on is now known as FalconK
03:04:54<FalconK>and apparently I forgot how to use IRC while I was gone
03:17:01<@JAA>Welcome back!
03:17:17<FalconK>thanks!
03:18:19<FalconK>technically I should be working on my thesis, but instead I'm working on trying to beat Squidwarc into being a more maintainable crawler
03:18:23<FalconK>has anyone tried that before?
03:59:00HP_Archiv joins
04:05:09HP_Archivist quits [Client Quit]
04:05:21HP_Archiv quits [Client Quit]
04:05:41HP_Archivist (HP_Archivist) joins
04:21:46<FalconK>experimentally, squidwarc is really, really, really slow, and I'm not sure if it can easily be made parallel either.
04:22:34<FalconK>but it looks like when it loads the page, its strategy is to wait $x=1500 milliseconds after the network goes idle, and then go to the next page
04:23:07<@JAA>Any browser-based archival is going to be really, really, really slow, unfortunately.
04:23:12<@JAA>Cf. brozzler and crocoite/chromebot.
04:23:18FalconK nods
04:23:46<@JAA>I try to reverse-engineer how stuff works either through dev tools or reading the JS code, then emulate it with qwarc.
04:24:17<@JAA>But that's not really 'high-fidelity', and getting it to cover all edge cases is really tricky.
04:24:24<FalconK>I suspect the approach I'm taking could be made fast by making it parallel
04:24:41<@JAA>You might want to order a shipping container full of RAM though.
04:24:45<FalconK>ha
04:24:56<FalconK>I have a box with 96gb of ram in a rack in fremont now that I'm using
04:25:17<@JAA>Sounds like you could run about 2 Chromium tabs then. :-)
04:25:23<FalconK>I could also start running a normal pipeline on that box
04:25:45<FalconK>though I seem to have lost my ability to write to the archivebot collection in the archive
04:26:12<HP_Archivist>Continued from #youtubearchive I do hope we find a way to get some of what we've archived on IA. Is it worth asking Jason Scott anyway?
04:26:52<@JAA>FalconK: The direct uploader was always having issues anyway, e.g. hitting the item size limit and truncating filenames.
04:27:38<FalconK>huh. I'd be interested in fixing that.
04:27:55<@Fusl>> a box with 96gb of ram
04:27:56<@JAA>Don't bother. The uploading system will be revamped entirely soon anyway.
04:28:00<FalconK>though the filename truncation was on purpose, because there's a maximum filename length limit
04:28:01<@Fusl>our tracker redis uses more than that by now :D
04:28:26<FalconK>yeah the redis is massive
04:30:02<HP_Archivist>JAA: Alright. I only say that because I've spent the last year dedicating time, we all have, and would hate to see that all for naught
04:30:31<FalconK>I suspect he was telling me not to bother, not you lol
04:31:01<@JAA>Yep
04:31:12<HP_Archivist>Oops
04:31:15<ivan>I have other contacts at IA who are responsive
04:31:56<ivan>I would appreciate help writing Rust and Svelte for the archive browser if anyone wants to later
04:32:29<ivan>or porting some of the uploader from Elixir
04:32:54<HP_Archivist>ivan: I don't know either language but can ping IA people if you need help with that
04:33:05<FalconK>blah, I don't know either of those either
04:33:05<ivan>IA people don't do your work for you
04:34:27<HP_Archivist>Well by contacting someone else other than JS, I was hoping someone would agree to take in most of what we've amassed from YT
04:34:39<ivan>there is no bottleneck on that
04:35:04<ivan>IA takes deleted stuff
04:35:51<HP_Archivist>Yeah but didn't JAA say they would only take maybe 1 PB
04:36:09<ivan>don't think they'll take all of it
04:36:41<ivan>when I say the bottleneck is writing software, it's literally only that
04:37:01<ivan>if someone would like to devote more attention to the channel and dashboard that would give me an extra hour or so a day
04:37:09<@JAA>Yeah, IA has repeatedly said they don't want people to mirror YT onto their systems, so it's unlikely they'll say 'yeah, of course' to 8 PB from ivan.
04:38:05<@JAA>There was a big thing a while ago about 'stolen revenue' and stuff because people were watching videos on IA instead of YT and getting around the ads that way.
04:38:18<@JAA>That's what caused the Mirrortube collection to get noindex'd, IIRC.
04:38:46godane joins
04:40:26godane2 quits [Ping timeout: 252 seconds]
04:40:35<HP_Archivist>Right. When I talked to JS about archived video for my Harry Potter project he said 'you're basically mirroring people's YouTube's. If you ship us a drive of everything we'll take it but will all be indexed'
04:40:49<HP_Archivist>non-indexed*
04:41:05godane1 joins
04:41:12<HP_Archivist>And that's why I did it all manually with ytdl.
04:42:11<HP_Archivist>ivan: I could set time aside to learn either language but I'm afraid I wouldn't be much help. There are number of members in here, nobody else knows or can help
04:43:08<@JAA>Meanwhile, I'm trying to figure out what the hell I did back in March when I archived that Bulgarian educational video site. lol
04:43:22godane quits [Ping timeout: 252 seconds]
04:44:01<FalconK>blah, squidwarc also generates non-gzipped WARCs
04:44:04<HP_Archivist>On avg I'd say there are between 10-15 active people here daily. Compared to the list of names on the right I would think more people could help with this. *shrugs*
04:46:51<FalconK>I think maybe instead of this I'll try and set up another pipeline and see how that goes / how much maintenance it demands.
04:48:55<@JAA>Let's discuss that in #archivebot.
04:55:46qw3rty_ joins
04:57:49Arcorann (Arcorann) joins
04:59:04qw3rty__ quits [Ping timeout: 240 seconds]
05:04:07teej quits [Client Quit]
05:14:48cadence quits [Client Quit]
05:36:45Ryz quits [Remote host closed the connection]
05:37:40Ryz (Ryz) joins
06:02:45godane1 quits [Read error: Connection reset by peer]
06:03:26godane1 joins
06:04:59godane2 joins
06:07:41systwi_ is now known as systwi
06:07:42godane1 quits [Ping timeout: 252 seconds]
07:10:15godane2 quits [Read error: Connection reset by peer]
07:11:30godane2 joins
07:21:28godane1 joins
07:21:28godane2 quits [Read error: Connection reset by peer]
07:25:13godane1 quits [Read error: Connection reset by peer]
07:26:28godane1 joins
07:26:51@arkiver quits [Excess Flood]
07:27:42arkiver (arkiver) joins
07:27:42@ChanServ sets mode: +o arkiver
07:58:24marked195 quits [Ping timeout: 240 seconds]
08:01:44marked195 joins
11:45:24nico_32 quits [Ping timeout: 240 seconds]
11:47:32nico_32 joins
12:03:00godane2 joins
12:05:24godane1 quits [Ping timeout: 240 seconds]
12:07:28godane1 joins
12:09:44godane2 quits [Ping timeout: 240 seconds]
12:09:54VADemon joins
14:08:33Datechnom quits [Read error: Connection reset by peer]
14:09:22Datechnom (Datechnom) joins
14:46:22<benjins>https://workspaceupdates.googleblog.com/2020/11/changes-to-google-workspace-storage.html "more than 4.3 million GB are added across Gmail, Drive, and Photos every day"
14:56:26Arcorann quits [Ping timeout: 252 seconds]
15:00:34Atom-- joins
15:02:44Atom quits [Ping timeout: 240 seconds]
15:04:53@arkiver quits [Excess Flood]
15:05:44arkiver (arkiver) joins
15:05:44@ChanServ sets mode: +o arkiver
17:00:55<ivan>That's not too bad
17:01:44<ivan>at 2x redundancy that's 477 new hard drives per day
17:35:54@arkiver quits [Excess Flood]
17:36:35arkiver (arkiver) joins
17:36:35@ChanServ sets mode: +o arkiver
18:23:54VADemon quits [Client Quit]
20:26:04qw3rty_ quits [Ping timeout: 240 seconds]
20:31:24qw3rty joins
20:59:32lunik1 quits [Read error: Connection reset by peer]
20:59:34lunik1 joins
21:31:12HP_Archivist quits [Client Quit]
21:33:15fionera quits [Ping timeout: 264 seconds]
22:09:03fionera joins
22:09:06fionera quits [Signing in (fionera)]
22:09:07fionera (Fionera) joins
22:09:14fionera is now known as RJHacker65978
22:38:09Arcorann (Arcorann) joins
22:38:27lunik16 joins
22:39:10lunik1 quits [Ping timeout: 252 seconds]
22:39:10lunik16 is now known as lunik1
23:03:51RJHacker65978 quits [Quit: RJHacker65978]
23:03:57fionera (Fionera) joins