01:39:49qwertyasdfuiopghjkl quits [Client Quit]
02:29:03Iki joins
03:01:32<Ryz>Wondering how much we barreled through right now~ o.o;
03:57:39<Wolfin>Soooo many sites
03:59:32<@JAA>omnomnomnom
03:59:51<@JAA>Almost 3k done
04:00:25<Wolfin>Awesome!
04:14:35<Ryz>Hmm, makes me wanna do a bit of a self-challenge on this :p
04:14:53<Ryz>Have all the websites finish within 12 hours since the start of running the jobs from queuebot >;o
04:17:06<Ryz>lol http://dashboard.at.ninjawedding.org/finished takes a notable moment now to load all the jobs xD
04:17:41<Ryz>And yeah, uhh, there's a couple of jobs that aren't run properly because of pending being 5 max for voiced users, so they have to be run manually~
04:17:52<Ryz>Most likely have to do that when all the jobs went through
04:35:10Wolfin nods
05:31:43<Ryz>JAA, pondering whether it's better to temporarily give queuebot OP S:
06:14:35<Ryz>Yeeeeeeeah, major desync issues? There's now constant hitting of the pending limit
06:33:09<Ryz>Oh, it's more of AK's pipelines having problems just enough for the queuebot to hit the pending constantly now... S:
07:01:49<AK>Ooh, what's the problem with the pipelines?
07:10:11<Ryz>I'm not sure if it's because it can't keep up with the ongoing jobs that are trying to be slotted through, but there's some delay from your pipeline at least from checking http://dashboard.at.ninjawedding.org/pipelines
07:10:23<Ryz>Probably should wait for JAA for further checking
07:10:52<Ryz>There seems to be a constant 2-3 minute delay oo;
07:11:12<Ryz>Both of the Igloo-EU-OSS pipelines as well but it seems to have recovered
07:12:10<AK>Oooh weird, I noticed that with hel3 when it was first setup, but it then seemed to be okay
07:12:54<AK>CPU +memory usage on the servers is still fairly low
07:12:58<AK>Wonder what's going on
07:14:36<Ryz>JAA said it's due to the latency of being located in Europe, which makes having to get the pending goods a bit more time consuming
07:15:12<Ryz>Although, I'm curious if it's solely because of the constant jobs going through or it's stuck like that now even when the jobs have been done tossing through
07:59:20tzt quits [Ping timeout: 250 seconds]
11:28:57Iki quits [Ping timeout: 244 seconds]
11:59:27AK quits [Quit: AK]
12:01:08AK (AK) joins
12:06:56AK quits [Client Quit]
12:09:11AK (AK) joins
13:22:50<@JAA>AK, Ryz: Yeah, just latency and slowness of dequeueing from Redis by pipelines that aren't in North America.
14:21:56<Sanqui><Aramaki> queueh2ibot: Sorry, all pipelines are currently full, and only opped users can add to the queue beyond 5 pending. Please try again later.
14:21:59<Sanqui>fyi
15:02:06tzt joins
15:28:44<@JAA>Yup, known, see above.
15:29:30<@JAA>When too many jobs finish within a short time frame and newly submitted jobs aren't dequeued quickly enough, pending can reach 5, and further jobs are blocked.
16:34:46<Ryz>So, excluding the ones that were skipped because pending limit, how much more to go?
16:39:42<@JAA>50-ish
16:40:47<@JAA>I'll deal with the pending limit ones.
16:46:10<AK>Guessing you're gonna grep for the pipelines are full message and then rerun from that?
16:47:54<@JAA>Nope, other way around, extract the ones that were acknowledged as queued by Aramaki and then grep -v those from the command list.
17:18:56<AK>Ahh
17:18:57<@JAA>412 to rerun
17:19:06<AK>That's a clever way of doing iot
17:19:27<@JAA>The problem with going your way is that you can't reliably match up the bot responses to the !a commands.
17:19:57<AK>I was thinking grep -C 2 which would show 2 above and below, then hope they match up
17:19:59<AK>But good point
17:20:00<@JAA>Since Aramaki sometimes got slightly backlogged with its messages.
17:20:45<@JAA>Rerun starts now.
17:21:26<@JAA>It also checks for the pending size now.
19:02:59<@JAA>Done
19:08:56<@JAA>Confirmed that all 6805 !a commands in the list ran successfully.
19:09:13<@JAA>(Doesn't mean all content was grabbed successfully, of course.)
19:09:35IDK (IDK) joins
19:12:31<Ryz>I'm curious on the not "all content was grabbed successfully" part o:
19:12:52<Ryz>Well, aside from which that's all what we got from wolfin's scraping of userpages under http://www3.sympatico.ca/
19:17:02<Wolfin>Yeah, my scrape was pretty deep for that set (all existing WBM runs, all archive.is runs, SERP from Google, Bing and Alexa, then linkchecker to traverse all the sites and find even more user directories) but if you had a user page with no in-links you could still be left out.
19:17:47<Wolfin>Beyond that, if sites are poorly constructed (weird iframes with JS loading, etc) wpull could have trouble finding all those in-links, too.
19:19:55<Wolfin>It's still going to be the best archive set to date, even with those issues.
20:14:17<@JAA>I mean I don't know whether AB recursed through everything on every site. Maybe there's weird JS shit etc.
20:14:31<@JAA>Impossible to verify without manually checking each site.
20:18:10<Ryz>I know I encountered a website that's JS ridden because it's "Made on a Mac" x_x;
20:34:16<Ryz>Another one that you have to navigate through Adobe Flash
21:39:55qwertyasdfuiopghjkl joins
21:50:08<@JAA>Researching old ISPs' web hostings is a depressing affair. Just added a bunch of Italian ones to the wiki page.
22:25:02<Ryz>More diggy ;-;
23:19:55Wolfin cracks fingers