| 00:04:46 | | hamouda quits [Quit: Ooops, wrong browser tab.] |
| 00:10:38 | | Webuser087640 joins |
| 00:14:10 | | Webuser087640 quits [Client Quit] |
| 00:19:31 | | etnguyen03 (etnguyen03) joins |
| 00:48:55 | | nine quits [Quit: See ya!] |
| 00:49:07 | | nine joins |
| 01:01:08 | | hyperreal quits [Remote host closed the connection] |
| 01:15:49 | <h2ibot> | AntiFrutigerAero edited List of websites excluded from the Wayback Machine/Partial exclusions (+31, Added http://www.angelfire.com/md/): https://wiki.archiveteam.org/?diff=62886&oldid=61284 |
| 01:15:50 | <h2ibot> | AntiFrutigerAero edited List of websites excluded from the Wayback Machine (+136, Added http://www.drweird.com/,…): https://wiki.archiveteam.org/?diff=62887&oldid=62828 |
| 01:15:51 | <h2ibot> | Gabrinori edited Alive... OR ARE THEY (+281, /* Alarm */ Add São Paulo Antiga): https://wiki.archiveteam.org/?diff=62888&oldid=62418 |
| 01:15:52 | <h2ibot> | ChippyTechOfficial edited Valhalla (+368): https://wiki.archiveteam.org/?diff=62889&oldid=59635 |
| 01:22:05 | <h2ibot> | KleaBot edited List of websites excluded from the Wayback Machine (+0, Reordered websites and/or updated count.): https://wiki.archiveteam.org/?diff=62890&oldid=62887 |
| 01:23:00 | <klea> | T31M, Exorcism: The Blice project has pooped up in projects.json. |
| 01:23:10 | <klea> | Detected in https://changes.nulldata.foo/diff/13574663-5cde-4f6a-ab11-ee320e8258ac. |
| 01:23:56 | <klea> | uhh, https://changes.nulldata.foo/diff/13574663-5cde-4f6a-ab11-ee320e8258ac?from_version=1781766804&to_version=1782696100#text |
| 01:39:24 | | McAfee leaves [Disconnected: Replaced by new connection] |
| 01:39:27 | | McAfee joins |
| 01:43:26 | | dabs quits [Read error: Connection reset by peer] |
| 01:46:13 | <Dango360> | klea: *popped up? 😅 |
| 01:46:29 | <klea> | Yeah, sorry, I'm sleepy. |
| 01:47:10 | <klea> | Pooped is considered a word in my client's dictionary, so my spell checker didn't save me. |
| 02:05:12 | <h2ibot> | Cooljeanius edited Deathwatch (+2, /* 2026-07 */ copyedit (further rewording still…): https://wiki.archiveteam.org/?diff=62891&oldid=62839 |
| 02:07:12 | <h2ibot> | Cooljeanius edited Deathwatch (-1, /* 2026-06 */ copyedit): https://wiki.archiveteam.org/?diff=62892&oldid=62891 |
| 02:13:18 | <nicolas17> | it is a word (: |
| 02:15:15 | <klea> | Not the one I wanted to use, but yeah. |
| 02:19:59 | <Dango360> | ooo |
| 02:21:13 | | etnguyen03 quits [Client Quit] |
| 02:21:43 | <Dango360> | oops, didn't mean to send that |
| 02:24:02 | <@arkiver> | blice is started |
| 02:24:25 | <@arkiver> | i'm not fully confident in the code yet, will do some more tests tomorrow, but should be good. we can always requeue items |
| 02:24:31 | <@arkiver> | so just starting it now anyway since deadline is close |
| 02:25:48 | | etnguyen03 (etnguyen03) joins |
| 02:29:11 | | @arkiver is off for some sleep |
| 02:30:59 | <cruller> | arkiver++ |
| 02:30:59 | <eggdrop> | [karma] 'arkiver' now has 118 karma! |
| 02:33:01 | | gatagoto (gatagoto) joins |
| 02:39:48 | | etnguyen03 quits [Remote host closed the connection] |
| 02:43:37 | <legoktm> | I don't see a docker image for blice yet? |
| 02:44:35 | <klea> | JAA, can you probe drone? |
| 03:15:46 | | DogsRNice quits [Read error: Connection reset by peer] |
| 03:32:51 | | Island quits [Read error: Connection reset by peer] |
| 03:43:51 | <Dango360> | times: item just has one url `202=302 https://blice.co.kr/web/viewer.kt?timesId=242208` |
| 03:43:52 | <Dango360> | is it meant to do that? 302 without anything else done to it? |
| 04:00:28 | | BennyOtt quits [Quit: Bye] |
| 04:00:52 | | BennyOtt (BennyOtt) joins |
| 04:07:55 | | thewinwin84 joins |
| 04:11:47 | | thewinwin8 quits [Ping timeout: 268 seconds] |
| 04:11:55 | | thewinwin84 is now known as thewinwin8 |
| 04:45:24 | <hexagonwin> | blice docker image still seems to be missing :( |
| 04:52:29 | | h|ca2 quits [Ping timeout: 268 seconds] |
| 04:53:37 | | h|ca2 (h) joins |
| 05:16:53 | | atphoenix__ (atphoenix) joins |
| 05:19:37 | | atphoenix_ quits [Ping timeout: 268 seconds] |
| 05:32:30 | | thoughtrise joins |
| 05:32:59 | <thoughtrise> | hi friends :) |
| 05:43:41 | | thoughtrise quits [Client Quit] |
| 06:07:17 | <Exorcism> | thanks kela! |
| 06:07:21 | <Exorcism> | klea* |
| 06:29:18 | | gatagoto quits [Ping timeout: 268 seconds] |
| 06:30:44 | | gatagoto (gatagoto) joins |
| 06:45:37 | | gatagoto quits [Client Quit] |
| 07:02:04 | | grill_ (grill) joins |
| 07:05:19 | | grill quits [Ping timeout: 248 seconds] |
| 07:20:49 | | lflare (lflare) joins |
| 07:30:50 | | lflare is now authenticated as * |
| 07:30:51 | | lflare is now known as RJHacker58012 |
| 07:30:52 | | lflare (lflare) joins |
| 07:32:10 | | lflare quits [Client Quit] |
| 07:32:56 | | lflare (lflare) joins |
| 07:33:26 | | RJHacker58012 quits [Ping timeout: 268 seconds] |
| 07:36:35 | <cruller> | FWIW, you can subscribe to #at-change by topic (ATDocker, ATWarrior, ATGit, etc.) via ntfy. |
| 07:42:40 | <cruller> | *changes |
| 08:13:38 | <h2ibot> | Usernam edited List of websites excluded from the Wayback Machine (+21, https://en.btdig.com/about/ =…): https://wiki.archiveteam.org/?diff=62893&oldid=62890 |
| 08:17:50 | | dendory quits [Ping timeout: 268 seconds] |
| 08:18:33 | | dendory (dendory) joins |
| 08:30:10 | | dendory quits [Ping timeout: 268 seconds] |
| 08:37:43 | | dendory (dendory) joins |
| 08:42:10 | | flotwig quits [Read error: Connection reset by peer] |
| 08:42:32 | | flotwig joins |
| 08:54:11 | | thewinwin88 joins |
| 08:57:55 | | thewinwin8 quits [Ping timeout: 268 seconds] |
| 08:57:59 | | thewinwin88 is now known as thewinwin8 |
| 09:06:56 | | Webuser523822 joins |
| 09:09:15 | | pabs quits [Read error: Connection reset by peer] |
| 09:10:29 | | pabs (pabs) joins |
| 09:10:57 | | Webuser523822 quits [Client Quit] |
| 09:33:15 | <yzqzss> | Hello everyone, I would like to propose an optional WARC-zstd extension spec for discussion. |
| 09:33:22 | <yzqzss> | The current WARC-zstd format already allows a single WARC record to consist of multiple zstd frames. Based on this, I have been experimenting with a record-local multi-frame layout: |
| 09:33:29 | <yzqzss> | 1. one zstd frame for the WARC record header; |
| 09:33:43 | <yzqzss> | 2. zero or more zstd frames for the record payload; |
| 09:33:58 | <yzqzss> | 3. one trailing skippable frame containing a zstd seekable table for the frames of this record. |
| 09:34:11 | <yzqzss> | ( For records with Content-Length: 0, the compact recommended layout is to put the WARC header and record trailer in the first frame, followed by a one-entry seekable table. In that case there is no separate payload frame.) |
| 09:34:40 | <yzqzss> | In my tests, separating the WARC header from the payload can slightly improve compression ratio. More importantly, the record-local seek table makes it possible to walk backward from the end of a WARC file and rapidly discover record boundaries without scanning the whole compressed warc.zst file. |
| 09:35:07 | <yzqzss> | For large HTTP payloads, such as MP4/MP3/ZIP/PDF or other range-friendly resources, the payload can also be split into multiple zstd frames. A replay system could then use the per-record seek table to load only the frames overlapping a requested byte range, instead of decompressing the whole warc record. |
| 09:35:40 | <yzqzss> | This should remain backward-compatible with compliant WARC-zstd readers: unsupported readers will simply ignore the skippable seek-table frames, and the concatenation of all zstd data frames still produces a valid WARC stream. The main compatibility concern is readers that assume “one WARC record = one zstd frame”; those implementations may need to be adjusted, since the existing WARC-zstd model already permits one record to span multiple |
| 09:35:40 | <yzqzss> | frames. |
| 09:35:56 | <yzqzss> | <- END -> |
| 09:44:19 | <Exorcism> | arkiver: error for blice (log): https://x0.at/7gYh.txt |
| 09:54:35 | <yzqzss> | What do you think of this idea? is it good to have :) |
| 10:18:04 | <yzqzss> | https://github.com/iipc/warc-specifications/issues/118 |
| 10:18:04 | <yzqzss> | just created an issue for this |
| 10:29:57 | <@arkiver> | yzqzss: are the reasons for introducing this only faster seeking and saving 0.1-1.0% of space? |
| 10:30:39 | <@arkiver> | it is often useful to have an example of a case in which this clearly benefits things |
| 10:31:54 | <yzqzss> | arkiver: no, not only that |
| 10:34:51 | <yzqzss> | it's good for replaying system, good for process records in parallel, good for quickly locating a record by reading minimal data without a cdx index. |
| 10:34:52 | <klea> | Are we still relying on Quad9 for DNS? |
| 10:35:19 | <@arkiver> | klea: yes, nearly nearly nearly ready for running out own DNS resolver |
| 10:35:20 | <klea> | Oh, that jumpscared me. <https://uptime.quad9.net/incident/938372> |
| 10:35:45 | <@arkiver> | i remember when i thought it would take me a week to get DoH and DoT in Wget-AT :P |
| 10:35:45 | <klea> | #quad9-status got a "quad9.net is down" which made me think their resolver was down, but no, only website. |
| 10:35:55 | <klea> | hmm. |
| 10:36:16 | <klea> | Next up, testing arm for Wget-AT? |
| 10:42:01 | <h2ibot> | KleaBot edited Main Page/Current Warrior Project (+98, Default projects are now robloxgroups (weight…): https://wiki.archiveteam.org/?diff=62894&oldid=62699 |
| 10:42:28 | <yzqzss> | arkiver: I conducted the tests using my few own warc files. I plan to public a script later so that you guys can evaluate the extension and see if it offers improvements on your warcs |
| 10:59:07 | | aliz joins |
| 11:00:19 | | Bleo18260072271962345522201107 quits [Quit: The Lounge - https://thelounge.chat] |
| 11:03:04 | | Bleo18260072271962345522201107 joins |
| 11:04:03 | <aliz> | Trying to start grabbing blice but it seems the docker image for blice-grab hasn't been added to atdr.meo.ws |
| 11:06:37 | <klea> | arkiver: ^, you made it a default project, so not sure what we're expected to see, but also, how did Exorcism get the error log shared if the image wasn't made? |
| 11:12:46 | | Webuser227848 joins |
| 11:24:15 | | tertu2 quits [Quit: so long...] |
| 11:24:35 | | tertu (tertu) joins |
| 11:36:50 | | Webuser062556 joins |
| 11:38:51 | | Webuser062556 quits [Client Quit] |
| 11:39:14 | | Webuser481930 joins |
| 11:39:55 | | Webuser481930 quits [Client Quit] |
| 11:39:59 | | Webuser767462 joins |
| 12:16:01 | <IDK> | blice website doesnt look very good |
| 12:21:34 | <@arkiver> | yzqzss: alright, thank you! |
| 12:32:51 | | Webuser227848 quits [Client Quit] |
| 12:44:09 | | twiswist quits [Read error: Connection reset by peer] |
| 12:44:19 | | twiswist (twiswist) joins |
| 12:45:29 | | Webuser259977 joins |
| 12:52:00 | | FiTheArchiver joins |
| 12:54:47 | <FiTheArchiver> | hey guys, was wondering if anyone could help me with this. so from 2014-2016 i had a twitter account and unfortunately my ipad from that time period wiped itself, and i lost the account to a suspension from a hack. i wanted to be able to document every tweet to the account because there's obviously some memories there, i started doing it a few months ago with nitter.net but the search on there has |
| 12:54:47 | <FiTheArchiver> | become awful, obviously same with actual twitter. i was wondering is there any other way i can archive this? or is twitter search just over forever now |
| 12:55:54 | <FiTheArchiver> | it's really important to me to archive everything from this time period in my life because i lost it all from my twitter and ipad. ugh it just sucks. and i can't even download the archiveteam tweet stream things from 2014 |
| 12:59:59 | <skankhunt42> | is blice only available through warrior? |
| 13:06:26 | <h2ibot> | Manu edited Discourse/inactive (+20, Queued https://discourse.federated.computer/): https://wiki.archiveteam.org/?diff=62895&oldid=62786 |
| 13:06:27 | <h2ibot> | Manu edited Discourse/archived (+107, Queued https://discourse.federated.computer/): https://wiki.archiveteam.org/?diff=62896&oldid=62708 |
| 13:07:13 | | FiTheArchiver quits [Read error: Connection reset by peer] |
| 13:08:24 | | FiTheArchiver joins |
| 13:08:26 | <h2ibot> | Manu edited Discourse/inactive (+14, Queued https://forum.otrscommunityedition.com/): https://wiki.archiveteam.org/?diff=62897&oldid=62895 |
| 13:08:27 | <h2ibot> | Manu edited Discourse/archived (+109, Queued https://forum.otrscommunityedition.com/): https://wiki.archiveteam.org/?diff=62898&oldid=62896 |
| 13:09:23 | | McAfee leaves [Error from remote client] |
| 13:09:25 | | McAfee joins |
| 13:10:27 | <h2ibot> | Manu edited Discourse/inactive (+17, Queued https://discourse.rahul.net/): https://wiki.archiveteam.org/?diff=62899&oldid=62897 |
| 13:10:28 | <h2ibot> | Manu edited Discourse/archived (+98, Queued https://discourse.rahul.net/): https://wiki.archiveteam.org/?diff=62900&oldid=62898 |
| 13:12:27 | <h2ibot> | Manu edited Discourse/archived (+96, Queued https://forum.abeedesk.com/): https://wiki.archiveteam.org/?diff=62901&oldid=62900 |
| 13:12:28 | <h2ibot> | Manu edited Discourse/inactive (+12, Queued https://forum.abeedesk.com/): https://wiki.archiveteam.org/?diff=62902&oldid=62899 |
| 13:13:52 | | thewinwin86 joins |
| 13:14:39 | <cruller> | According to the tracker, the answer is no. They are likely using a Docker image they built themselves. |
| 13:16:02 | <skankhunt42> | https://github.com/ArchiveTeam/blice-grab looks like at least a repo is there but no image was built. I will check that out later maybe, looks like its not that much time left :/ |
| 13:16:06 | <cruller> | Among recent ones, ktoon was like this too, IIRC. |
| 13:17:35 | | thewinwin8 quits [Ping timeout: 248 seconds] |
| 13:17:40 | | thewinwin86 is now known as thewinwin8 |
| 13:18:28 | <h2ibot> | Manu edited Discourse/inactive (-5): https://wiki.archiveteam.org/?diff=62905&oldid=62902 |
| 13:25:46 | <h2ibot> | Manu edited Discourse/inactive (+17, Queued https://forum.e-handel.info/): https://wiki.archiveteam.org/?diff=62906&oldid=62905 |
| 13:25:47 | <h2ibot> | Manu edited Discourse/archived (+98, Queued https://forum.e-handel.info/): https://wiki.archiveteam.org/?diff=62907&oldid=62901 |
| 13:38:54 | <skankhunt42> | not exactly sure how do build docker images myself, didn't find a wiki entry for that. only option I see is running the pipeline directly on the host for now |
| 13:38:57 | | McAfee leaves |
| 13:39:42 | | McAfee joins |
| 13:42:06 | <skankhunt42> | nvm its actually super easy lol, just docker build. |
| 13:55:05 | | nexussfan (nexussfan) joins |
| 14:00:16 | <th3z0l4|m> | putting a little effort on blice rn |
| 14:07:14 | | TunaLobster44 quits [Quit: So long and thanks for all the fish] |
| 14:10:58 | <@imer> | cruller: skankhunt42: poked drone, image should be available in a few seconds |
| 14:11:17 | <skankhunt42> | ah neat, thanks! I was just done setting up a github fork + action :D |
| 14:12:05 | <klea> | imer: Should I have pinged you instead of only pinging JAA and arkiver? |
| 14:12:44 | | TunaLobster44 joins |
| 14:14:09 | <@imer> | klea: sure |
| 14:15:35 | | TunaLobster44 quits [Client Quit] |
| 14:17:09 | <skankhunt42> | thanks imer, set up on a few nodes. still think this might be a tough one to beat. you know about rate limits or something? I don't want to burn my ip too fast |
| 14:20:42 | <skankhunt42> | seems containers are crashing sometimes: Lua runtime error: blice.lua:864: attempt to concatenate local 'items' (a nil value) |
| 14:21:31 | <skankhunt42> | and Lua runtime error: blice.lua:613: Expected value but found invalid token at character 1 ([C]: in function 'decode' || blice.lua:656: in function <blice.lua:206>.) |
| 14:21:53 | <skankhunt42> | Lua runtime error: blice.lua:656: Expected value but found invalid token at character 1 |
| 14:24:41 | <cruller> | imer++ |
| 14:24:41 | <eggdrop> | [karma] 'imer' now has 38 karma! |
| 14:25:08 | <klea> | imer++ |
| 14:25:08 | <eggdrop> | [karma] 'imer' now has 39 karma! |
| 14:25:10 | <klea> | imer++ |
| 14:25:11 | <eggdrop> | [karma] 'imer' now has 40 karma! |
| 14:26:21 | <klea> | skankhunt42: Could you take the full logs for the different errors, in order, save them as text files, and upload them to https://transfer.archivete.am/ ? |
| 14:26:40 | <@imer> | skankhunt42: haven't looked at blice at all yet so far, so no idea. (and what klea said, more context before the error would be good) |
| 14:28:47 | <@arkiver> | thanks for the reports skankhunt42 , checking |
| 14:36:43 | <skankhunt42> | klea: https://transfer.archivete.am/zOk7O/merged_logs.txt (hope I didn't strip too much) |
| 14:36:44 | <eggdrop> | inline (for browser viewing): https://transfer.archivete.am/inline/zOk7O/merged_logs.txt |
| 14:37:39 | | grill_ is now known as grill |
| 14:37:45 | <skankhunt42> | imer++ |
| 14:37:46 | <eggdrop> | [karma] 'imer' now has 41 karma! |
| 14:37:51 | <skankhunt42> | for the image <3 |
| 14:38:13 | <klea> | Yeah, that seems to be interesting, the script is trying to submit an empty list of items. |
| 14:39:03 | <klea> | I guess what happens when you ^C^V other repos whilst being sleepy :) |
| 14:39:12 | <klea> | arkiver++ |
| 14:39:13 | <eggdrop> | [karma] 'arkiver' now has 119 karma! |
| 14:39:35 | <klea> | Soon arkiver will return a 200 OK in eggdrop's karma score :) |
| 14:39:49 | <skankhunt42> | arkiver++ |
| 14:39:49 | <eggdrop> | [karma] 'arkiver' now has 120 karma! |
| 14:41:02 | <skankhunt42> | thanks for all your efforts btw, I'm pretty new to all this stuff and I love it. If I had more connections and resources, I'd throw a lot more in. currently having access to a public /26 v4 but I assume this would just lead to range bans, right |
| 14:43:12 | <@arkiver> | skankhunt42: no idea, some sites just don't care. sometimes it also doesn't matter since the site is going away anyway. of course there can be cases in which a block/ban/limit affects access to other sites form the same provider or on the same IP too |
| 14:44:41 | <skankhunt42> | okay, lets see then. :) |
| 14:46:20 | | unknownsrc quits [Ping timeout: 268 seconds] |
| 14:48:35 | <skankhunt42> | yeah. got blocked already. |
| 14:48:55 | | aliz quits [Quit: Ooops, wrong browser tab.] |
| 14:56:40 | <skankhunt42> | I've seen a generic "web firewall block" but I think it doesn't block assets on cds.blice.co.kr (main page is super slow right now anyway, maybe already OL or bad peering on my end) |
| 14:57:45 | <@arkiver> | skankhunt42: what status code does it give? |
| 15:00:09 | <@arkiver> | skankhunt42: reported error should be fixed, minimum version is increased |
| 15:00:25 | <skankhunt42> | 200 unfortunately. https://transfer.archivete.am/zZi79/index.html |
| 15:00:25 | <eggdrop> | inline (for browser viewing): https://transfer.archivete.am/inline/zZi79/index.html |
| 15:00:47 | <skankhunt42> | arkiver++ |
| 15:00:47 | <eggdrop> | [karma] 'arkiver' now has 121 karma! |
| 15:03:15 | <klea> | Time to do search for "Web firewall security policies" in the results? |
| 15:03:36 | | unknownsrc (unknownsrc) joins |
| 15:03:46 | <skankhunt42> | weirdly this WAF only blocks access to https://blice.co.kr which forwards to https://blice.co.kr/home.kt?needGaLog=false (showing that index.html I uploaded). Accessing cds or even "https://blice.co.kr/mw/comment.kt?timesId=3373193&novelId=69663&commentType=times&quickViewYn=N" (I just grabbed that from the log) seems to still work? |
| 15:05:47 | <skankhunt42> | or maybe I was just lucky. not 100% sure when the WAF triggers. but yeah, maybe I submitted a lot of that 200 garbage results. |
| 15:09:46 | <@arkiver> | skankhunt42: another updateis out |
| 15:10:10 | <skankhunt42> | cds.blice.co.kr/download?file=/epub/noveltimes/drm/2025/02/10/drme_1739172952207.epub looks like there is no block on CDS, at least I couldn't trigger yet |
| 15:10:24 | | Island joins |
| 15:15:18 | <skankhunt42> | arkiver looks like everything gets rate limited now. how is that calculated? |
| 15:15:36 | <IDK> | uh oh, looks like my AB jobs might not have worked |
| 15:15:43 | <@arkiver> | skankhunt42: yes i paused things |
| 15:15:48 | <@arkiver> | we'll resume in a bit |
| 15:15:51 | <skankhunt42> | understood :) |
| 15:16:40 | <@arkiver> | skankhunt42: just double checking, you're sure about the 200 on the blocked page? or could it have been a mixup? |
| 15:17:48 | | klea wonders if it's easier to connect to the target and check WARCs. |
| 15:18:10 | <skankhunt42> | I only did a quick wget and yeah, HTTP request sent, awaiting response... 200 OK |
| 15:18:16 | <@arkiver> | alright |
| 15:18:27 | <@arkiver> | klea: no, WARC are sent on to IA fast, i'll check the CDX |
| 15:18:50 | <klea> | Woah, I thought it had some slowdown before getting to IA. |
| 15:19:00 | | Webuser907632 joins |
| 15:19:25 | <skankhunt42> | hope I didn't cause too much trouble :x |
| 15:19:43 | <@arkiver> | skankhunt42: it's trouble, but not caused by you. |
| 15:19:45 | <@arkiver> | :) |
| 15:19:48 | <skankhunt42> | <3 |
| 15:19:54 | <klea> | Trouble caused by Blince :) |
| 15:20:15 | <IDK> | #BliceOfShit |
| 15:20:16 | | h|ca2 quits [Ping timeout: 248 seconds] |
| 15:20:29 | <Webuser907632> | f |
| 15:21:00 | <skankhunt42> | they probably cant handle that much traffic. I wonder if the WAF is so overloaded that it doesn't even process details like detect time, IP and url? |
| 15:21:05 | <@arkiver> | resumed now |
| 15:21:15 | <skankhunt42> | new code too? |
| 15:21:20 | <@arkiver> | yes and forced as minimum |
| 15:21:51 | | moth3 quits [Ping timeout: 248 seconds] |
| 15:22:49 | <skankhunt42> | should I worry about my scaling or is it handled on tracker site? |
| 15:22:52 | | arch quits [Remote host closed the connection] |
| 15:23:05 | | arch (arch) joins |
| 15:23:12 | <klea> | The tracker limits the max amount of items that will flow out, AFAIK. |
| 15:23:12 | | Matthww quits [Quit: The Lounge - https://thelounge.chat] |
| 15:23:20 | <klea> | s/amount/rate/ |
| 15:24:00 | <@arkiver> | skankhunt42: this should be handled now, but it could make items go to claims very fast |
| 15:24:13 | | nexussfan quits [Remote host closed the connection] |
| 15:25:10 | <IDK> | does anyone know how many concurrency is safe before you hit 200? |
| 15:25:22 | | Matthww joins |
| 15:25:47 | <skankhunt42> | I'd assume 0 to not hit 200 :D |
| 15:25:54 | <IDK> | :p |
| 15:26:06 | <IDK> | I mean waf |
| 15:27:06 | <klea> | IDK: I'd say 50000000000 may make the machine crash, and not hit 200, but we don't want that. |
| 15:27:07 | <h2ibot> | Skankhunt42 edited Blice (+181): https://wiki.archiveteam.org/?diff=62910&oldid=62851 |
| 15:27:08 | <h2ibot> | Arkiver changed the user rights of User:Skankhunt42 |
| 15:27:08 | <skankhunt42> | I am running about 100 concurrent from 3 nodes I guess. seeing some "Server returned bad response", not sure if that covers the WAF. |
| 15:27:09 | <h2ibot> | Hans5958 moved Template:Tracker item to Template:Tracker project: https://wiki.archiveteam.org/?title=Template%3ATracker%20project |
| 15:27:53 | <@arkiver> | skankhunt42: maybe, does it have status code 200? |
| 15:28:27 | | h|ca2 (h) joins |
| 15:29:05 | <skankhunt42> | looks like it: https://transfer.archivete.am/hLYKl/merged_logs.txt |
| 15:29:06 | <eggdrop> | inline (for browser viewing): https://transfer.archivete.am/inline/hLYKl/merged_logs.txt |
| 15:29:50 | <skankhunt42> | but its fetching a lot of assets and finds items (I assume that means its working right) |
| 15:30:13 | <Webuser907632> | I'm getting 200 and have way less concurrent |
| 15:30:31 | <skankhunt42> | 200 with WAF? or just normal 200 and OK? |
| 15:30:55 | <skankhunt42> | not sure how I would spot hitting the firewall block in the logs |
| 15:30:55 | | arch quits [Ping timeout: 248 seconds] |
| 15:32:07 | <IDK> | 200 would be normal, what we have here is 200 and its actually banned 😭 |
| 15:32:14 | <IDK> | 10 linodes coming right up |
| 15:32:58 | <skankhunt42> | I don't like it when they don't use status codes. at least use 418 or something funny |
| 15:32:59 | <@arkiver> | we don't need a ton more workers on this |
| 15:33:05 | <@arkiver> | the site is not holding up great already |
| 15:34:02 | <@arkiver> | Webuser907632: what do you see on the web page? |
| 15:34:46 | <Webuser907632> | Server returned bad response. Sleeping 2 seconds. |
| 15:34:52 | <Webuser907632> | If that's what you're even asking |
| 15:35:35 | | McAfee leaves |
| 15:36:27 | | arch (arch) joins |
| 15:38:02 | | arch quits [Remote host closed the connection] |
| 15:38:26 | | arch (arch) joins |
| 15:43:01 | <IDK> | Webuser907632 if that is 0, then it means the server is overloaded and isnt returning anything |
| 15:43:24 | <IDK> | aka starting with 0 = [URL] |
| 15:44:40 | <Webuser907632> | I figured we were hugging their server to death, yeah |
| 15:44:49 | | McAfee joins |
| 15:44:51 | <skankhunt42> | yeah indeed |
| 15:47:17 | <skankhunt42> | accessing from a clean IP (not used for blice): wget needed 4 tries to get "/web/homescreen/main.kt?service=WEBNOVEL&genre=romance" |
| 15:48:00 | <Webuser907632> | 100% getting worse |
| 15:48:01 | <IDK> | anyways, great its almost 1am in korea |
| 15:48:17 | <skankhunt42> | maybe we get all items before breakfast |
| 15:48:35 | <IDK> | or else arkiver could lower the speed before 8am in korea |
| 15:49:25 | <klea> | Potato |
| 15:49:28 | <IDK> | i doubt KT would care about a dying website but who knows |
| 15:51:08 | <nicolas17> | I just uploaded a .part (partial download from browser) to IA so that's another check I need to add to my scripts -_- |
| 15:55:30 | <klea> | How come you got a partial download out of a at most 1GB file? |
| 15:55:31 | <klea> | Oh wait. |
| 15:55:35 | <klea> | Samsung-- |
| 15:55:35 | <eggdrop> | [karma] 'Samsung' now has -2 karma! |
| 15:58:10 | <nicolas17> | 26MB NOTICE.html, connection dropped halfway |
| 15:59:09 | <klea> | Oh, those are the small files you download manually, bypassing the crowdsource. |
| 16:00:04 | <klea> | (And thus the tooling that checks file sizes.) |
| 16:11:24 | | McAfee leaves |
| 16:13:24 | | Webuser767462 quits [Quit: Ooops, wrong browser tab.] |
| 16:19:41 | | thewinwin81 joins |
| 16:23:46 | | thewinwin8 quits [Ping timeout: 268 seconds] |
| 16:23:48 | | thewinwin81 is now known as thewinwin8 |
| 16:23:54 | | FiTheArchiver quits [Quit: Leaving] |
| 16:25:54 | <fuzzy80211> | arkiver i can turn up blice if you want to move warriors away otherwise i will stay away |
| 16:27:38 | <th3z0l4|m> | wow im getting heavy rate limited by the tracker |
| 16:27:56 | <klea> | IIRC blice already has enough workers, since there's a few people who cranked up concurrency, and a few Warriors should be doing it too. |
| 16:28:21 | <klea> | th3z0l4|m: The ratelimit is global to the project, not per user, but yeah. |
| 16:29:22 | <fuzzy80211> | klea correct. comment was if someone wanted to move warrior default to a different project |
| 16:29:38 | <klea> | Oh. |
| 16:29:55 | <klea> | Sorry, I understood it another way. |
| 16:30:06 | <that_lurker> | one fuzzy = all warrior workers :-P |
| 16:30:24 | <th3z0l4|m> | klea: i know, since some more people joined and other more raised concurrency im moving a few nodes away :) |
| 16:30:49 | <klea> | hmm, 50% of the Warriors are on it. |
| 16:31:04 | <klea> | And yeah, nice to have fuzzys. |
| 16:31:18 | <h2ibot> | Manu edited Discourse/inactive (-5): https://wiki.archiveteam.org/?diff=62913&oldid=62906 |
| 16:32:18 | <h2ibot> | Manu edited Discourse/inactive (-10): https://wiki.archiveteam.org/?diff=62914&oldid=62913 |
| 16:32:31 | | McAfee joins |
| 16:33:21 | <@arkiver> | fuzzy80211: done, blice is no more a default |
| 16:33:23 | | MrMcNuggets (MrMcNuggets) joins |
| 16:34:02 | | michaelblob76413413 joins |
| 16:34:26 | <justauser> | If it holds 4K items/minute, it's not a potato. |
| 16:34:52 | | MrMcNugg1 quits [Ping timeout: 268 seconds] |
| 16:35:19 | <h2ibot> | KleaBot edited Main Page/Current Warrior Project (-98, Default projects are now robloxgroups (weight…): https://wiki.archiveteam.org/?diff=62916&oldid=62894 |
| 16:36:06 | | michaelblob7641341 quits [Ping timeout: 268 seconds] |
| 16:36:07 | | michaelblob76413413 is now known as michaelblob7641341 |
| 16:36:47 | <fuzzy80211> | sounds good. will wait and see what happens to the IRSR and spin up if needed |
| 16:48:54 | | hyperreal (hyperreal) joins |
| 16:50:21 | <h2ibot> | Manu edited Discourse/archived (+107, Queued https://community.abs-consulting.com/): https://wiki.archiveteam.org/?diff=62917&oldid=62907 |
| 16:50:22 | <h2ibot> | Manu edited Discourse/inactive (+69, Queued https://community.abs-consulting.com/): https://wiki.archiveteam.org/?diff=62918&oldid=62914 |
| 16:52:22 | <h2ibot> | Manu edited Discourse/inactive (+17, Queued…): https://wiki.archiveteam.org/?diff=62921&oldid=62918 |
| 16:52:23 | <h2ibot> | Manu edited Discourse/archived (+124, Queued…): https://wiki.archiveteam.org/?diff=62922&oldid=62917 |
| 16:53:05 | | ThreeHM quits [Read error: Connection reset by peer] |
| 16:53:22 | <h2ibot> | Manu edited Discourse/archived (+0): https://wiki.archiveteam.org/?diff=62923&oldid=62922 |
| 16:56:44 | | ThreeHM (ThreeHeadedMonkey) joins |
| 17:01:52 | | hyperreal quits [Client Quit] |
| 17:15:25 | <h2ibot> | Skankhunt42 edited Blice (+24): https://wiki.archiveteam.org/?diff=62926&oldid=62910 |
| 17:15:46 | | hyperreal (hyperreal) joins |
| 17:27:27 | <h2ibot> | Justauser edited Discourse/active (+81, https://community.developer.atlassian.com/): https://wiki.archiveteam.org/?diff=62929&oldid=62884 |
| 18:18:20 | | Webuser202395 joins |
| 18:23:55 | | apache2_ quits [Remote host closed the connection] |
| 18:23:58 | | apache2 joins |
| 18:47:30 | <justauser> | Allegedly Facebook is *still* doing UA sniffing with favors to Googlebot? https://is-a.cat/@tofu/116815819340160459 |
| 18:54:12 | | Webuser907632 quits [Quit: Ooops, wrong browser tab.] |
| 18:59:10 | | DogsRNice joins |
| 19:09:42 | <h2ibot> | Manu edited Webring (+33, found https://literring.neocities.org/): https://wiki.archiveteam.org/?diff=62931&oldid=62852 |
| 19:21:59 | | h|ca2 quits [Ping timeout: 268 seconds] |
| 19:38:16 | <IDK> | lmao do they not check the IP address |
| 20:12:45 | | Wohlstand1 (Wohlstand) joins |
| 20:15:07 | | Wohlstand1 is now known as Wohlstand |
| 20:45:59 | | moth3 joins |
| 21:00:34 | | JohnnyJ joins |
| 21:04:43 | | h|ca2 (h) joins |
| 21:16:46 | | Starchives__ (Starchives) joins |
| 21:20:23 | | Starchives_ quits [Ping timeout: 268 seconds] |
| 21:35:51 | | dabs joins |
| 21:36:40 | | dabs quits [Remote host closed the connection] |
| 21:36:59 | | dabs joins |
| 21:42:53 | | IHaveAnIdea joins |
| 21:44:24 | <IHaveAnIdea> | Can bowlroll.net be saved please? Bowlroll is a file hosting service, mainly used for sharing 3d models. The url structure is a simple incremental number: https://bowlroll.net/file/123456 Unfortunately, in order to download the file, you need to click a button which submit a form. For this reason, archive.org has not saved any files from that site |
| 21:44:51 | | IHaveAnIdea quits [Client Quit] |
| 21:46:38 | | etnguyen03 (etnguyen03) joins |
| 21:52:32 | <pokechu22> | Downloads on e.g. https://bowlroll.net/file/355662 require a POST to https://bowlroll.net/api/file/355662/download-check to generate a signed URL for https://bowlroll.net/file/355662/download-execution/TribalClassV1.0.1-distr.zip?one_time_key=XXXX to actually save it. Archivbot can't do POSTs; it'd need to be a custom project |
| 21:52:45 | <pokechu22> | Is the site shutting down, or just proactive? |
| 22:03:54 | | dabs quits [Read error: Connection reset by peer] |
| 22:15:40 | | nexussfan (nexussfan) joins |
| 22:20:08 | <@arkiver> | pff they left so fast |
| 22:20:44 | <klea> | Do we do projects for proactive things? |
| 22:20:53 | <klea> | (DPoS projects, to be precise) |
| 22:24:12 | <@arkiver> | it is possible |
| 22:24:20 | <@arkiver> | but there needs to be a very good reason |
| 22:25:03 | <klea> | I guess because it's more work to keep up a project than to run AB jobs? |
| 22:25:19 | <@arkiver> | partially, but also simply size |
| 22:25:32 | <klea> | Ah, yeah, true. |
| 23:08:16 | | Wohlstand quits [Client Quit] |
| 23:09:00 | | HP_Archivist quits [Quit: Leaving] |
| 23:17:25 | | etnguyen03 quits [Client Quit] |
| 23:20:36 | | cyan_box joins |
| 23:20:57 | | allani0 quits [Quit: The Lounge - https://thelounge.chat] |
| 23:23:27 | | cyanbox_ quits [Ping timeout: 248 seconds] |
| 23:24:44 | | allani03 joins |
| 23:38:25 | | hamouda joins |
| 23:53:29 | | Webuser202395 quits [Quit: Ooops, wrong browser tab.] |