| 00:34:45 | | error quits [Remote host closed the connection] |
| 00:35:09 | | error (disinfecttm) joins |
| 01:03:53 | | error quits [Remote host closed the connection] |
| 01:04:11 | | error (disinfecttm) joins |
| 01:14:23 | | MrMcNugg1 quits [Ping timeout: 248 seconds] |
| 01:17:02 | | etnguyen03 quits [Client Quit] |
| 01:20:33 | | etnguyen03 (etnguyen03) joins |
| 01:39:04 | | gatagoto (gatagoto) joins |
| 02:04:02 | | etnguyen03 quits [Remote host closed the connection] |
| 02:08:17 | | chunkynutz6018160 joins |
| 02:10:23 | | chunkynutz601816 quits [Ping timeout: 248 seconds] |
| 02:10:24 | | chunkynutz6018160 is now known as chunkynutz601816 |
| 02:25:54 | | ericgallager joins |
| 02:37:57 | | thewinwin89 joins |
| 02:42:13 | | thewinwin8 quits [Ping timeout: 268 seconds] |
| 02:42:16 | | thewinwin89 is now known as thewinwin8 |
| 02:43:19 | | L9330x3sLFE joins |
| 02:51:02 | | L9330x3sLFE quits [Client Quit] |
| 02:52:59 | | L9330x3sLFE joins |
| 02:56:20 | | L9330x3sLFE quits [Client Quit] |
| 02:58:58 | | L9330x3sLFE joins |
| 03:03:20 | <h2ibot> | Wickedplayer494 edited TV Time (+0, Is dead): https://wiki.archiveteam.org/?diff=63168&oldid=62959 |
| 03:14:16 | | Island quits [Read error: Connection reset by peer] |
| 03:44:57 | | MrMcNuggets (MrMcNuggets) joins |
| 04:07:28 | | DogsRNice quits [Read error: Connection reset by peer] |
| 04:55:54 | <Pendonym> | Any way to archive https://fogcam.org/ every 20s for the image |
| 05:12:43 | | nexussfan quits [Quit: Konversation terminated!] |
| 05:37:05 | <nicolas17> | Pendonym: WBM playback would be hard |
| 05:37:29 | <nicolas17> | they use `"fogcam2.jpg" + new Date()` probably for cache-busting purposes |
| 05:40:59 | <nicolas17> | also since I loaded the page I haven't seen the image ever change |
| 05:41:06 | <nicolas17> | does it actually change every 20s? |
| 05:45:59 | | Cupping12854 quits [Ping timeout: 268 seconds] |
| 05:46:48 | | Cupping12854 joins |
| 05:50:55 | <nicolas17> | ok yeah it updates super irregularly |
| 06:01:24 | | nulldata quits [Ping timeout: 268 seconds] |
| 06:05:13 | | nulldata (nulldata) joins |
| 06:13:33 | | MPThLee quits [Quit: bye] |
| 06:14:59 | | MPThLee (MPThLee) joins |
| 06:29:35 | | benjins3 quits [Ping timeout: 248 seconds] |
| 07:22:11 | | L9330x3sLFE quits [Ping timeout: 268 seconds] |
| 07:28:30 | | L9330x3sLFE joins |
| 07:34:40 | | h|ignition quits [Ping timeout: 272 seconds] |
| 07:36:16 | | h|ignition (h) joins |
| 07:37:27 | | _Apgap joins |
| 07:41:18 | | Boppen_ quits [Ping timeout: 268 seconds] |
| 07:41:38 | | h|ignition quits [Ping timeout: 272 seconds] |
| 07:42:35 | | h|ignition (h) joins |
| 08:02:27 | | simon816 quits [Quit: ZNC 1.10.2 - https://znc.in] |
| 08:04:40 | | BennyOtt quits [Quit: Bye] |
| 08:04:44 | | gatagoto quits [Ping timeout: 268 seconds] |
| 08:04:47 | <@arkiver> | imer: yep, we just went with the #// target, all good! :) |
| 08:06:09 | <h2ibot> | Rexma edited List of websites excluded from the Wayback Machine (+21, add xypb.net): https://wiki.archiveteam.org/?diff=63169&oldid=63154 |
| 08:07:54 | | BennyOtt (BennyOtt) joins |
| 08:10:43 | | simon816 (simon816) joins |
| 08:16:10 | <h2ibot> | KleaBot edited List of websites excluded from the Wayback Machine (+0, Reordered websites and/or updated count.): https://wiki.archiveteam.org/?diff=63170&oldid=63169 |
| 08:17:27 | <L9330x3sLFE> | Is there a difference between Discrub and DiscordChatExporter when it comes to archiving? How about measures against rate limiting? |
| 08:29:22 | | benjins3 joins |
| 08:30:07 | | Wohlstand1 (Wohlstand) joins |
| 08:32:29 | | Wohlstand1 is now known as Wohlstand |
| 09:14:55 | | ThreeHM quits [Ping timeout: 248 seconds] |
| 09:18:13 | | ThreeHM (ThreeHeadedMonkey) joins |
| 09:19:11 | | thewinwin88 joins |
| 09:22:55 | | thewinwin8 quits [Ping timeout: 248 seconds] |
| 09:22:58 | | thewinwin88 is now known as thewinwin8 |
| 09:33:14 | <klea> | L9330x3sLFE: → #discard |
| 09:36:58 | <L9330x3sLFE> | Thanks, klea |
| 10:19:57 | | nate1766 joins |
| 10:51:27 | | L9330x3sLFE quits [Ping timeout: 248 seconds] |
| 10:58:52 | | Webuser110307 joins |
| 10:58:56 | | Webuser110307 quits [Client Quit] |
| 11:00:20 | | Bleo18260072271962345522201107 quits [Quit: The Lounge - https://thelounge.chat] |
| 11:03:09 | | Bleo18260072271962345522201107 joins |
| 11:17:35 | | superkuh quits [Ping timeout: 248 seconds] |
| 11:18:22 | | BennyOtt quits [Ping timeout: 268 seconds] |
| 11:21:39 | | BennyOtt (BennyOtt) joins |
| 11:38:35 | | superkuh joins |
| 11:47:26 | | superkuh_ joins |
| 11:49:03 | | superkuh quits [Ping timeout: 248 seconds] |
| 11:50:54 | | superkuh joins |
| 11:52:17 | | superkuh_ quits [Ping timeout: 268 seconds] |
| 12:00:18 | | superkuh quits [Ping timeout: 268 seconds] |
| 12:01:03 | | superkuh joins |
| 12:06:00 | | superkuh_ joins |
| 12:07:11 | | superkuh quits [Ping timeout: 248 seconds] |
| 12:10:45 | | superkuh joins |
| 12:12:31 | | superkuh_ quits [Ping timeout: 248 seconds] |
| 12:15:40 | | superkuh_ joins |
| 12:17:34 | | superkuh quits [Ping timeout: 268 seconds] |
| 12:24:25 | | Webuser719442 joins |
| 12:27:48 | | nine quits [Quit: See ya!] |
| 12:28:01 | | nine joins |
| 12:39:45 | | nate1766 quits [Client Quit] |
| 12:41:52 | <h2ibot> | Cruller edited Mastodon (+237, /* Dead and dying instances */ Added…): https://wiki.archiveteam.org/?diff=63171&oldid=62499 |
| 12:48:15 | | superkuh_ quits [Ping timeout: 248 seconds] |
| 13:09:02 | <justauser> | arkiver: https://schoolwiki.in/ has switched from Anubis to apparently https://www.prophaze.com/platform/bot-mitigation/ ? |
| 13:28:34 | | superkuh joins |
| 13:29:20 | | Webuser094094 joins |
| 13:32:15 | | Webuser094094 quits [Client Quit] |
| 13:49:27 | | superkuh quits [Ping timeout: 268 seconds] |
| 13:49:53 | | superkuh joins |
| 13:59:00 | | nexussfan (nexussfan) joins |
| 14:04:42 | | cyanbox_ joins |
| 14:08:34 | | cyanbox quits [Ping timeout: 268 seconds] |
| 14:21:50 | <cruller|irc> | bokete.jp, which is set to shut down on 2026-09-30, is a service with over 115 million posts and an 18-year history. It appears to require DPoS, but it is excluded from WBM and employs CloudFront's geoblocking: https://globalping.io/?measurement=2sm9XcEih875ENrMc00020lue&collapse=none |
| 14:23:43 | <cruller|irc> | Fortunately, it seems the U.S. is not blocked? |
| 14:27:41 | | superkuh quits [Ping timeout: 268 seconds] |
| 14:28:07 | <@arkiver> | cruller|irc: very nice! |
| 14:28:28 | <@arkiver> | cruller|irc: can you make sure it's on the deathwatch page? |
| 14:28:33 | <@arkiver> | justauser: this evolves so quickly nowadays :/ |
| 14:28:46 | | Island joins |
| 14:29:36 | <cruller|irc> | Since the vast majority of Japanese speakers live in Japan, Japanese-language websites can employ geoblocking without adverse effects :/ |
| 14:29:46 | | superkuh joins |
| 14:48:25 | | cyan_box joins |
| 14:52:21 | | cyanbox_ quits [Ping timeout: 268 seconds] |
| 15:04:40 | | superkuh_ joins |
| 15:05:51 | | superkuh quits [Ping timeout: 248 seconds] |
| 15:15:47 | | leo60228 quits [Ping timeout: 268 seconds] |
| 15:20:43 | | superkuh_ quits [Ping timeout: 268 seconds] |
| 15:20:44 | | superkuh joins |
| 15:27:22 | | Umbire (Umbire) joins |
| 15:38:27 | | Webuser719442 quits [Quit: Ooops, wrong browser tab.] |
| 15:50:31 | | thewinwin82 joins |
| 15:54:23 | | thewinwin8 quits [Ping timeout: 248 seconds] |
| 15:54:26 | | thewinwin82 is now known as thewinwin8 |
| 16:00:47 | | moth3 quits [Ping timeout: 248 seconds] |
| 16:09:26 | | pseudorizer quits [Ping timeout: 268 seconds] |
| 16:14:01 | | hackbug quits [Remote host closed the connection] |
| 16:15:27 | | ScenarioPlanet0 (ScenarioPlanet) joins |
| 16:17:19 | | ScenarioPlanet quits [Ping timeout: 248 seconds] |
| 16:17:20 | | ScenarioPlanet0 is now known as ScenarioPlanet |
| 16:18:54 | | pseudorizer (pseudorizer) joins |
| 16:51:12 | | hackbug joins |
| 17:10:11 | | Cupping12854 quits [Quit: Ping timeout (120 seconds)] |
| 17:10:24 | | Cupping12854 joins |
| 17:14:55 | | khaoohs quits [Read error: Connection reset by peer] |
| 17:25:17 | | Cupping12854 quits [Ping timeout: 268 seconds] |
| 17:28:54 | | eroc1990 joins |
| 17:31:19 | | superkuh_ joins |
| 17:32:48 | | eroc1990 quits [Client Quit] |
| 17:33:01 | | eroc1990 joins |
| 17:33:04 | | L9330x3sLFE joins |
| 17:33:18 | | superkuh quits [Ping timeout: 268 seconds] |
| 17:37:23 | | khaoohs joins |
| 17:40:39 | | eroc1990 quits [Client Quit] |
| 17:42:11 | | Cupping12854 joins |
| 17:55:40 | | msfjarvis quits [Quit: Lurker 1.1.0 (the truth is out there) https://lurker.chat] |
| 17:56:05 | | msfjarvis joins |
| 17:57:21 | | superkuh_ quits [Ping timeout: 268 seconds] |
| 17:58:08 | | superkuh joins |
| 18:18:56 | | ThreeHM quits [Read error: Connection reset by peer] |
| 18:21:43 | | ThreeHM (ThreeHeadedMonkey) joins |
| 18:30:30 | | gatagoto (gatagoto) joins |
| 18:34:55 | | L9330x3sLFE quits [Ping timeout: 248 seconds] |
| 18:37:50 | | L9330x3sLFE joins |
| 18:46:57 | | nate1766 joins |
| 19:22:53 | | hamouda joins |
| 19:24:14 | | Erebus quits [Remote host closed the connection] |
| 19:24:28 | | Erebus (Erebus) joins |
| 19:25:16 | | Erebus quits [Excess Flood] |
| 19:26:21 | | Erebus (Erebus) joins |
| 20:02:31 | | Erebus quits [Client Quit] |
| 20:03:03 | | Erebus (Erebus) joins |
| 20:06:02 | | Erebus quits [Client Quit] |
| 20:06:27 | | Erebus (Erebus) joins |
| 20:09:25 | | Erebus quits [Remote host closed the connection] |
| 20:09:41 | | Erebus (Erebus) joins |
| 20:15:42 | | Erebus quits [Remote host closed the connection] |
| 20:15:57 | | Erebus (Erebus) joins |
| 20:18:31 | <@JAA> | Vito`: I still need to finish sorting and uploading some data from CBS News Radio. The existing items aren't organised in any way, and indeed they aren't in the WBM yet. But yes, marking as offline and saved is good. |
| 20:20:58 | | Erebus quits [Remote host closed the connection] |
| 20:21:14 | | Erebus (Erebus) joins |
| 20:25:33 | <nicolas17> | arkiver: I grabbed a cdx from a random archivebot item |
| 20:25:47 | | Erebus quits [Remote host closed the connection] |
| 20:25:48 | <nicolas17> | 14.7GiB of duplicated data out of 108GiB total |
| 20:26:03 | | Erebus (Erebus) joins |
| 20:26:17 | <nicolas17> | time to import a few hundred cdxs and see how it goes |
| 20:32:05 | <klea> | ouch, I have a folder with almost a terabyte of stuff now, just from getting the IA item that had a tar, not removing the tar file, and splitting each file, and preparing my own CDXs. |
| 20:32:47 | <klea> | Sad, nicolas17's script breaks with the CDXs this tool I'm using makes. |
| 20:32:50 | <klea> | "CDX a b m s k V T g u" |
| 20:35:43 | <klea> | https://transfer.archivete.am/inline/EBKbJ/storify_cdx_matches.txt |
| 20:35:49 | <klea> | Someone asked for those~ |
| 20:36:04 | <klea> | Webuser843256 ^ |
| 20:36:49 | <nicolas17> | damn archivebot works a lot |
| 20:37:21 | <nicolas17> | archiveteam_archivebot_go_20260701*.cdx.gz (a single day) -> 2GB of CDXs |
| 20:37:24 | <nicolas17> | gzipped |
| 20:39:04 | <nicolas17> | 14GB of CDX :| |
| 20:43:01 | <nicolas17> | SQLite DB is gonna grow fast |
| 20:46:20 | <@JAA> | You can easily trade space for compute with this. I've been meaning to do a thorough analysis at some point. |
| 20:46:45 | <nicolas17> | JAA: I'm doing this in a dumb way for an early test |
| 20:46:58 | <nicolas17> | eg. importing every CDX field even though I don't need them for this analysis |
| 20:47:06 | <@JAA> | Mhm |
| 20:47:44 | <@JAA> | My plan was to do a full month. |
| 20:48:19 | <nicolas17> | I also don't *know* how big the payload alone is, or how much the record would shrink if the payload was replaced with a revisit record |
| 20:48:28 | <nicolas17> | but I estimate it with "min(record_size) group by checksum" |
| 20:49:18 | <@JAA> | Or how much the record would grow as a revisit record, for tiny payloads. :-) |
| 20:49:25 | <nicolas17> | true |
| 20:51:41 | <nicolas17> | select sum(s)/1024./1024/1024 from (select (count(*)-1) * min(record_size) s from cdx group by checksum); |
| 20:52:06 | <nicolas17> | checksums that only appear once multiply the record size by zero |
| 20:52:52 | <nicolas17> | that gives me 2677GB dup in 6349GB data?? |
| 20:56:50 | | nate1766 quits [Client Quit] |
| 21:00:23 | <klea> | Oh, that's what nicolas17 is ingesting CDXs for. |
| 21:00:25 | <klea> | nicolas17++ |
| 21:00:26 | <eggdrop> | [karma] 'nicolas17' now has 29 karma! |
| 21:06:21 | <nicolas17> | most repeated checksum 3I42H3S6NNFQ2MSVX7XZKYAYSCX5QBYJ (empty file) |
| 21:06:26 | <nicolas17> | I should probably ignore that one |
| 21:08:21 | <nicolas17> | second most repeated in this dataset F2OCM4GZ5UA7SH6I4QW42H7SEPQSBB75, 634k times, corresponding to https://forum.moparisthebest.com/ |
| 21:08:48 | <nicolas17> | wonder if this is an error message and archival actually failed |
| 21:09:08 | <@JAA> | Or some bogus redirect to the homepage? |
| 21:13:46 | <nicolas17> | not an HTTP redirect but a rewrite |
| 21:13:54 | <nicolas17> | https://forum.moparisthebest.com/user_avatar/forum.moparisthebest.com/zege7/45/576_2.png |
| 21:13:59 | <nicolas17> | it's like every URL returns the same payload |
| 21:14:54 | <nicolas17> | select sum(record_size)/1024./1024/1024 from cdx where checksum='F2OCM4GZ5UA7SH6I4QW42H7SEPQSBB75' -> 2.82GiB |
| 21:14:54 | <@JAA> | Ah, crappy server that doesn't 404, yeah. |
| 21:15:08 | <pokechu22> | Yeah, that was a very broken server |
| 21:15:22 | <pokechu22> | ignores were eventually added but it was doing a lot of that nonsense |
| 21:16:02 | <@JAA> | There will undoubtedly be many weird isolated cases like this in the data. But that's not really what's interesting. |
| 21:16:04 | <pokechu22> | lots of URLs returning the home page, and the home page uses relative links so you get *more* nonsense |
| 21:19:15 | <Yna> | holy smokes not moparisthebest, what a throwback |
| 21:19:17 | | Erebus quits [Remote host closed the connection] |
| 21:19:32 | | Erebus (Erebus) joins |
| 21:27:52 | <nicolas17> | wonder how much storage I'd save by base32-decoding the checksums :P |
| 21:43:52 | <nicolas17> | ok, largest space-wasting duplicates in these 24h |
| 21:44:17 | <nicolas17> | OVIKZB6W3XWZ5HRTZRDVUK6KGT5I4XMD is a 9GB file from mirrors.lolinet.com which we saved 20 times |
| 21:45:14 | <nicolas17> | actually the top 4 worst ones are all from lolinet |
| 21:45:36 | <nicolas17> | different URLs |
| 21:47:18 | <klea> | I feel like what you're doing from only one day might have lots of cases of specific jobs' issues, but kind of showcases how a per-job-dedupe could be useful to implement to avoid disk space issues around AB operators not ignoring things on time. |
| 21:49:22 | <nicolas17> | I need some optimizations before I can add more data |
| 21:50:57 | <nicolas17> | even creating an index on cdx(checksum) didn't seem to help that much |
| 21:52:12 | <nicolas17> | to know how much a *per-job* dedup would help compared to something more global, I'll have to reimport this, because I left out the warc filename :D |
| 21:53:06 | | nine quits [Quit: See ya!] |
| 21:53:20 | | nine joins |
| 21:53:34 | <nicolas17> | for a moment I thought maybe I can do the analysis for each .cdx.gz separately |
| 21:54:01 | <nicolas17> | but each .cdx.gz (ie. each IA item) has stuff from multiple jobs, and jobs can be split across multiple warcs on different items |
| 21:55:33 | | Shard768 quits [Remote host closed the connection] |
| 21:56:31 | | jinn6 quits [Ping timeout: 248 seconds] |
| 21:56:41 | <klea> | IIRC There's a CDX per file in the item? |
| 21:56:59 | <klea> | (of course, you'd have to get the entire job set, which may span multiple items/days/months.) |
| 21:57:13 | <nicolas17> | yeah I guess I could download those CDXs instead |
| 21:57:32 | | jinn6 (jinn6) joins |
| 21:57:42 | <klea> | Or you could download all the WARCs, keep them in your big vault, and make your own cdxs! |
| 21:57:51 | <nicolas17> | this is how I did the download: ia search 'archiveteam_archivebot_go_20260701*' | jq -r '"https://archive.org/download/" + .identifier + "/" + .identifier + ".cdx.gz"' | aria2c -i - |
| 21:58:33 | <nicolas17> | jank everywhere (: |
| 22:01:00 | | Shard768 (Shard) joins |
| 22:06:13 | <nicolas17> | ...I might have gotten rate-limited trying to get the list of smaller CDXs |
| 22:11:15 | | etnguyen03 (etnguyen03) joins |
| 22:12:17 | | Shard768 quits [Client Quit] |
| 22:13:32 | | Shard768 (Shard) joins |
| 22:14:12 | | Shard768 quits [Client Quit] |
| 22:16:47 | | Shard768 (Shard) joins |
| 22:17:54 | <nicolas17> | mirrors.lolinet.com consists of 2836 warcs :D |
| 22:17:54 | | Shard768 quits [Client Quit] |
| 22:20:03 | | Shard768 (Shard) joins |
| 22:27:46 | <nicolas17> | ok |
| 22:28:58 | <nicolas17> | arkiver: https://archive.fart.website/archivebot/viewer/job/20260622131900djo4a this whole job has 20.46TiB, of which 7.67TiB is unique data and 12.80TiB is duplicated |
| 22:32:36 | <nicolas17> | and I *suspect* the large duplicate data is in different warcs, so per-warc dedup wouldn't be enough help |
| 22:37:12 | <nicolas17> | https://archive.fart.website/archivebot/viewer/job/202607010056355mnz4 351GiB, of which 194GiB is unique and 156GiB is duplicated |
| 22:37:22 | <nicolas17> | I'm taking requests for other jobs to look into |
| 22:40:35 | | etnguyen03 quits [Client Quit] |
| 23:04:25 | | etnguyen03 (etnguyen03) joins |
| 23:24:35 | | etnguyen03 quits [Client Quit] |
| 23:25:42 | <pokechu22> | arkiver: https://www.axllent.org/ has a WASM challenge for outdated browser versions/spoofed browser UAs (affects some of AB's user-agents). -u curl works though. (cc Ryz) |
| 23:33:21 | | Hackerpcs quits [Quit: Hackerpcs] |
| 23:34:17 | | Hackerpcs (Hackerpcs) joins |
| 23:36:41 | <nicolas17> | sooo how do revisit records work? can they reference stuff in a different warc? |
| 23:37:12 | <klea> | IIRC yes. |
| 23:37:44 | <klea> | nicolas17: Have some reading: https://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/#revisit and https://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/#example-of-revisit-record |
| 23:38:50 | <nicolas17> | going out of my way to archive 16GB Apple stuff myself with wget-at to prevent duplication instead of just throwing it in archivebot, starts to feel like a waste of time when I see an archivebot job with 12.8TB of dupes :/ |
| 23:40:07 | <nicolas17> | going out of my way to archive 16GB Apple stuff myself with wget-at to prevent duplication instead of just throwing it in archivebot, starts to feel like a waste of time when I see an archivebot job with 12.8TB of dupes :/ |
| 23:40:10 | <nicolas17> | oops |
| 23:40:25 | <klea> | katia bonk++ |
| 23:40:26 | <eggdrop> | [karma] 'katia bonk' now has 4 karma! |
| 23:41:00 | <nicolas17> | why |
| 23:41:13 | <klea> | mirrors.lolinet.com dupes? |
| 23:41:37 | <nicolas17> | ah, not much she could have done about it, ideally what we need is archivebot supporting dedup |
| 23:41:43 | | Wohlstand quits [Remote host closed the connection] |
| 23:41:45 | <klea> | Though, arguably it could be said it's not her issue. |
| 23:41:52 | <klea> | katia bonk-- |
| 23:41:53 | <eggdrop> | [karma] 'katia bonk' now has 3 karma! |
| 23:41:57 | <klea> | ArchiveBot-- |
| 23:41:58 | <eggdrop> | [karma] 'ArchiveBot' now has 4 karma! |
| 23:42:06 | <klea> | deduplication++ |
| 23:42:07 | <eggdrop> | [karma] 'deduplication' now has 8 karma! |
| 23:48:24 | | hamouda quits [Quit: Ooops, wrong browser tab.] |