00:34:45error quits [Remote host closed the connection]
00:35:09error (disinfecttm) joins
01:03:53error quits [Remote host closed the connection]
01:04:11error (disinfecttm) joins
01:14:23MrMcNugg1 quits [Ping timeout: 248 seconds]
01:17:02etnguyen03 quits [Client Quit]
01:20:33etnguyen03 (etnguyen03) joins
01:39:04gatagoto (gatagoto) joins
02:04:02etnguyen03 quits [Remote host closed the connection]
02:08:17chunkynutz6018160 joins
02:10:23chunkynutz601816 quits [Ping timeout: 248 seconds]
02:10:24chunkynutz6018160 is now known as chunkynutz601816
02:25:54ericgallager joins
02:37:57thewinwin89 joins
02:42:13thewinwin8 quits [Ping timeout: 268 seconds]
02:42:16thewinwin89 is now known as thewinwin8
02:43:19L9330x3sLFE joins
02:51:02L9330x3sLFE quits [Client Quit]
02:52:59L9330x3sLFE joins
02:56:20L9330x3sLFE quits [Client Quit]
02:58:58L9330x3sLFE joins
03:03:20<h2ibot>Wickedplayer494 edited TV Time (+0, Is dead): https://wiki.archiveteam.org/?diff=63168&oldid=62959
03:14:16Island quits [Read error: Connection reset by peer]
03:44:57MrMcNuggets (MrMcNuggets) joins
04:07:28DogsRNice quits [Read error: Connection reset by peer]
04:55:54<Pendonym>Any way to archive https://fogcam.org/ every 20s for the image
05:12:43nexussfan quits [Quit: Konversation terminated!]
05:37:05<nicolas17>Pendonym: WBM playback would be hard
05:37:29<nicolas17>they use `"fogcam2.jpg" + new Date()` probably for cache-busting purposes
05:40:59<nicolas17>also since I loaded the page I haven't seen the image ever change
05:41:06<nicolas17>does it actually change every 20s?
05:45:59Cupping12854 quits [Ping timeout: 268 seconds]
05:46:48Cupping12854 joins
05:50:55<nicolas17>ok yeah it updates super irregularly
06:01:24nulldata quits [Ping timeout: 268 seconds]
06:05:13nulldata (nulldata) joins
06:13:33MPThLee quits [Quit: bye]
06:14:59MPThLee (MPThLee) joins
06:29:35benjins3 quits [Ping timeout: 248 seconds]
07:22:11L9330x3sLFE quits [Ping timeout: 268 seconds]
07:28:30L9330x3sLFE joins
07:34:40h|ignition quits [Ping timeout: 272 seconds]
07:36:16h|ignition (h) joins
07:37:27_Apgap joins
07:41:18Boppen_ quits [Ping timeout: 268 seconds]
07:41:38h|ignition quits [Ping timeout: 272 seconds]
07:42:35h|ignition (h) joins
08:02:27simon816 quits [Quit: ZNC 1.10.2 - https://znc.in]
08:04:40BennyOtt quits [Quit: Bye]
08:04:44gatagoto quits [Ping timeout: 268 seconds]
08:04:47<@arkiver>imer: yep, we just went with the #// target, all good! :)
08:06:09<h2ibot>Rexma edited List of websites excluded from the Wayback Machine (+21, add xypb.net): https://wiki.archiveteam.org/?diff=63169&oldid=63154
08:07:54BennyOtt (BennyOtt) joins
08:10:43simon816 (simon816) joins
08:16:10<h2ibot>KleaBot edited List of websites excluded from the Wayback Machine (+0, Reordered websites and/or updated count.): https://wiki.archiveteam.org/?diff=63170&oldid=63169
08:17:27<L9330x3sLFE>Is there a difference between Discrub and DiscordChatExporter when it comes to archiving? How about measures against rate limiting?
08:29:22benjins3 joins
08:30:07Wohlstand1 (Wohlstand) joins
08:32:29Wohlstand1 is now known as Wohlstand
09:14:55ThreeHM quits [Ping timeout: 248 seconds]
09:18:13ThreeHM (ThreeHeadedMonkey) joins
09:19:11thewinwin88 joins
09:22:55thewinwin8 quits [Ping timeout: 248 seconds]
09:22:58thewinwin88 is now known as thewinwin8
09:33:14<klea>L9330x3sLFE: → #discard
09:36:58<L9330x3sLFE>Thanks, klea
10:19:57nate1766 joins
10:51:27L9330x3sLFE quits [Ping timeout: 248 seconds]
10:58:52Webuser110307 joins
10:58:56Webuser110307 quits [Client Quit]
11:00:20Bleo18260072271962345522201107 quits [Quit: The Lounge - https://thelounge.chat]
11:03:09Bleo18260072271962345522201107 joins
11:17:35superkuh quits [Ping timeout: 248 seconds]
11:18:22BennyOtt quits [Ping timeout: 268 seconds]
11:21:39BennyOtt (BennyOtt) joins
11:38:35superkuh joins
11:47:26superkuh_ joins
11:49:03superkuh quits [Ping timeout: 248 seconds]
11:50:54superkuh joins
11:52:17superkuh_ quits [Ping timeout: 268 seconds]
12:00:18superkuh quits [Ping timeout: 268 seconds]
12:01:03superkuh joins
12:06:00superkuh_ joins
12:07:11superkuh quits [Ping timeout: 248 seconds]
12:10:45superkuh joins
12:12:31superkuh_ quits [Ping timeout: 248 seconds]
12:15:40superkuh_ joins
12:17:34superkuh quits [Ping timeout: 268 seconds]
12:24:25Webuser719442 joins
12:27:48nine quits [Quit: See ya!]
12:28:01nine joins
12:39:45nate1766 quits [Client Quit]
12:41:52<h2ibot>Cruller edited Mastodon (+237, /* Dead and dying instances */ Added…): https://wiki.archiveteam.org/?diff=63171&oldid=62499
12:48:15superkuh_ quits [Ping timeout: 248 seconds]
13:09:02<justauser>arkiver: https://schoolwiki.in/ has switched from Anubis to apparently https://www.prophaze.com/platform/bot-mitigation/ ?
13:28:34superkuh joins
13:29:20Webuser094094 joins
13:32:15Webuser094094 quits [Client Quit]
13:49:27superkuh quits [Ping timeout: 268 seconds]
13:49:53superkuh joins
13:59:00nexussfan (nexussfan) joins
14:04:42cyanbox_ joins
14:08:34cyanbox quits [Ping timeout: 268 seconds]
14:21:50<cruller|irc>bokete.jp, which is set to shut down on 2026-09-30, is a service with over 115 million posts and an 18-year history. It appears to require DPoS, but it is excluded from WBM and employs CloudFront's geoblocking: https://globalping.io/?measurement=2sm9XcEih875ENrMc00020lue&collapse=none
14:23:43<cruller|irc>Fortunately, it seems the U.S. is not blocked?
14:27:41superkuh quits [Ping timeout: 268 seconds]
14:28:07<@arkiver>cruller|irc: very nice!
14:28:28<@arkiver>cruller|irc: can you make sure it's on the deathwatch page?
14:28:33<@arkiver>justauser: this evolves so quickly nowadays :/
14:28:46Island joins
14:29:36<cruller|irc>Since the vast majority of Japanese speakers live in Japan, Japanese-language websites can employ geoblocking without adverse effects :/
14:29:46superkuh joins
14:48:25cyan_box joins
14:52:21cyanbox_ quits [Ping timeout: 268 seconds]
15:04:40superkuh_ joins
15:05:51superkuh quits [Ping timeout: 248 seconds]
15:15:47leo60228 quits [Ping timeout: 268 seconds]
15:20:43superkuh_ quits [Ping timeout: 268 seconds]
15:20:44superkuh joins
15:27:22Umbire (Umbire) joins
15:38:27Webuser719442 quits [Quit: Ooops, wrong browser tab.]
15:50:31thewinwin82 joins
15:54:23thewinwin8 quits [Ping timeout: 248 seconds]
15:54:26thewinwin82 is now known as thewinwin8
16:00:47moth3 quits [Ping timeout: 248 seconds]
16:09:26pseudorizer quits [Ping timeout: 268 seconds]
16:14:01hackbug quits [Remote host closed the connection]
16:15:27ScenarioPlanet0 (ScenarioPlanet) joins
16:17:19ScenarioPlanet quits [Ping timeout: 248 seconds]
16:17:20ScenarioPlanet0 is now known as ScenarioPlanet
16:18:54pseudorizer (pseudorizer) joins
16:51:12hackbug joins
17:10:11Cupping12854 quits [Quit: Ping timeout (120 seconds)]
17:10:24Cupping12854 joins
17:14:55khaoohs quits [Read error: Connection reset by peer]
17:25:17Cupping12854 quits [Ping timeout: 268 seconds]
17:28:54eroc1990 joins
17:31:19superkuh_ joins
17:32:48eroc1990 quits [Client Quit]
17:33:01eroc1990 joins
17:33:04L9330x3sLFE joins
17:33:18superkuh quits [Ping timeout: 268 seconds]
17:37:23khaoohs joins
17:40:39eroc1990 quits [Client Quit]
17:42:11Cupping12854 joins
17:55:40msfjarvis quits [Quit: Lurker 1.1.0 (the truth is out there) https://lurker.chat]
17:56:05msfjarvis joins
17:57:21superkuh_ quits [Ping timeout: 268 seconds]
17:58:08superkuh joins
18:18:56ThreeHM quits [Read error: Connection reset by peer]
18:21:43ThreeHM (ThreeHeadedMonkey) joins
18:30:30gatagoto (gatagoto) joins
18:34:55L9330x3sLFE quits [Ping timeout: 248 seconds]
18:37:50L9330x3sLFE joins
18:46:57nate1766 joins
19:22:53hamouda joins
19:24:14Erebus quits [Remote host closed the connection]
19:24:28Erebus (Erebus) joins
19:25:16Erebus quits [Excess Flood]
19:26:21Erebus (Erebus) joins
20:02:31Erebus quits [Client Quit]
20:03:03Erebus (Erebus) joins
20:06:02Erebus quits [Client Quit]
20:06:27Erebus (Erebus) joins
20:09:25Erebus quits [Remote host closed the connection]
20:09:41Erebus (Erebus) joins
20:15:42Erebus quits [Remote host closed the connection]
20:15:57Erebus (Erebus) joins
20:18:31<@JAA>Vito`: I still need to finish sorting and uploading some data from CBS News Radio. The existing items aren't organised in any way, and indeed they aren't in the WBM yet. But yes, marking as offline and saved is good.
20:20:58Erebus quits [Remote host closed the connection]
20:21:14Erebus (Erebus) joins
20:25:33<nicolas17>arkiver: I grabbed a cdx from a random archivebot item
20:25:47Erebus quits [Remote host closed the connection]
20:25:48<nicolas17>14.7GiB of duplicated data out of 108GiB total
20:26:03Erebus (Erebus) joins
20:26:17<nicolas17>time to import a few hundred cdxs and see how it goes
20:32:05<klea>ouch, I have a folder with almost a terabyte of stuff now, just from getting the IA item that had a tar, not removing the tar file, and splitting each file, and preparing my own CDXs.
20:32:47<klea>Sad, nicolas17's script breaks with the CDXs this tool I'm using makes.
20:32:50<klea>"CDX a b m s k V T g u"
20:35:43<klea>https://transfer.archivete.am/inline/EBKbJ/storify_cdx_matches.txt
20:35:49<klea>Someone asked for those~
20:36:04<klea>Webuser843256 ^
20:36:49<nicolas17>damn archivebot works a lot
20:37:21<nicolas17>archiveteam_archivebot_go_20260701*.cdx.gz (a single day) -> 2GB of CDXs
20:37:24<nicolas17>gzipped
20:39:04<nicolas17>14GB of CDX :|
20:43:01<nicolas17>SQLite DB is gonna grow fast
20:46:20<@JAA>You can easily trade space for compute with this. I've been meaning to do a thorough analysis at some point.
20:46:45<nicolas17>JAA: I'm doing this in a dumb way for an early test
20:46:58<nicolas17>eg. importing every CDX field even though I don't need them for this analysis
20:47:06<@JAA>Mhm
20:47:44<@JAA>My plan was to do a full month.
20:48:19<nicolas17>I also don't *know* how big the payload alone is, or how much the record would shrink if the payload was replaced with a revisit record
20:48:28<nicolas17>but I estimate it with "min(record_size) group by checksum"
20:49:18<@JAA>Or how much the record would grow as a revisit record, for tiny payloads. :-)
20:49:25<nicolas17>true
20:51:41<nicolas17>select sum(s)/1024./1024/1024 from (select (count(*)-1) * min(record_size) s from cdx group by checksum);
20:52:06<nicolas17>checksums that only appear once multiply the record size by zero
20:52:52<nicolas17>that gives me 2677GB dup in 6349GB data??
20:56:50nate1766 quits [Client Quit]
21:00:23<klea>Oh, that's what nicolas17 is ingesting CDXs for.
21:00:25<klea>nicolas17++
21:00:26<eggdrop>[karma] 'nicolas17' now has 29 karma!
21:06:21<nicolas17>most repeated checksum 3I42H3S6NNFQ2MSVX7XZKYAYSCX5QBYJ (empty file)
21:06:26<nicolas17>I should probably ignore that one
21:08:21<nicolas17>second most repeated in this dataset F2OCM4GZ5UA7SH6I4QW42H7SEPQSBB75, 634k times, corresponding to https://forum.moparisthebest.com/
21:08:48<nicolas17>wonder if this is an error message and archival actually failed
21:09:08<@JAA>Or some bogus redirect to the homepage?
21:13:46<nicolas17>not an HTTP redirect but a rewrite
21:13:54<nicolas17>https://forum.moparisthebest.com/user_avatar/forum.moparisthebest.com/zege7/45/576_2.png
21:13:59<nicolas17>it's like every URL returns the same payload
21:14:54<nicolas17>select sum(record_size)/1024./1024/1024 from cdx where checksum='F2OCM4GZ5UA7SH6I4QW42H7SEPQSBB75' -> 2.82GiB
21:14:54<@JAA>Ah, crappy server that doesn't 404, yeah.
21:15:08<pokechu22>Yeah, that was a very broken server
21:15:22<pokechu22>ignores were eventually added but it was doing a lot of that nonsense
21:16:02<@JAA>There will undoubtedly be many weird isolated cases like this in the data. But that's not really what's interesting.
21:16:04<pokechu22>lots of URLs returning the home page, and the home page uses relative links so you get *more* nonsense
21:19:15<Yna>holy smokes not moparisthebest, what a throwback
21:19:17Erebus quits [Remote host closed the connection]
21:19:32Erebus (Erebus) joins
21:27:52<nicolas17>wonder how much storage I'd save by base32-decoding the checksums :P
21:43:52<nicolas17>ok, largest space-wasting duplicates in these 24h
21:44:17<nicolas17>OVIKZB6W3XWZ5HRTZRDVUK6KGT5I4XMD is a 9GB file from mirrors.lolinet.com which we saved 20 times
21:45:14<nicolas17>actually the top 4 worst ones are all from lolinet
21:45:36<nicolas17>different URLs
21:47:18<klea>I feel like what you're doing from only one day might have lots of cases of specific jobs' issues, but kind of showcases how a per-job-dedupe could be useful to implement to avoid disk space issues around AB operators not ignoring things on time.
21:49:22<nicolas17>I need some optimizations before I can add more data
21:50:57<nicolas17>even creating an index on cdx(checksum) didn't seem to help that much
21:52:12<nicolas17>to know how much a *per-job* dedup would help compared to something more global, I'll have to reimport this, because I left out the warc filename :D
21:53:06nine quits [Quit: See ya!]
21:53:20nine joins
21:53:34<nicolas17>for a moment I thought maybe I can do the analysis for each .cdx.gz separately
21:54:01<nicolas17>but each .cdx.gz (ie. each IA item) has stuff from multiple jobs, and jobs can be split across multiple warcs on different items
21:55:33Shard768 quits [Remote host closed the connection]
21:56:31jinn6 quits [Ping timeout: 248 seconds]
21:56:41<klea>IIRC There's a CDX per file in the item?
21:56:59<klea>(of course, you'd have to get the entire job set, which may span multiple items/days/months.)
21:57:13<nicolas17>yeah I guess I could download those CDXs instead
21:57:32jinn6 (jinn6) joins
21:57:42<klea>Or you could download all the WARCs, keep them in your big vault, and make your own cdxs!
21:57:51<nicolas17>this is how I did the download: ia search 'archiveteam_archivebot_go_20260701*' | jq -r '"https://archive.org/download/" + .identifier + "/" + .identifier + ".cdx.gz"' | aria2c -i -
21:58:33<nicolas17>jank everywhere (:
22:01:00Shard768 (Shard) joins
22:06:13<nicolas17>...I might have gotten rate-limited trying to get the list of smaller CDXs
22:11:15etnguyen03 (etnguyen03) joins
22:12:17Shard768 quits [Client Quit]
22:13:32Shard768 (Shard) joins
22:14:12Shard768 quits [Client Quit]
22:16:47Shard768 (Shard) joins
22:17:54<nicolas17>mirrors.lolinet.com consists of 2836 warcs :D
22:17:54Shard768 quits [Client Quit]
22:20:03Shard768 (Shard) joins
22:27:46<nicolas17>ok
22:28:58<nicolas17>arkiver: https://archive.fart.website/archivebot/viewer/job/20260622131900djo4a this whole job has 20.46TiB, of which 7.67TiB is unique data and 12.80TiB is duplicated
22:32:36<nicolas17>and I *suspect* the large duplicate data is in different warcs, so per-warc dedup wouldn't be enough help
22:37:12<nicolas17>https://archive.fart.website/archivebot/viewer/job/202607010056355mnz4 351GiB, of which 194GiB is unique and 156GiB is duplicated
22:37:22<nicolas17>I'm taking requests for other jobs to look into
22:40:35etnguyen03 quits [Client Quit]
23:04:25etnguyen03 (etnguyen03) joins
23:24:35etnguyen03 quits [Client Quit]
23:25:42<pokechu22>arkiver: https://www.axllent.org/ has a WASM challenge for outdated browser versions/spoofed browser UAs (affects some of AB's user-agents). -u curl works though. (cc Ryz)
23:33:21Hackerpcs quits [Quit: Hackerpcs]
23:34:17Hackerpcs (Hackerpcs) joins
23:36:41<nicolas17>sooo how do revisit records work? can they reference stuff in a different warc?
23:37:12<klea>IIRC yes.
23:37:44<klea>nicolas17: Have some reading: https://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/#revisit and https://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/#example-of-revisit-record
23:38:50<nicolas17>going out of my way to archive 16GB Apple stuff myself with wget-at to prevent duplication instead of just throwing it in archivebot, starts to feel like a waste of time when I see an archivebot job with 12.8TB of dupes :/
23:40:07<nicolas17>going out of my way to archive 16GB Apple stuff myself with wget-at to prevent duplication instead of just throwing it in archivebot, starts to feel like a waste of time when I see an archivebot job with 12.8TB of dupes :/
23:40:10<nicolas17>oops
23:40:25<klea>katia bonk++
23:40:26<eggdrop>[karma] 'katia bonk' now has 4 karma!
23:41:00<nicolas17>why
23:41:13<klea>mirrors.lolinet.com dupes?
23:41:37<nicolas17>ah, not much she could have done about it, ideally what we need is archivebot supporting dedup
23:41:43Wohlstand quits [Remote host closed the connection]
23:41:45<klea>Though, arguably it could be said it's not her issue.
23:41:52<klea>katia bonk--
23:41:53<eggdrop>[karma] 'katia bonk' now has 3 karma!
23:41:57<klea>ArchiveBot--
23:41:58<eggdrop>[karma] 'ArchiveBot' now has 4 karma!
23:42:06<klea>deduplication++
23:42:07<eggdrop>[karma] 'deduplication' now has 8 karma!
23:48:24hamouda quits [Quit: Ooops, wrong browser tab.]