00:04:18etnguyen03 (etnguyen03) joins
00:15:24MrMcNuggets (MrMcNuggets) joins
00:31:53^ quits [Read error: Connection reset by peer]
00:33:22^ (^) joins
00:37:04MrMcNuggets quits [Client Quit]
00:44:13BearFortress quits [Ping timeout: 268 seconds]
01:05:48valdikss quits [Ping timeout: 268 seconds]
01:06:37<klea>arkiver: https://nopecha.com/demo/ might be interesting to support all those, preferably with local models, since we don't have money to call out APIs.
01:08:10valdikss joins
01:08:19MrMcNuggets (MrMcNuggets) joins
01:09:14MrMcNuggets quits [Read error: Connection reset by peer]
01:09:43MrMcNuggets (MrMcNuggets) joins
01:39:49thalia quits [Client Quit]
01:56:33Guest58 joins
02:01:54Island quits [Read error: Connection reset by peer]
02:10:33<nicolas17>who admins our wikis?
02:11:30<klea>By admins you mean shell access or wiki-level admin access?
02:12:13<klea>If shell, Jason Scott and maybe ark has for the fileformats wiki, and for sure JAA has for wiki.a.o internetarchive.a.o, and maybe others have access to that shared hosting.
02:12:31<nicolas17>https://lists.wikimedia.org/hyperkitty/list/mediawiki-announce@lists.wikimedia.org/thread/6JITMBXZ6TXX3MW32BN33Z3JNHPJGJZS/
02:13:17<nicolas17>I think if import/importupload are already restricted to trusted users, we should be fine
02:13:37<@JAA>Ack
02:14:48<klea>nicolas17++
02:14:49<eggdrop>[karma] 'nicolas17' now has 28 karma!
02:16:06BennyOtt quits [Ping timeout: 268 seconds]
02:18:01Webuser374065 joins
02:18:54Webuser374065 quits [Client Quit]
02:33:47BearFortress joins
02:35:36MrMcNugg1 (MrMcNuggets) joins
02:39:32MrMcNuggets quits [Ping timeout: 268 seconds]
02:41:32etnguyen03 quits [Remote host closed the connection]
03:18:43^ quits [Remote host closed the connection]
03:18:51^ (^) joins
03:55:55<c3manu>JAA, pokechu22, arkiver: https://forums.rpgmakerweb.com/ just got banned (refused connections) a second time in #archivebot, at around the same downloaded size/number of URLs as first time. that will require some kind of workaround or even DPoS to be saved before it gets deleted in december
03:55:57DogsRNice quits [Read error: Connection reset by peer]
03:56:21<c3manu>size-wise https://gelbooru.com/ might also require DPoS (see https://xcancel.com/gelbooru/status/2069889788908331451 )
04:57:02didyousayboop joins
05:00:15<didyousayboop>Have folks here already seen that long running gay forum Data Lounge is shutting down on July 31, 2026? Closing announcement: https://datalounge.com/closing
05:01:10<didyousayboop>Reddit post for context: https://www.reddit.com/r/DataHoarder/comments/1udgnft/looking_for_a_hoarder_who_can_archive_a_piece_of/
05:01:43<pokechu22>I don't think we have heard about that
05:02:29<didyousayboop>I see it’s already been added to Deathwatch: https://wiki.archiveteam.org/?title=Deathwatch#2026-07
05:02:37<didyousayboop>Hi pokechu22, glad to see you as always
05:03:03<didyousayboop>Would it be possible to unleash ArchiveBot on datalounge.com?
05:04:17<pokechu22>Checking... Looks like we previously ran it in 2017.
05:04:24<didyousayboop>The closing announcement says the site has been up for 31 years and has over 36 million posts, but I can’t tell if that’s a joke
05:04:44<didyousayboop>Oh, that’s awesome! Particularly as I heard a lot of pre-2022 content was taken down
05:11:32<pokechu22>Hmm, I don't recognize the forum software on there currently, and they don't seem to have a sitemap
05:13:20<pokechu22>Looks like https://datalounge.com/thread/36555880 redirects to https://datalounge.com/thread/36555880-european-heat-wave!
05:13:29nexussfan quits [Remote host closed the connection]
05:14:52hexa- quits [Quit: Disconnected]
05:14:56<didyousayboop>It might use its own bespoke software? https://en.wikipedia.org/wiki/DataLounge
05:15:42hexa- (hexa-) joins
05:15:43<didyousayboop>Apparently it’s owned and operated by a company that designs websites: https://en.wikipedia.org/wiki/Mediapolis_(company)
05:15:53<pokechu22>... but https://datalounge.com/thread/36555080 is https://datalounge.com/thread/36554908-datalounge-is-dead-to-me!!-part-2#36555080. So it's not 36 million threads, at least... but I'm not sure if there's a good way to discover threads without bruteforcing all of the post IDs
05:16:31<pokechu22>https://datalounge.com/thread/36554908-datalounge-is-dead-to-me!!-part-2 doesn't actually seem to link back to 36555080 though
05:28:42<didyousayboop>I find it hard to parse the things people say about this site. There seem to be years and years of built up inside jokes and memes and “lore”. So, half the time, I have no idea what people are talking about. Or whether they’re being serious or joking around.
05:29:49<didyousayboop>If it was a serious remark, the admin might have meant posts as in comments rather than threads.
05:30:03<didyousayboop>Comments as in replies.
05:30:04<pokechu22>I started an archivebot job, but I don't know if it's actually going to find old threads
05:30:13<didyousayboop>Hmm, okay! Thank you!
05:30:34<pokechu22>Yeah, just based on the number there it seems like they probably have 36 million posts including replies
05:37:09<didyousayboop>Makes sense
05:42:41<pokechu22>JAA: https://datalounge.com/ might be a good candidate for qwarc bruteforcing thread IDs, assuming there isn't any rate-limiting?
05:44:59<pokechu22>The AB job hasn't ran into any rate-limiting so far but it's also picking up a ton of outlinks
05:47:42didyousayboop quits [Client Quit]
05:47:48didyousayboop (didyousayboop) joins
05:51:19<pokechu22>The formatting info on https://www.datalounge.com/faq doesn't indicate support for inline images or anything, just links.
05:57:23<pokechu22>I tried a second archivebot job with https://www.datalounge.com/ as --no-offsite and at con=9, d=0. No rate-limiting, but some older threads are "expired" or give 404s (e.g. https://www.datalounge.com/thread/11204830-let-s-pretend-we-re-umpy-threads / https://www.datalounge.com/thread/11204830, https://www.datalounge.com/thread/21874700). Seems like around before 20M is where
05:57:25<pokechu22>things are expired
05:57:40<pokechu22>qwarc almost certainly is a good choice for that
07:09:38traxys quits [Ping timeout: 268 seconds]
07:28:08wessel1512 quits [Ping timeout: 268 seconds]
08:23:39traxys (traxys) joins
08:29:35traxys quits [Ping timeout: 248 seconds]
08:30:01traxys (traxys) joins
08:33:33traxys8 (traxys) joins
08:35:58traxys quits [Ping timeout: 268 seconds]
08:35:58traxys8 is now known as traxys
08:43:42amphitryon_ joins
08:47:12amphitryon quits [Ping timeout: 248 seconds]
08:51:50BennyOtt (BennyOtt) joins
08:54:10Webuser1910735 joins
08:55:02BennyOtt quits [Client Quit]
08:55:37BennyOtt (BennyOtt) joins
08:56:03<Webuser1910735>Hi, I'm a researcher looking for pre-2023 geolocated Twitter data — full tweet text, user IDs, and GPS coordinates. I found the Twitter Stream Grab on the Internet Archive but it's access-restricted. Does anyone know of other ways to access this kind of data for research?
09:03:20systwo (systwi) joins
09:03:52BennyOtt quits [Client Quit]
09:04:41Webuser1910735 quits [Client Quit]
09:04:53traxys1 (traxys) joins
09:06:10<hexagonwin>arkiver: yes, there are many. for example this https://blice.co.kr/web/detail.kt?novelId=113318
09:06:11systwi quits [Ping timeout: 268 seconds]
09:06:45BennyOtt (BennyOtt) joins
09:06:50<hexagonwin>the ones you can easily find on main page is usually made by professional authors. but this platform seems to host many user-created contents and those seem to be mostly accessible
09:07:20<hexagonwin>(so we should just scan all possible IDs, there's no other way to discover them)
09:07:50<hexagonwin>i found that specific example from this post written on its discussion board btw https://blice.co.kr/web/documents/detail.kt?documentSeq=36196&start=1
09:08:25<h2ibot>Brad edited Deathwatch (+282, Added ExpertCare): https://wiki.archiveteam.org/?diff=62839&oldid=62825
09:08:39traxys quits [Ping timeout: 268 seconds]
09:08:39traxys1 is now known as traxys
09:09:11<hexagonwin>arkiver: and for joongang and jtbc, we don't know yet. but their financial situation seems to be very bad.
09:09:32arch quits [Remote host closed the connection]
09:09:56arch (arch) joins
09:10:42<hexagonwin>(btw, i'm quite busy until about july 3rd so i can't check here often, sorry)
09:11:46<hexagonwin>klea: i tried nopecha on my personal project earlier this year, it works very well (it's the cheapest one i found that works with hcaptcha)
09:13:33<hexagonwin>i used automated chromium browser (pydoll) and their browser addon.. free usage without account (100/day) can be used infinitely by rotating IPs
09:14:17<hexagonwin>i ended up paying for the $20/month plan (20000/day) tho. the solving speed seemed identical for all free, starter($5) and basic($20)
09:19:31evergreen4 joins
09:22:50evergreen quits [Ping timeout: 268 seconds]
09:22:50evergreen4 is now known as evergreen
09:29:12thewinwin89 joins
09:33:19thewinwin8 quits [Ping timeout: 268 seconds]
09:33:22thewinwin89 is now known as thewinwin8
10:04:58Webuser976790 joins
10:05:05Webuser976790 quits [Client Quit]
11:00:21Bleo18260072271962345522201107 quits [Quit: The Lounge - https://thelounge.chat]
11:03:11Bleo18260072271962345522201107 joins
11:25:17DLoader_ is now known as DLoader
11:27:01<klea>hexagonwin: I meant to reimplement it locally, since we're not full of money.
11:27:29<klea>The issue is that it's starting to look like needing to run headless browsers rather than Wget-AT then.
11:43:46<cruller>Is the primary purpose to (by)pass anti-bot softwares rather than to execute JavaScript?
11:45:16APOLLO03a joins
11:45:17APOLLO03 quits [Ping timeout: 268 seconds]
11:45:30<klea>I mean, technically executing JS we'd possibly get more complete captures, but that's likely to be CPU intensive, and given most pages prompt you to not be a bot due to the LLMs, I think possibly yes.
11:56:03<cruller>I see. I don't know how high the CPU load is for browsing without JS, but it's likely far higher than for wget-at.
12:01:45symmret (symmret) joins
12:04:15pabs quits [Read error: Connection reset by peer]
12:05:20pabs (pabs) joins
12:06:47<cruller>Incidentally, brozzler has a headful option, which is likely effective for bypassing anti-bot measures. However, the extent of its effectiveness and the CPU load are unknown too.
12:14:51<klea>We could try to do some tests :p
12:18:23<cruller>Reasonably quantitative experiments might be necessary.
12:24:42<cruller>https://pixel.withgoogle.com/ was shut down without notice: https://jetstream.blog/2026/06/26/pixel-simulator-shut-down/
12:36:57<h2ibot>Cruller edited Deathwatch/Dead as a Doornail (+221, /* 2026 */ Add Pixel phone Simulator): https://wiki.archiveteam.org/?diff=62840&oldid=62799
12:52:51<hexagonwin>klea i mean that would be good, but i don't think we can realistically implement a reliable hcaptcha solver/bypass..
12:53:27<hexagonwin>there's already a great recaptcha solver based on audio though ('buster' on FF AMO)
12:53:35unknownsrc quits [Read error: Connection reset by peer]
12:54:02unknownsrc (unknownsrc) joins
13:00:37parallaxstellar quits [Remote host closed the connection]
13:00:44parallaxstellar (parallaxstellar) joins
13:15:25<@arkiver>hexagonwin: alright i'll have a close look at joongang and jtbc too then
13:16:14<@arkiver>c3manu: can https://forums.rpgmakerweb.com/ be archived with a "simple" crawl? if yes, we can put it in #Y . is there a deadline?
13:17:29<masterx244|m>arkiver: 11th december is the deadline, see the announcement post. JAA waited until the closure a week ago because from that point to the december deadline the site is readonly so no posts shifting around possible
13:18:17<@arkiver>alright, we'll launch it on #Y
13:28:44<TheTechRobo>cruller: FWIW mnbot uses the headful option of Brozzler
13:30:33<cruller>That sounds good!
13:48:51<justauser>I mentioned DataLounge and did a few checks. Doesn't seem easily spiderable (right now?). AB job ran in 2017 and was of nontrivial size.
13:52:32<justauser>Much of the Pixel code seems to still be on the page?
14:04:51nexussfan (nexussfan) joins
14:09:14thewinwin85 joins
14:09:35thewinwin8 quits [Ping timeout: 268 seconds]
14:09:40thewinwin85 is now known as thewinwin8
14:30:16<cruller>justauser: That seems to be the case, but it's a bit scripty, so archiving is difficult. What I came up with is to override or block the resources that causes the redirection appropriately, perform browsing and interactions, and then do !ao the discovered URLs.
14:31:53<justauser>Can't find what exactly trigger the redirect.
14:32:32<justauser>I can Devtools-block "support.google.com", preventing it from happening, but the page still won't load properly.
14:33:27<cruller>Me too. Why is the modern web so complicated?
14:34:40<justauser>Why would simulating a phone be easy?
14:34:52<justauser>It's no ordinary page - it's an application.
14:36:58scotrod21 joins
14:38:39scotrod2 quits [Ping timeout: 248 seconds]
14:38:40scotrod21 is now known as scotrod2
14:43:31<cruller>Well, it is certainly inevitable that the main function is scripty.
14:44:00<cruller>However, it wouldn't be surprising if the code for performing the redirection were even simpler.
14:49:20Island joins
14:53:50<skankhunt42>hey! I am running a few containers for kakaotv right now. Unfortunately, I only have about 20G disk space. It looks like the temp files don't get deleted after upload. After restarting the container, space was free again. I was thinking about checking avail diskspace with a systemd timer and just restart the stack, is there a better option?
15:04:28<justauser>A beautiful thing that I completely fail to understand: https://mrchicken.nexussfan.cz/ .
15:04:28<justauser>An antibot that consists, *entirely*, of an overlay covering the site for 5 seconds and not disappearing with 1% probability.
15:04:28<justauser>Disable JS - and everything works fine. /cc arkiver just in case?
15:05:56Dango360 quits [Quit: The Lounge - https://thelounge.chat]
15:06:22<TheTechRobo>skankhunt42: That sounds like a bug, which files are not getting deleted?
15:07:42<skankhunt42>Argh, I should've saved the logs. Unsure, it was stuff in /grab/data (iirc). Given the size of about 15GB, I just assumed that it was old stuff already uploaded. Now it just works again. I will keep an eye on it :) Sorry!
15:08:47nexussfan quits [Ping timeout: 268 seconds]
15:09:54Dango360 (Dango360) joins
15:12:35<justauser>Perhaps it was a huge video in process of being up/down-loaded?
15:14:56<skankhunt42>might be, I increased storage to 50gb just to be sure :)
15:15:24<cruller>https://mrchicken.nexussfan.cz/antibot.js timeSpeed.textContent = Math.floor(Math.random() * 1000); lol
15:27:17n9nes quits [Ping timeout: 268 seconds]
15:27:29n9nes joins
15:38:02<cruller>If AT were to create a tool like https://github.com/TheGP/untidetect-tools#detection-tests for its own use, it would be a good idea to display the results not only as text but also as QR codes.
15:43:43n9nes quits [Ping timeout: 248 seconds]
15:43:49n9nes joins
15:50:06ThreeHM quits [Ping timeout: 268 seconds]
15:51:31<Dango360>`if (Math.floor(Math.random() * 100) == 50)`
16:03:30ThreeHM (ThreeHeadedMonkey) joins
16:13:59dabs joins
16:14:50dabs quits [Remote host closed the connection]
16:15:09dabs joins
16:16:17<@JAA>pokechu22, didyousayboop: Yeah, sounds like a good fit for qwarc. Have you seen any threads with pagination?
16:30:59hackbug77 joins
16:34:30hackbug quits [Ping timeout: 268 seconds]
16:41:48nine quits [Quit: See ya!]
16:42:02nine joins
16:43:58Webuser600015 joins
16:46:34<Webuser600015>https://xcancel.com/gelbooru/status/2069889788908331451
16:46:34<Webuser600015>is it emergency?
16:48:01<justauser>We know about this one.
16:48:26<Webuser600015>oh...
16:55:19Webuser478445 joins
16:57:22<pokechu22>JAA: I haven't; https://www.datalounge.com/thread/36553611-a-note-from-muriel and https://datalounge.com/thread/36554908-datalounge-is-dead-to-me!!-part-2 are both 600+ posts without pagination. https://ab2f.archivingyoursh.it/2bgzyucn32ywsd847qr6nmp66.jsonl. Looking at https://datalounge.com/thread/36368243-the-iran-war-main-open-thread +
16:57:24<pokechu22>https://datalounge.com/thread/36390737-the-iran-war-%E2%80%94-main-open-thread-vol.-2 + https://datalounge.com/thread/36412766-the-iran-war-main-open-thread-part-3 it seems like they usually get split at 600 but I'm not sure if anything enforces that
16:58:58<pokechu22>https://www.datalounge.com/thread/36276187-savannah-guthrie-s-mother-is-missing-part-4-the-search-for-mother-guthrie-continues... also splits at 600 (just looking at threads with part in the URL)
17:10:03nicolas17_ (nicolas17) joins
17:11:30nicolas17 quits [Ping timeout: 268 seconds]
17:12:47unknownsrc quits [Ping timeout: 248 seconds]
17:13:33nicolas17_ is now known as nicolas17
17:14:39Webuser600015 quits [Client Quit]
17:21:55unknownsrc (unknownsrc) joins
17:23:14<justauser>arkiver: A custom bot blocker at https://git.nolog.cz/ /cc #:3
17:32:44<@JAA>pokechu22: Thanks, that makes it easier.
17:33:52<@JAA>Looks like they serve everything on datalounge.com and www.datalounge.com without redirects and funnily treat each other as offsite.
17:34:00<@JAA>www seems to be the canonical one.
18:10:05n9nes quits [Ping timeout: 268 seconds]
18:11:40n9nes joins
18:15:40dabs quits [Read error: Connection reset by peer]
18:15:46Webuser615183 joins
18:16:17Webuser615183 quits [Client Quit]
18:19:52Webuser143737 joins
18:21:22Shard76 quits [Quit: Im doing something rq. Il brb]
18:21:45Shard76 (Shard) joins
18:22:39n9nes quits [Ping timeout: 248 seconds]
18:25:47Webuser698633 joins
18:25:51Webuser698633 quits [Client Quit]
18:34:11thewinwin88 joins
18:38:27thewinwin8 quits [Ping timeout: 268 seconds]
18:38:32thewinwin88 is now known as thewinwin8
18:39:02n9nes joins
18:40:27nine quits [Client Quit]
18:40:41nine joins
18:52:16DogsRNice joins
18:56:02Guest58 quits [Quit: My Mac has gone to sleep. ZZZzzz…]
19:12:11arch quits [Remote host closed the connection]
19:12:33arch (arch) joins
19:29:10McAfee leaves [Disconnected: Hibernating too long]
19:31:27pabs quits [Read error: Connection reset by peer]
19:33:28pabs (pabs) joins
19:34:17arch quits [Remote host closed the connection]
19:34:30arch (arch) joins
19:37:02n9nes quits [Ping timeout: 268 seconds]
19:37:55nine quits [Client Quit]
19:38:09nine joins
19:38:11n9nes joins
19:50:09arch quits [Remote host closed the connection]
19:50:28arch (arch) joins
19:55:10nine quits [Client Quit]
19:55:23nine joins
19:59:03nine quits [Client Quit]
19:59:15nine joins
20:07:44McAfee joins
20:17:51n9nes quits [Ping timeout: 248 seconds]
20:20:50n9nes joins
20:22:01thewinwin85 joins
20:25:45thewinwin8 quits [Ping timeout: 268 seconds]
20:25:51thewinwin85 is now known as thewinwin8
20:28:31n9nes quits [Ping timeout: 248 seconds]
20:29:24DLoader_ (DLoader) joins
20:30:04n9nes joins
20:33:46DLoader quits [Ping timeout: 268 seconds]
21:01:14hackbug77 quits [Remote host closed the connection]
21:02:12McAfee leaves
21:04:13hackbug joins
21:06:33McAfee joins
21:23:12lunik1 quits [Quit: :x]
21:23:43n9nes quits [Ping timeout: 268 seconds]
21:24:19n9nes joins
21:25:13lunik1 joins
21:28:34McAfee leaves
21:30:04Webuser143737 quits [Quit: Ooops, wrong browser tab.]
21:30:45McAfee joins
22:15:53etnguyen03 (etnguyen03) joins
22:40:03steering is now known as RJHacker46102
22:40:43RJHacker46102 is now known as steering
22:40:44steering is now known as RJHacker27932
22:45:56RJHacker27932 quits [Quit: Reconnecting]
22:47:28steering7253 (steering) joins
22:50:19McAfee leaves
22:51:08McAfee joins
22:58:34etnguyen03 quits [Client Quit]
23:10:40Wohlstand (Wohlstand) joins
23:16:12McAfee leaves
23:47:35thewinwin88 joins
23:51:06Church quits [Ping timeout: 268 seconds]
23:51:11thewinwin8 quits [Ping timeout: 248 seconds]
23:51:13thewinwin88 is now known as thewinwin8
23:55:18etnguyen03 (etnguyen03) joins
23:58:15<didyousayboop>I guess because the entire site (Data Lounge) is designed around infinite scroll without very little navigation to speak of, it’s difficult for Archive Bot to crawl it?
23:59:44<didyousayboop>Or maybe I misunderstand?