00:17:35Matthww5 quits [Ping timeout: 252 seconds]
00:22:30AlsoHP_Archivist (HP_Archivist) joins
00:23:05HP_Archivist quits [Ping timeout: 252 seconds]
00:39:06Wuzado quits [Read error: Connection reset by peer]
00:39:17Wuzado joins
00:39:17Wuzado quits [K-Lined]
01:57:42<@JAA>eBay seems to have loginwalled searching for completed listings (sold or unsold). The item pages themselves are still accessible. A quick web search suggests that this started in July.
02:00:01pie_ quits [Remote host closed the connection]
02:06:43<nicolas17>webuser's request on #archiveteam
02:06:47<nicolas17>curl -H "Referer: https://qu.ax/LUne4.mp4/" -v https://qu.ax/x/LUne4.mp4 # this works
02:07:06grill quits [Ping timeout: 255 seconds]
02:07:37<pokechu22>so !ao https://qu.ax/LUne4.mp4/ should probably work in AB
02:07:39grill (grill) joins
02:08:06<nicolas17>hm interesting, let's try it
02:08:16<pokechu22>Yeah, looks like it worked (it got ~100MB)
02:08:36<nicolas17>without the referer, the video URL 302-redirects to the webpage
02:09:22<nicolas17>cool
02:54:53Umbire quits [Ping timeout: 252 seconds]
02:59:25Webuser247347 joins
03:01:10<pokechu22>Webuser819111: All of them seem to have saved successfully
03:01:23<pokechu22>er sorry, Webuser247347: ^
03:01:55<nicolas17>pokechu22: oh damn, I mixed things up and thought we needed an !a
03:02:02<Webuser247347><pokechu22> How are you guys able to save the qu.ax videos but I can't? Do you just modify the server headers or something to display the raw video/file?
03:02:05<nicolas17>I could have done !ao<
03:02:32<Webuser247347>Whenever I save them on the wayback machine, it redirects to the html page displaying the video and the video fails to save
03:02:44<nicolas17>Webuser247347: we use archivebot to get the webpage, and archivebot sees there's a video in the page and requests the video *with* the Referer: header
03:02:52<pokechu22>Archivebot behaves slightly differently from web.archive.org/save/ (SPN), and it's able to set the proper headers in this case (more specifically, saving the html page displaying the video ends up loading the video embed with the proper headers)
03:03:11<nicolas17>does SPN even follow <video>?
03:03:41<pokechu22>I think it doesn't do HLS, but archivebot doesn't do HLS either. It might do mp4s in some cases?
03:10:43etnguyen03 quits [Remote host closed the connection]
03:11:39<Webuser247347>I know you guys are not associated with archive.org but does anyone know if the wayback machine changed "live web proxy calls" to save page now? or are they two separate things? I notice some pages say live web proxy calls in the "about this capture" section and others say "save page now"
03:13:38<pokechu22>https://archive.org/details/liveweb says it's a variant of save page now, but it looks like it's mostly old captures at this point?
03:14:54<Webuser247347>I wonder if that variant still exists
03:15:19<Webuser247347>Also, I have another one that I am requesting to be archived https://qu.ax/pNYG4.mp4
03:16:30<pokechu22>Done
03:19:01<pokechu22>(and as mentioned before, archivebot data can take a few days to get uploaded and indexed onto web.archive.org)
03:25:44Cupping128544 quits [Quit: Ping timeout (120 seconds)]
03:25:59Cupping128544 joins
03:26:58<Webuser247347>I'm okay with that
03:27:50<Webuser247347>Just curious, does the internet archive allow you to get your own .warc files injected into the wayback machine?
03:28:14<Webuser247347>Like, say I wanted to archive a page and I uploaded a warc file and asked them to include it on the wayback machine
03:31:51Island quits [Read error: Connection reset by peer]
03:34:33notarobot345 quits [Quit: The Lounge - https://thelounge.chat]
03:37:03DogsRNice quits [Read error: Connection reset by peer]
03:39:09<nicolas17>Webuser247347: no
03:39:22notarobot345 joins
03:40:28<Webuser247347>nicolas17 How do you guys get your warcs injected into the wayback machine? Do you build trust with archive.org or did they just decide to inject your archives into there?
03:40:48notarobot345 quits [Client Quit]
03:41:08<pokechu22>Existing trust
03:42:45<nicolas17>well, for these gu.ax URLs we're not injecting our own *local* warcs
03:42:55<nicolas17>we're submitting URLs to #archivebot (which *also* requires trust and permission), then archivebot downloads the URLs and uploads WARCs somehow
03:54:17Webuser247347 quits [Client Quit]
03:55:16Cupping128544 quits [Client Quit]
03:56:12lev (lev) joins
03:56:52Cupping128544 joins
03:57:58Umbire joins
04:02:22notarobot345 joins
04:06:16notarobot345 quits [Client Quit]
04:07:27BornOn420 quits [Read error: Connection reset by peer]
04:07:27lev quits [Remote host closed the connection]
04:08:31BornOn420 (BornOn420) joins
04:10:28notarobot345 joins
04:10:38Cupping128544 quits [Client Quit]
04:11:10Cupping128544 joins
04:11:51BornOn420 quits [Remote host closed the connection]
04:12:58BornOn420 (BornOn420) joins
04:21:47Umbire quits [Ping timeout: 252 seconds]
04:54:45goecho quits [Quit: 👋]
04:56:55nicolas17 quits [Ping timeout: 268 seconds]
05:23:31Hackerpcs quits [Quit: Hackerpcs]
05:24:29Hackerpcs (Hackerpcs) joins
05:30:22nexussfan quits [Remote host closed the connection]
05:31:57Webuser055856 joins
05:32:28Webuser055856 quits [Client Quit]
05:42:59invadeuze (invadeuz) joins
05:44:21invadeuze quits [Client Quit]
05:59:04beastbg8 quits [Read error: Connection reset by peer]
06:01:10Matthww5 joins
06:10:16cascode_ joins
06:14:37cascode quits [Ping timeout: 268 seconds]
06:14:42Erebus quits [Ping timeout: 240 seconds]
06:15:20cascode joins
06:18:09Erebus (Erebus) joins
06:18:56cascode_ quits [Ping timeout: 268 seconds]
06:26:27BitByBit410082 quits [Ping timeout: 252 seconds]
06:29:12BitByBit410082 (BitByBit) joins
06:31:22Erebus quits [Ping timeout: 240 seconds]
06:50:40BornOn420 quits [Remote host closed the connection]
06:51:46BornOn420 (BornOn420) joins
06:57:32Erebus (Erebus) joins
07:04:39BornOn420 quits [Remote host closed the connection]
07:05:03BornOn420 (BornOn420) joins
07:25:21BornOn420 quits [Remote host closed the connection]
07:26:26BornOn420 (BornOn420) joins
07:34:36benjins3_ joins
07:37:51benjins3 quits [Ping timeout: 255 seconds]
07:44:48Umbire joins
08:02:54cascode quits [Read error: Connection reset by peer]
08:03:09cascode joins
08:07:15cascode_ joins
08:07:15cascode quits [Read error: Connection reset by peer]
08:11:37beastbg8 joins
08:13:20BornOn420 quits [Remote host closed the connection]
08:14:30BornOn420 (BornOn420) joins
08:18:52BornOn420 quits [Remote host closed the connection]
08:19:56BornOn420 (BornOn420) joins
09:00:21Bleo182600722719623455222011079 quits [Quit: The Lounge - https://thelounge.chat]
09:03:08Bleo182600722719623455222011079 joins
09:24:00h|ca2 quits [Quit: WeeChat 4.10.0]
09:27:02h|ca2 (h) joins
09:38:39AK quits [Quit: AK]
09:45:52McAfee leaves [Disconnected: Replaced by new connection]
09:45:55McAfee joins
09:56:55AK (AK) joins
10:05:00nune quits [Remote host closed the connection]
10:05:01nune joins
11:07:32pepsi2024 joins
12:25:48h|ignition quits [Ping timeout: 272 seconds]
12:28:00d02 joins
12:33:46McAfee leaves
12:54:02d02 quits [Client Quit]
13:37:07h|ignition (h) joins
13:38:27Island joins
13:44:29McAfee joins
13:54:09Radzig quits [Ping timeout: 252 seconds]
14:02:57Radzig joins
14:14:14thewinwin8702 joins
14:18:05thewinwin870 quits [Ping timeout: 268 seconds]
14:18:16thewinwin8702 is now known as thewinwin870
14:44:20Webuser763684 quits [Quit: Ooops, wrong browser tab.]
14:45:30Wohlstand (Wohlstand) joins
14:58:35Webuser892358 joins
15:00:12Wohlstand quits [Ping timeout: 255 seconds]
15:18:39Cupping128544 quits [Ping timeout: 255 seconds]
15:20:09Cupping128544 joins
15:33:35<h2ibot>Kittyluv edited Amino (+35): https://wiki.archiveteam.org/?diff=63871&oldid=61133
15:33:36<h2ibot>JustAGrook edited Wattpad (+383, vital signs. theres two pages of the same…): https://wiki.archiveteam.org/?diff=63873&oldid=63745
15:33:37<h2ibot>WASDwasd edited YouTube (+0): https://wiki.archiveteam.org/?diff=63874&oldid=63668
15:33:38<h2ibot>Kittyluv edited Wikimedia Commons (+42): https://wiki.archiveteam.org/?diff=63875&oldid=63591
15:34:35<h2ibot>Kittyluv edited ArchiveTeam wiki (+890): https://wiki.archiveteam.org/?diff=63876&oldid=42135
15:35:35<h2ibot>Arkiver edited Amino (-8, Merge edit by…): https://wiki.archiveteam.org/?diff=63877&oldid=63871
15:35:36<h2ibot>Belankt294842 created Talk:ShoutWiki (+293, /* August 31, 2026. */ new section): https://wiki.archiveteam.org/?oldid=63878
15:35:37<h2ibot>Grill edited Discourse/active (+432, talk.jekyllrb.com discourse.elm-lang.org…): https://wiki.archiveteam.org/?diff=63879&oldid=63736
15:37:55<@arkiver>anyone have an idea for a channel for weebly?
15:38:34<justauser>#dweebly ?
15:38:35<h2ibot>Arkiver uploaded File:Weebly-icon.png: https://wiki.archiveteam.org/?title=File%3AWeebly-icon.png
15:38:58<@arkiver>justauser: that was fast, yes let's do it
15:40:01<justauser>Dweeb is a game with fun sounds. Probably enough to be considered humorous.
15:42:36<h2ibot>Justauser edited Chrome Web Store (-12, Manifest v2 extensions removed): https://wiki.archiveteam.org/?diff=63881&oldid=63035
15:42:58Cupping128544 quits [Client Quit]
15:43:31Cupping128544 joins
15:46:33ThreeHM quits [Ping timeout: 255 seconds]
15:48:21ThreeHM (ThreeHeadedMonkey) joins
16:02:57Cupping128544 quits [Client Quit]
16:03:11Cupping128544 joins
16:05:17Cupping128544 quits [Client Quit]
16:05:29Cupping128544 joins
16:15:36BornOn420 quits [Remote host closed the connection]
16:16:03BornOn420 (BornOn420) joins
16:50:30Webuser801697 joins
16:58:07Doranwen quits [Remote host closed the connection]
16:58:29Doranwen (Doranwen) joins
17:12:13Webuser0458530 joins
17:14:43<h2ibot>IDKhowToEdit edited YouTube (+230, Added some information about AT project): https://wiki.archiveteam.org/?diff=63883&oldid=63874
17:15:43<h2ibot>IDKhowToEdit edited YouTube (+6, fix): https://wiki.archiveteam.org/?diff=63884&oldid=63883
17:18:12Webuser892358 quits [Client Quit]
17:22:43<h2ibot>IDKhowToEdit edited YouTube (+2, fix): https://wiki.archiveteam.org/?diff=63885&oldid=63884
17:33:50nicolas17 (nicolas17) joins
17:43:29Goofybally quits [Killed (NickServ (GHOST command used by Goofybally9!~Goofyball@145.82.36.26))]
17:43:34Goofybally joins
17:46:42unknownsrc quits [Ping timeout: 255 seconds]
17:50:19unknownsrc (unknownsrc) joins
17:59:50thewinwin8708 joins
18:03:32Wohlstand (Wohlstand) joins
18:03:47thewinwin870 quits [Ping timeout: 268 seconds]
18:03:56thewinwin8708 is now known as thewinwin870
18:14:08Webuser003593 joins
18:14:29nine quits [Ping timeout: 252 seconds]
18:15:07<Webuser003593>Hello all, just a quick FYI : https://usenet-rewind.com/ Search Usenet newsgroup archives — text posts from 1981 to present. nearly a billion mesages
18:15:48nine joins
18:16:48<h2ibot>Nemo bis edited Talk:ShoutWiki (+199, /* August 31, 2026. */ +re): https://wiki.archiveteam.org/?diff=63886&oldid=63878
18:30:00msfjarvis quits [Quit: Lurker 2.2.1 (the truth is out there) https://lurker.chat]
18:30:58msfjarvis joins
18:37:06Wohlstand quits [Ping timeout: 255 seconds]
18:39:06snvy (snvy) joins
18:41:36Umbire quits [Ping timeout: 255 seconds]
18:49:31Wohlstand (Wohlstand) joins
18:59:14Boppen_ quits [Read error: Connection reset by peer]
19:01:36ericgallager quits [Read error: Connection reset by peer]
19:04:36Boppen joins
19:10:10<snvy>hello, I'd love to help with pushing forward the zapytaj.onet.pl archival [no way it finishes only with the current archivebot job alone]. I have the website structure and some of its quirks documented, along with a list of all valid urls of the first 1 million question ids. should i just go straight up with creating the wiki page as a starting
19:10:10<snvy>point or should I get some approval before moving things forward? thx [and sorry if question is obvious, first time doing something like this :D]
19:11:13Webuser003593 quits [Client Quit]
19:11:32<@arkiver>snvy: yes please! put all information you have on the wiki page
19:12:01<@arkiver>deadline is 2026-09-30 i see?
19:12:39<@arkiver>just had a look - nice site!
19:12:52<@arkiver>snvy: please ping me when you have the information on the wiki
19:14:09Webuser0458530 quits [Client Quit]
19:17:07<snvy>arkiver: yeah, deadline is correct, thankfully full site archive seems to be doable with the time left. will start work on the wiki page in a few moments, will get back to you
19:28:37Webuser558151 joins
19:32:36thewinwin8702 joins
19:36:36<Webuser558151>Hello, I was just about to ask in regards to zapytaj.onet.pl as a coincidence -- I would like to contribute my internet connection to running an archive warrior VM if a project would be created for it (I would simply run the warrior for any kind of project but, I am concerned of extensive flash wear for the projects that write large amounts of data
19:36:36<Webuser558151>to disk in a loop. and I know trying to modify it myself to use a ramdisk or whatever is a big no no)
19:36:54thewinwin870 quits [Ping timeout: 268 seconds]
19:37:04thewinwin8702 is now known as thewinwin870
19:38:25<skankhunt42>Webuser558151: If you run a single grab and you are sure, that your /dev/shm can hold the data for that project, I think thats fine. Probably not in regards to the warrior.
19:39:57<pokechu22>My impression is that modifying the project script is a big problem, but running the project VM in a specific configuration is less of an issue (though I know there still is some stuff regarding ARM, and I don't have authority to say any specific configuration is fine)
19:40:29<skankhunt42>basically what pokechu22 said <3
19:41:14<skankhunt42>I think (saying without authority too) you are fine if you don't modify pipelines nor do something spicy with networking.
19:42:08<Webuser558151>my concern is of running out of ramdisk space and submitting broken data as a result
19:43:25<pokechu22>I would expect that the job would just fail (or hang and then be eventually reclaimed when the tracker sees it hasn't gotten any data back after a day) in that case instead of submitting broken data, same as if you ran out of normal disk space or lost power. But I'm not an expert
19:43:27<skankhunt42>yeah, you should make sure that the amount of data wont exceed your ramdisk. otherwise you'd probably end up doing items for nothing, since, iirc, the item will just fail and needs to be requeued.
19:43:50<skankhunt42>pokechu22++
19:43:50<eggdrop>[karma] 'pokechu22' now has 432 karma!
19:45:57<skankhunt42>personally, I use zram to avoid swapping to disk and reduce wear, but thats a personal decision.
19:49:02Umbire joins
20:00:27Umbire quits [Ping timeout: 252 seconds]
20:00:36BornOn420 quits [Remote host closed the connection]
20:01:19BornOn420 (BornOn420) joins
20:04:56linuxgemini2 (linuxgemini) joins
20:06:12linuxgemini quits [Ping timeout: 255 seconds]
20:06:12linuxgemini2 is now known as linuxgemini
20:31:18lev (lev) joins
20:47:22lev quits [Ping timeout: 240 seconds]
20:47:51<IDK>Webuser558151: in my experience, most jobs does negligible amount of wear
20:48:24<IDK>the only project I had issues with wear is #down-the-tube
20:54:28ericgallager joins
21:03:32Wohlstand quits [Client Quit]
21:10:26thewinwin8706 joins
21:13:00brk275 quits [Read error: Connection reset by peer]
21:13:07brk275 joins
21:13:41Doranwen quits [Remote host closed the connection]
21:14:09thewinwin870 quits [Ping timeout: 252 seconds]
21:14:14thewinwin8706 is now known as thewinwin870
21:15:04Doranwen (Doranwen) joins
21:15:19invadeuz (invadeuz) joins
21:33:02LddPotato quits [Remote host closed the connection]
21:58:45<invadeuz>Hi, I had a question about the archiving of noblogs : 2 days ago pokechu22 gave me this link https://archive.fart.website/archivebot/viewer/job/2o9yg . Do you have an idea if / when the archiving bot will end it's crawl ? A/I will shutdown in 8 days so it's urgent
22:03:52<invadeuz>(it was the crawl of 8136 entries here : https://transfer.archivete.am/mA5Go/noblogs.org_additional_subdomains_sep_2026.txt )
22:03:52<eggdrop>inline (for browser viewing): https://transfer.archivete.am/inline/mA5Go/noblogs.org_additional_subdomains_sep_2026.txt
22:14:59<@JAA>invadeuz: You can monitor it live on http://archivebot.com/ . Note that there's a delay of about 2 days on the uploads to IA, and then it'll take a bit longer still to show up in the viewer.
22:15:28<snvy>arkiver: dropped all the documentation and links i found on https://wiki.archiveteam.org/index.php?title=Zapytaj_Onet [pending review]. might be a bit messy, sorry about that [tired after long day, I'll maybe clean it up tomorrow]
22:22:53DogsRNice joins
22:48:54etnguyen03 (etnguyen03) joins
22:53:57<pokechu22>invadeuz: That one has made pretty significant progress and been through most of the sitemaps at this point; it's mostly grabbing outlinks now. I also started an extra job with all outlinks ignored for all noblogs.org subdomains, which will end up at https://archive.fart.website/archivebot/viewer/job/5l38c (and uses
22:53:59<pokechu22>https://transfer.archivete.am/lCZHE/noblogs.org_all_subdomains_sep_2026_no_outlinks.txt as the list); that one is still working through posts
22:53:59<eggdrop>inline (for browser viewing): https://transfer.archivete.am/inline/lCZHE/noblogs.org_all_subdomains_sep_2026_no_outlinks.txt
22:59:38<invadeuz>JAA thanks, I didn't have the info on de 2 days delay, it makes sense
23:00:51<invadeuz>pokechu22 I had a question about outlinks : does it means that totally different sites are captured as well in the archives ?
23:01:45<pokechu22>Yes (e.g. if a post on noblogs.org links to a news article, we capture that news article, as well as any images in that article)
23:02:30<pokechu22>That's also (part of) why the original noblogs.org subdomain crawls are so big
23:03:57<invadeuz>and so your extra job, what does it do ? why outlinks were ignored in the first place ?
23:04:49superkuh joins
23:05:42<invadeuz>aaah I think I understand
23:05:45<pokechu22>The new one was because the clear shutdown date was announced, and it seemed worthwhile to have one where the WARCs only contain noblogs.org URLs so it would be smaller (but also would cover everything, rather than it being split across a bunch of confusingly-named jobs)
23:06:12<pokechu22>it's technically redundant
23:07:15<invadeuz>ok I understand, thanks ! It will be super useful indeed, it's a great idea
23:09:18nexussfan (nexussfan) joins
23:14:30<invadeuz>thanks a lot for your work, it would have been hard to do it on our own in such a short time, and the redundancy on the wayback machine is great to have
23:31:00etnguyen03 quits [Client Quit]
23:56:35etnguyen03 (etnguyen03) joins