00:09:32<h2ibot>Usernam edited List of websites excluded from the Wayback Machine (+201): https://wiki.archiveteam.org/?diff=63154&oldid=63118
00:12:48rohvani quits [Quit: Ping timeout (120 seconds)]
00:13:03rohvani joins
00:13:23etnguyen03 (etnguyen03) joins
00:18:26moth3 quits [Remote host closed the connection]
00:18:53moth3 joins
00:20:09daxxy quits [Quit: bye]
00:23:26moth3 quits [Remote host closed the connection]
00:23:41moth3 joins
00:28:11apergos joins
00:28:40h|ignition quits [Ping timeout: 272 seconds]
00:29:21h|ignition (h) joins
00:30:25<apergos>probably extremely late to the party, but is it safe to assume that people are already aware of the shutdown at the end of the month of https://en.wikipedia.org/wiki/DataLounge ?
00:58:22error2 quits [Ping timeout: 240 seconds]
00:59:44error (disinfecttm) joins
01:04:25Tom1992 joins
01:07:25Tom1992 quits [Client Quit]
01:22:58<mindrecent>apergos: it's on deathwatch but I haven't seen anything else about it (I'm new though).
01:30:57<IDK>love how data.world is returning 429 even on brand new IP
01:32:50Tom1992 joins
01:40:23<Tom1992>Excuse me, how can I find a recent status of archiving https://fami-net.no-ip.org/mirrors/qt/unofficial_builds?
01:42:37tzt quits [Remote host closed the connection]
01:42:57tzt (tzt) joins
02:38:09etnguyen03 quits [Remote host closed the connection]
03:05:41nine quits [Quit: See ya!]
03:05:55nine joins
03:13:17Island quits [Read error: Connection reset by peer]
03:26:40Hackerpcs quits [Remote host closed the connection]
03:27:33Hackerpcs (Hackerpcs) joins
03:33:06Tom1992 quits [Client Quit]
04:27:06DogsRNice quits [Read error: Connection reset by peer]
04:35:35nexussfan quits [Quit: Konversation terminated!]
04:46:29<h2ibot>Manu edited Distributed recursive crawls (+75, Candidates: Add https://www.militar.org.ua/): https://wiki.archiveteam.org/?diff=63155&oldid=63142
04:48:15Yakov2 (Yakov) joins
04:49:51Yakov quits [Ping timeout: 248 seconds]
04:49:52Yakov2 is now known as Yakov
05:01:43thewinwin82 joins
05:05:27thewinwin8 quits [Ping timeout: 268 seconds]
05:05:33thewinwin82 is now known as thewinwin8
05:07:44Pendonym quits [Quit: Pendonym]
05:08:55Pendonym (Pendonym) joins
05:10:39khaoohs_ quits [Remote host closed the connection]
05:11:04khaoohs_ joins
05:35:44sg-72 quits [Quit: Leaving]
06:35:23Tom1992 joins
06:39:37apergos quits [Quit: Ooops, wrong browser tab.]
06:56:20Tom1992 quits [Client Quit]
07:04:05michaelblob7641341 quits [Quit: yoop]
07:07:36michaelblob7641341 joins
07:53:42ThreeHM quits [Read error: Connection reset by peer]
07:56:57ThreeHM (ThreeHeadedMonkey) joins
08:02:34nine quits [Quit: See ya!]
08:02:47nine joins
08:24:13thewinwin87 joins
08:28:20thewinwin8 quits [Ping timeout: 268 seconds]
08:28:23thewinwin87 is now known as thewinwin8
09:09:26<@arkiver>nicolas17: doesn't AB do that? i'm not sure
09:09:39<@arkiver>#Y can let you do that eventually, but remember there is a bloom filter
09:10:05<@arkiver>hmm but i think i can add an option to turn the bloom filter off but then disallow recursive archiving
09:10:11<@arkiver>nicolas17: what do you think of that? ^
09:26:16systwo (systwi) joins
09:28:16goro (goro) joins
09:30:37systwi quits [Ping timeout: 268 seconds]
09:31:01simon816 quits [Quit: ZNC 1.10.2 - https://znc.in]
09:40:43simon816 (simon816) joins
09:41:12amphitryon_ joins
09:43:43amphitryon0 quits [Ping timeout: 248 seconds]
09:59:27sg72 joins
10:07:51Yna joins
10:10:31<Yna>Getting 429'd hardcode from dataworld
10:34:33scotrod2 quits [Quit: The Lounge - https://thelounge.chat]
10:35:06scotrod2 joins
10:35:35scotrod2 quits [Client Quit]
10:36:39scotrod2 joins
11:00:20Bleo18260072271962345522201107 quits [Quit: The Lounge - https://thelounge.chat]
11:00:24BitByBit4100 (BitByBit) joins
11:03:05Bleo18260072271962345522201107 joins
11:03:37Radzig2 joins
11:05:25emily quits [Quit: ZNC 1.10.2 - https://znc.in]
11:06:20pseudorizer (pseudorizer) joins
11:07:26Radzig quits [Ping timeout: 268 seconds]
11:07:26Radzig2 is now known as Radzig
11:10:56pseudorizer quits [Client Quit]
11:11:26pseudorizer (pseudorizer) joins
11:24:28Webuser039857 joins
11:35:32goro quits [Remote host closed the connection]
11:38:42goro (goro) joins
11:44:43<h2ibot>Brad edited WhatsApp (+147, Added more technical infomation and 3rd party…): https://wiki.archiveteam.org/?diff=63156&oldid=52461
11:45:40inedia quits [Ping timeout: 268 seconds]
12:04:02systwo is now known as systwi
12:46:04<fuzzy80211>arkiver may want a few more retries. the few l9gs i had before the code update looked like i generally needed between 5 and 10 minutes to get past the 429. new code is aborting after 4 mins
12:46:49<fuzzy80211>on dataworld
12:58:29nine quits [Quit: See ya!]
12:58:42nine joins
13:03:13<@arkiver>fuzzy80211: understood, adding that
13:11:10<Yna>Yeah, submitting very little so far
13:27:38skankhunt424 (skankhunt42) joins
13:30:30skankhunt42 quits [Ping timeout: 268 seconds]
13:30:30skankhunt424 is now known as skankhunt42
13:47:54TheEnbyperor quits [Read error: Connection reset by peer]
13:56:55Umbire (Umbire) joins
13:58:10Webuser087791 joins
13:58:27Webuser087791 quits [Client Quit]
14:02:02<Yna>Nevermind, it does not appear to be working at all
14:07:15inedia joins
14:09:17<h2ibot>Cruller edited Data.world (+987, Organized the shutdown into a section. (To make…): https://wiki.archiveteam.org/?diff=63157&oldid=63144
14:11:17<h2ibot>Cruller edited Deathwatch (-14, /* 2026-07 */ Changed data.world into an…): https://wiki.archiveteam.org/?diff=63158&oldid=63128
14:11:19<@arkiver>thanks cruller
14:14:43nexussfan (nexussfan) joins
14:15:57<cruller|irc>:D
14:16:17<cruller|irc>BTW, I found another archival discussion/project: Quick! Let's save our TV Time discussions, together, before they're gone (July 15th) https://old.reddit.com/r/TVTime/comments/1usu5k1/quick_lets_save_our_tv_time_discussions_together/
14:35:58traxys quits [Quit: The Lounge - https://thelounge.chat]
14:37:52thewinwin85 joins
14:40:33traxys (traxys) joins
14:40:47thewinwin8 quits [Ping timeout: 248 seconds]
14:40:51thewinwin85 is now known as thewinwin8
14:45:59thewinwin83 joins
14:48:12thewinwin8 quits [Ping timeout: 268 seconds]
14:48:14thewinwin83 is now known as thewinwin8
14:57:35<klea>Uhh, that's a deadline of 1d?
14:58:04<Yna>Sounds reasonable
14:58:39<@arkiver>looking into audioleaf.com for a project
14:59:41<@arkiver>for nicotalk, does anyone just get the nicotalk.com string on http://www.nicotalk.com/ ?
14:59:44<@arkiver>not sure if it's still up
15:00:03<Yna>Yes
15:00:41<justauser>Same for me.
15:01:04<@arkiver>i ee
15:01:06<@arkiver>i see
15:01:54<nicolas17>arkiver: afaik archivebot doesn't deduplicate data unfortunately (if two URLs in the same list return identical files, the data will be stored twice in the warc)
15:02:01<@arkiver>look like datalounge (shutting down end of this month) was archived with AB https://archive.fart.website/archivebot/viewer/job/2026062605192221mqp
15:02:27<@arkiver>nicolas17: yes and also since each URL is an individual item, #Y would not work
15:03:06<justauser>Doesn't really work in AB.
15:03:17<justauser>JAA wanted to qwarc it?
15:03:43<klea>Wget-AT would also work, but it needs someone with collection:web access to do it.
15:05:38Island joins
15:10:40<@arkiver>JAA: is it correct AB does not deduplicate records URL-agnostically?
15:11:53matoro quits [Quit: https://quassel-irc.org - Chat comfortably. Anywhere.]
15:12:12matoro joins
15:14:46<klea>AB only dedupes URLs from the same job, and even if on the same job, if there's two different URLs which return the same content, it'll store the same content twice. (AFAIK)
15:18:27<@arkiver>understood
15:19:05<@arkiver>disk space is so expensive nowadays, i wonder if we should start calculating duplicate data per AB job and potential savings
15:19:16<@arkiver>then run somethig to deduplicate URL-agnostically when savings can be high
15:19:32<klea>Note IIRC it will refetch on non successful status codes a few times IIRC?, not exactly sure how that works, but IIRC some codes are considered success when sometimes that's not desired, and then people need to ask JAA to requeue stuff.
15:19:54ericgallager quits [Quit: This computer has gone to sleep]
15:20:14<klea>IIRC there's an issue for certain URLs to be ran only once, and they are added to a ignore set when possible.
15:20:39<nicolas17>oh yeah that too
15:20:54<klea>(The badvideos ignore set)
15:20:55<nicolas17>there's stuff like Google Fonts which are archived on multiple jobs
15:21:01<@arkiver>are the response records with bad status code saved in the WARC?
15:21:06<klea>IIRC yes.
15:21:08<justauser>I think so.
15:21:11<nicolas17>and now WBM has thousands of copies of each (same URL)
15:21:38<justauser>404 and 403 are considered successes and are not retried.
15:22:30<justauser>#jseater has a global deduplication database IIRC.
15:22:50<klea>warcprox to be specific does that.
15:23:26<klea>IIRC JAA's #UncleSamsArchive project for the Epstein files also has a dedupe db, or makes revisit records.
15:24:03<klea>Also, it'd be neat to have the source code for the qwarc jobs for that, but the source is included in the warcs, which are ARId, so that's fun.
15:25:54<klea>https://github.com/ArchiveTeam/ArchiveBot/issues/443 Is the dedupe for specific URLs (currently those URLs are thrown into the badvideos ignore set, but there's a queue of stuff to add to the ignore sets because time)
15:25:54<klea>https://github.com/ArchiveTeam/ArchiveBot/issues/464 Refused connections not retried.
15:25:54<klea>No issue (I think?) for the 403 and 404 as success.
15:27:34<nicolas17>hm I agree with arkiver that we should measure the extent of the problem first...
15:27:46<klea>True.
15:27:49amphitryon_ is now known as amphitryon
15:31:11<cruller|irc>arkiver: That URL http://www.nicotalk.com/ has been in that state for a very long time. The site is alive. It's a very small site, but the internal links aren't working, so I'll collect the URLs and share them.
15:34:17nine quits [Client Quit]
15:34:30nine joins
15:34:48<@arkiver>ah
15:34:52<@arkiver>hopefully that can go into AB
15:43:12goro quits [Remote host closed the connection]
15:43:35TheEnbyperor_ joins
15:44:39goro (goro) joins
15:51:43BitByBit4100 quits [Ping timeout: 248 seconds]
16:00:15ericgallager joins
16:09:36ericgallager quits [Ping timeout: 268 seconds]
16:10:23Cuphead2527480 (Cuphead2527480) joins
16:11:59goro quits [Ping timeout: 248 seconds]
16:14:27<c3manu>can someone here help me with this S3 bucket thing? the manuals for Oneplus aren't discovered due to the JS-heavy website
16:14:52<c3manu>i downloaded a PDF using this URL, and would like to get a list with all of them :) https://s3.eu-west-3.amazonaws.com/par-sow-cms-oppo-com/cms/userManual/manual/fee86647-7fe1-4de9-a522-873a998e41ff/OnePlus%2015%20Quick%20Guide.pdf
16:15:04<c3manu>found that specific PDF from https://service.oneplus.com/dk/user-manual#/detail?equipmentModelId=1979231480117637121
16:20:42moth3 quits [Ping timeout: 268 seconds]
16:21:23<c3manu>i know about the scripts in little-things, but i don't know which one expects what kind of URL
16:22:28<justauser>The bucket doesn't appear open?
16:24:12<c3manu>what would it looks like if it weren't?
16:24:29<justauser>Like a list of files in XML.
16:24:43<c3manu>on which URL?
16:24:52<c3manu>https://s3.eu-west-3.amazonaws.com/par-sow-cms-oppo-com/ ?
16:25:19<justauser>I think so...
16:26:28<klea>Oh, another possible use of the deduped thing would be for S3 buckets, which have two different URLs.
16:26:56<c3manu>dangit
16:28:16MrMcNugg1 quits [Quit: WeeChat 4.9.2]
16:29:17<klea>Sample open buckets: https://s3-us-west-1.amazonaws.com/umbrella-static/ (also available as https://umbrella-static.s3-us-west-1.amazonaws.com/)
16:30:23<c3manu>thanks you both
16:33:42Webuser324338 joins
16:37:39MrMcNuggets (MrMcNuggets) joins
16:38:16<c3manu>i'd be open for suggestions on how to gather a list of those pdf files for #archivebot ...one that's better than "manually" :)
16:38:27Webuser324338 quits [Client Quit]
16:39:02<c3manu>mostly for the manuals/pdfs on https://service.oneplus.com/
16:43:31earthnative quits [Ping timeout: 268 seconds]
16:45:20<fuzzy80211>arkiver on dataworld, not sure how long one should wait. still seeing aborts after waiting 28 minutes
16:46:19<@arkiver>fuzzy80211: i am starting to get the feeling there is some kind of global limit, either in general, or with one of our headers or something
16:46:31<@arkiver>which means i should probably decreasse item rate
16:46:39<@arkiver>will do that now and see what happens
16:46:57<fuzzy80211>with sleeping that long, not sure rate limit has a practical affect
16:48:14<@arkiver>perhaps, but maybe they count 429 responses as well in their limit
16:48:54<fuzzy80211>if thats the case, should you put in a higher429 immediately instead of starting at 2s
16:49:14<fuzzy80211>start at 120?
16:59:43<Yna>Yeah im having trouble getting anything
17:02:21<justauser>Transfer bork?
17:02:45<justauser>500 on uploads.
17:07:26earthnative (earthnative) joins
17:11:03edideaur7 joins
17:12:31<edideaur7>hi, wanted to let y'all know krakenfiles.com is shutting down as per their announcement 2 days ago. Thank you.
17:14:21<justauser>Announcement link?
17:21:28<Yna>Not seeing one, either
17:25:38<edideaur7>used to be on their site, not sure what happened to it. uploading a screengrab now,
17:28:42<edideaur7>getting 500 errors from transfer.archiveteam.org, https://files.catbox.moe/q6o0iz.png
17:30:27nine quits [Client Quit]
17:30:40nine joins
17:30:45<edideaur7>either way, i'm not sure if distributed archival is possible as they enforce CAPCHA on download pages.
17:30:50edideaur7 quits [Client Quit]
18:04:00<@JAA>arkiver: AB doesn't currently write revisit records under any circumstances (including when the same URL is fetched multiple times, e.g. via redirects).
18:05:52<pokechu22>AB also uses gz, not zstd
18:07:58McAfee joins
18:08:00<HP_Archivist>https://transfer.archivete.am down for anyone else? won't accept an upload
18:09:45<pokechu22>I've seen others say that yeah
18:11:23<klea>I got "Could not save metadata" when testing it now.
18:11:28<klea>JAA ^
18:12:05<@JAA>Yep, already investigating.
18:20:11Cuphead2527480 quits [Client Quit]
18:21:30<@JAA>Should be fixed.
18:25:07<klea>Seems to work.
18:30:41<HP_Archivist>Thank you JAA
18:35:39Grzesiek11 quits [Remote host closed the connection]
18:36:33DogsRNice joins
18:43:44Grzesiek11 (Grzesiek11) joins
18:48:07Grzesiek11 quits [Remote host closed the connection]
18:50:05Grzesiek11 (Grzesiek11) joins
18:58:04QZxV4 (QZxV4) joins
19:54:56h|ca2 quits [Ping timeout: 248 seconds]
20:01:31h|ca2 (h) joins
20:02:44McAfee leaves
20:07:16Webuser843256 joins
20:09:39<Webuser843256>Hi! I don't know how this all works but I was wondering if you could help me with a request. I'm trying to track down the contents of a Storify.com link from 2015 and was told that ArchiveTeam scraped Storify before it went down. The link is http://storify.com/KIRORadio/shutdowna14-protests-block-streets-in-downtown-se. Any help/leads would be much
20:09:39<Webuser843256>appreciated!
20:21:56<@JAA>Webuser843256: Hi, I'm the one who ran that back in the day. That post should be archived, but it's not currently accessible in any sensible way. But if this content is extremely important to you and you have a very high pain tolerance, I can give you some hints at where to dig. Make sure to bring a big showel or, better, an excavator...
20:26:34<Webuser843256>Hey, thanks so much! Any hints would be amazing.
20:27:52<Webuser843256>I know a guy with a backhoe...
20:28:59<@JAA>It's in the 650 GiB tar in https://archive.org/details/Archiveteam_Storify_Raw_Data_2018 and probably closer to the beginning (i.e. lower WARC number) than the end. Should've been archived at about 2018-05-10 09:50.
20:30:25<@JAA>I still want to sort that data out and get it into the WBM, but it's not a trivial task.
20:35:38<klea>hmm, why is the data in a tar to avoid getting WBM indexed?
20:39:44fuzzy80211 quits [Killed (NickServ (GHOST command used by fuzzy8021))]
20:39:50fuzzy80211 (fuzzy80211) joins
20:40:31<Webuser843256>Yeah it doesn't sound trivial! Thank you. I'm looking at the file listing and it only shows timestamps from June 2018, which aren't the same as the capture dates (if it was captured in May). I figure you don't have the WARC number for that capture...?
20:41:32klea downloads the entire tar, to untar and gencdx it.
20:42:26<@JAA>I don't. You can try to look at the beginnings of each file to figure out which one contains the right time window.
20:42:42<@JAA>The timestamps are 'wrong' due to an intermediate transfer of the data.
20:42:47<@JAA>(IIRC)
20:42:55h|ca2 quits [Ping timeout: 248 seconds]
20:43:07<@JAA>Preserving file metadata is hard.
20:46:05<klea>fun, ETA 19h
20:48:44a782s2z joins
20:48:48a782s2z quits [Remote host closed the connection]
20:49:05<c3manu>excavator for JAA on christmas, got it *checks notes*
20:49:57<@JAA>:-)
20:54:49h|ca2 (h) joins
21:00:18thewinwin82 joins
21:03:12Webuser751706 joins
21:03:23Webuser751706 quits [Client Quit]
21:03:27<Webuser843256>Thanks all! I'm going offline for a bit but I'll watch the logs of this channel & let you know if I have any luck. 👷‍♀️🪏
21:04:06Webuser843256 quits [Client Quit]
21:04:22thewinwin8 quits [Ping timeout: 268 seconds]
21:04:25thewinwin82 is now known as thewinwin8
21:24:03McAfee joins
21:29:23BitByBit4100 (BitByBit) joins
21:54:25khaoohs_ quits [Read error: Connection reset by peer]
22:02:24<klea>c3manu: curl -X POST https://sow-cms-fr.oneplus.com/oppo-api/officialWebsiteEquipmentUserManual/v1/list --data-binary '{}' -H 'Content-Type: application/json'
22:03:10<klea>https://transfer.archivete.am/dIVzF/sow-cms-fr.oneplus.com_oppo-api_officialWebsiteEquipmentUserManual_v1_list.json.zst
22:04:17<klea>Then make a bunch of POSTs to https://sow-cms-fr.oneplus.com/oppo-api/officialWebsiteEquipmentUserManual/v1/detail with a json like {"region":"en","equipmentModelId":"1996263580649701377","langId":"3082","isoLanguageCode":"en-US","sourceRoute":"1","brandCode":"12"} but changing the equipmentModelId.
22:08:30<klea>The list seemed to be incomplete with {}, so I replaced it with {"region":"es","equipmentModelId":"1996263580649701377","langId":"3082","isoLanguageCode":"en-US","sourceRoute":"1","brandCode":"12"} and made https://transfer.archivete.am/oyoqG/sow-cms-fr.oneplus.com_oppo-api_officialWebsiteEquipmentUserManual_v1_list_parms_browser.json
22:08:31<eggdrop>inline (for browser viewing): https://transfer.archivete.am/inline/oyoqG/sow-cms-fr.oneplus.com_oppo-api_officialWebsiteEquipmentUserManual_v1_list_parms_browser.json
22:11:33<klea>From that I did `jq -r '.data[][].equipmentModelId'|sort|uniq`, making https://transfer.archivete.am/inline/Kn5io/oneplus_modelids.txt (zst)
22:18:52<klea>I put a script to run to get the jsons for each item, will go to eep now.
22:27:54<klea>And I made https://transfer.archivete.am/tI7tN/oneplus.tar.zst, with the list in oneplus/urls.txt, left it there because it has some ' ' and I'm not sure how AB needs it.
22:27:54<eggdrop>inline (for browser viewing): https://transfer.archivete.am/inline/tI7tN/oneplus.tar.zst,
22:29:00<klea>I hope that's useful. :)
22:33:44etnguyen03 (etnguyen03) joins
22:57:34khaoohs joins
23:07:24Umbire quits [Quit: Umbire zaps a wand of digging!]
23:22:53moth3 joins