| 00:06:19 | | etnguyen03 (etnguyen03) joins |
| 01:46:49 | | nexussfan (nexussfan) joins |
| 01:58:40 | | justauser|ttt quits [Quit: Lost terminal] |
| 02:05:33 | | Lolo quits [Ping timeout: 268 seconds] |
| 02:15:14 | | Lolo (Lolo) joins |
| 02:34:09 | | benjins3_ joins |
| 02:37:21 | | benjins3 quits [Ping timeout: 252 seconds] |
| 02:39:07 | | benjins3__ joins |
| 02:42:39 | | benjins3_ quits [Ping timeout: 255 seconds] |
| 02:57:53 | | oxtyped quits [Ping timeout: 252 seconds] |
| 03:00:35 | | lev (lev) joins |
| 03:03:29 | | lev quits [Remote host closed the connection] |
| 03:06:37 | | oxtyped joins |
| 03:10:53 | | etnguyen03 quits [Remote host closed the connection] |
| 03:32:07 | | Island quits [Read error: Connection reset by peer] |
| 03:35:14 | | thewinwin870494633 joins |
| 03:39:17 | | thewinwin87049463 quits [Ping timeout: 268 seconds] |
| 03:39:26 | | thewinwin870494633 is now known as thewinwin87049463 |
| 03:56:58 | | DogsRNice_ quits [Read error: Connection reset by peer] |
| 04:12:41 | <AccelArch> | justauser: forum.turtlecraft.gg — wpull (AB/grab-site) tá travando aí, mas testei com um crawler de browser real (Browsertrix) e passou liso, sem precisar de cookie manual. Vou rodar um crawl completo (~25-30k páginas). O bloqueio é só detecção de bot/desafio JS, ou tem conteúdo atrás de login que eu vou perder rodando deslogado? Acesso público parece completo até agora, mas quero co |
| 04:13:38 | | Medowar quits [Quit: ZNC 1.9.1+deb2+b3 - https://znc.in] |
| 04:14:57 | | Medowar (Medowar) joins |
| 04:19:21 | <AccelArch> | justauser: sorry, wrong language above! English: forum.turtlecraft.gg — wpull (AB/grab-site) is getting blocked there, but I tested with a real-browser crawler (Browsertrix) and it got through fine, no manual cookie needed. |
| 04:19:45 | <AccelArch> | Running a full crawl (~25-30k pages). Also confirmed memberlist/profile pages require login, but forum/topics seem publicly accessible. Is the block just bot/JS-challenge detection, or is there login-gated content I'd be missing by crawling logged out? |
| 04:46:37 | | lev (lev) joins |
| 04:51:02 | | lev quits [Ping timeout: 240 seconds] |
| 05:10:06 | | BornOn420 quits [Remote host closed the connection] |
| 05:21:03 | | lev (lev) joins |
| 05:23:13 | | lev quits [Remote host closed the connection] |
| 05:27:19 | | nexussfan quits [Remote host closed the connection] |
| 05:49:24 | | Pendonym quits [Ping timeout: 268 seconds] |
| 05:50:13 | | Pendonym (Pendonym) joins |
| 05:59:05 | | Pendonym quits [Read error: Connection reset by peer] |
| 06:14:47 | | Mateon1 quits [Ping timeout: 252 seconds] |
| 06:17:19 | | skankhunt42 quits [Quit: Ping timeout (120 seconds)] |
| 06:17:32 | | camrod63629 quits [Quit: no-clipped reality] |
| 06:17:47 | | camrod63629 (camrod) joins |
| 06:18:51 | | skankhunt42 (skankhunt42) joins |
| 06:21:36 | | scotrod2 quits [Quit: The Lounge - https://thelounge.chat] |
| 06:22:22 | | scotrod2 joins |
| 06:22:51 | | scotrod2 quits [Client Quit] |
| 06:22:58 | | Webuser155154 joins |
| 06:23:37 | | Webuser155154 quits [Client Quit] |
| 06:24:16 | | Pendonym (Pendonym) joins |
| 06:24:34 | | scotrod2 joins |
| 06:27:15 | <hexagonwin> | btw, the 'Checking your browser, please wait...' page on forum.turtlecraft.gg contains a simple script to set some cookie, so it should be trivial to scrape without a browser if making a custom crawler i believe |
| 06:34:53 | | Jon quits [Quit: ZNC 1.9.1+deb2+b3 - https://znc.in] |
| 06:35:18 | | Jon joins |
| 06:35:18 | | Jon is now authenticated as Jon |
| 06:49:50 | | Umbire is now known as RJHacker76648 |
| 06:49:54 | | Umbire (Umbire) joins |
| 06:52:16 | | Umbire is now authenticated as * |
| 06:52:16 | | Umbire quits [Killed (zoidberg.hackint.org (Nickname regained by services))] |
| 06:52:19 | | Umbire (Umbire) joins |
| 06:53:45 | | RJHacker76648 quits [Ping timeout: 255 seconds] |
| 07:04:33 | | scrl__ quits [Ping timeout: 255 seconds] |
| 07:04:45 | | systwiii (systwi) joins |
| 07:08:20 | | systwo quits [Ping timeout: 268 seconds] |
| 07:10:48 | | khaoohs quits [Ping timeout: 268 seconds] |
| 07:11:10 | | jinn6 quits [Read error: Connection reset by peer] |
| 07:11:27 | | jinn6 (jinn6) joins |
| 07:14:55 | | cyanbox joins |
| 07:33:53 | | khaoohs joins |
| 08:21:43 | | APOLLO03 quits [Ping timeout: 268 seconds] |
| 08:29:48 | | benjins3__ quits [Read error: Connection reset by peer] |
| 08:30:27 | | benjins3__ joins |
| 08:36:06 | | HP_Archivist quits [Quit: Leaving] |
| 08:40:51 | | michaelblob764134108363059 joins |
| 08:41:36 | | michaelblob76413410836305 quits [Read error: Connection reset by peer] |
| 08:41:36 | | michaelblob764134108363059 is now known as michaelblob76413410836305 |
| 08:42:07 | | mjpowersjr quits [Read error: Connection reset by peer] |
| 08:42:28 | | mjpowersjr joins |
| 08:44:21 | | pepsi2024 joins |
| 08:49:20 | | Umbire quits [Client Quit] |
| 09:00:09 | | superkuh quits [Ping timeout: 252 seconds] |
| 09:00:21 | | Bleo182600722719623455222011079 quits [Quit: The Lounge - https://thelounge.chat] |
| 09:02:00 | | superkuh joins |
| 09:03:05 | | Bleo182600722719623455222011079 joins |
| 09:21:06 | | APOLLO03 joins |
| 09:24:58 | <@JAA> | AccelArch: Browsertrix writes bad WARCs. |
| 09:45:12 | | h|ignition quits [Ping timeout: 272 seconds] |
| 09:51:21 | | HP_Archivist (HP_Archivist) joins |
| 09:57:01 | <hexagonwin> | it's like the only option if someone wants an easily runnable browser based crawler though.. |
| 10:00:21 | | BornOn420 (BornOn420) joins |
| 10:00:26 | | h|ignition (h) joins |
| 10:01:42 | <@JAA> | Well yeah, the setups that write proper WARCs are more complicated because it's impossible to do it the way Browsertrix does it. |
| 10:02:55 | | BornOn420 quits [Client Quit] |
| 10:04:26 | <@JAA> | In particular, doing it without a proxy isn't possible, and that necessarily complicates the setup. |
| 10:05:00 | | nune quits [Remote host closed the connection] |
| 10:05:02 | | nune joins |
| 10:09:29 | <hexagonwin> | i've tried to get warcprox running properly a few times but it's really buggy for some reason and stops working after running a few hours |
| 10:09:38 | <cruller> | IMO, Zeno's --headless option is easy to use, but I haven't verified whether the WARC files it generates are correct. |
| 10:10:04 | <hexagonwin> | and considering there's no good way of viewing those WARCs other than using browsertrix/webrecorder tools, i don't think it's a bad option for personal archival |
| 10:10:49 | <hexagonwin> | personally made WARCs don't even get into the wayback machine anyway so.. |
| 10:11:12 | <@JAA> | Writing and reading are almost entirely orthogonal problems, and the data corruption issues are much less critical on the latter part. |
| 10:12:14 | <hexagonwin> | does browsertrix actually corrupt data? wasn't it more about how it writes whats provided by the browser's CDP instead of raw binary data provided by the web server and that's kinda against WARC specs? |
| 10:12:22 | <@JAA> | That's corruption. |
| 10:13:27 | <@JAA> | And it's not 'kinda' against the spec. It plainly and clearly is. |
| 10:13:40 | <hexagonwin> | if those two differ, doesn't that mean the data provided by web server is already corrupt? (and if the web server sent valid data, shouldn't the CDP output and raw data be equal?) |
| 10:13:49 | <@JAA> | No |
| 10:14:01 | <@JAA> | CDP doesn't expose the raw data, only a parsed representation. |
| 10:15:11 | <@JAA> | That's exactly why no approach of writing WARCs from within a browser can ever work (unless the browser APIs change, which I don't see happening). |
| 10:20:53 | <hexagonwin> | there are certain cases where having a full browser based crawler is the only option though. has there been any attempts on editing chromium to expose raw data on CDP then? |
| 10:22:31 | <@JAA> | Not that I'm aware of. |
| 10:23:00 | <@JAA> | Having to compile Chromium sounds more painful than dealing with warcprox. :-) |
| 10:23:06 | <@JAA> | You'd also have to disable HTTP/2 and HTTP/3. |
| 10:27:47 | | pepsi2024 quits [Ping timeout: 252 seconds] |
| 10:29:25 | <klea> | That sounds like it'd be easier to take something like Chromium and stick it in some kind of container, pre-configured with warcprox, and ship all that to the user as a executable. |
| 10:29:33 | <hexagonwin> | WARCs don't support HTTP/2, 3? |
| 10:29:45 | <hexagonwin> | warcprox was buggy in my usage but maybe because i used it with firefox |
| 10:29:55 | <hexagonwin> | (IA seems to only use it with chromium?) |
| 10:30:11 | <@JAA> | No, only HTTP/1.1. Depending on how you read the spec, not even HTTP/1.0, and definitely not HTTP/0.9. |
| 10:30:16 | <hexagonwin> | damn |
| 10:30:34 | <hexagonwin> | well i guess that should be addressed in a spec update at some point |
| 10:31:07 | <@JAA> | Yeah. It really should've been part of the 2017 revision in my opinion. |
| 10:36:26 | | Mateon1 joins |
| 10:37:04 | <@arkiver> | hexagonwin: is this about archiving https://forum.turtlecraft.gg/ ? |
| 10:37:56 | | Wohlstand1 (Wohlstand) joins |
| 10:40:19 | | Wohlstand1 is now known as Wohlstand |
| 10:44:30 | | pepsi2024 joins |
| 11:17:26 | <AccelArch> | JAA: Thanks for the heads-up. |
| 11:41:22 | | d02 joins |
| 11:50:09 | | mjpowersjr quits [Ping timeout: 268 seconds] |
| 12:17:00 | | BornOn420 (BornOn420) joins |
| 12:21:16 | <AccelArch> | Are you guys also having trouble uploading files? |
| 12:36:26 | <AccelArch> | It's back to normal now. |
| 12:40:36 | | devnoname120 joins |
| 12:42:52 | <devnoname120> | Hey, is there a chance you could recursively archive this website? https://files.olebeck.com/ |
| 12:42:52 | <devnoname120> | It’s a trove of firmware that are hard to find anywhere, would be a huge loss if it goes down. Thanks a lot! |
| 12:44:22 | | BornOn420 quits [Remote host closed the connection] |
| 12:45:15 | | Erebus quits [*.net *.split] |
| 12:46:14 | | Erebus (Erebus) joins |
| 12:54:26 | | BornOn420 (BornOn420) joins |
| 13:12:40 | | thewinwin870494632 joins |
| 13:15:14 | | d02 quits [Client Quit] |
| 13:16:29 | | thewinwin87049463 quits [Ping timeout: 268 seconds] |
| 13:16:32 | | thewinwin870494632 is now known as thewinwin87049463 |
| 13:17:02 | | Island joins |
| 13:41:42 | | ATinySpaceMarine quits [Quit: https://quassel-irc.org - Chat comfortably. Anywhere.] |
| 13:44:18 | | lev (lev) joins |
| 13:45:02 | | lev quits [Remote host closed the connection] |
| 13:55:14 | | d02 joins |
| 13:56:34 | <katia> | devnoname120, i staarted a job for this on archivebot.com - you can follow progress there |
| 14:02:58 | | devnoname120 quits [Client Quit] |
| 14:10:21 | | mjpowersjr joins |
| 14:13:00 | | ATinySpaceMarine joins |
| 14:35:14 | | Webuser720452 joins |
| 14:40:42 | <Webuser720452> | Hi! Question about the askfm project. The #dontaskfm channel mentioned on https://github.com/ArchiveTeam/askfm-grab#other-problems seems to be invite-only. |
| 14:40:42 | <Webuser720452> | I'm trying to recover ask.fm Q&A for my family member. The Wayback has only a few pages, which is the find_max() probe pattern from askfm.lua, not a content crawl. Looks like queue_pages() queued the rest and the project ended before they were claimed. |
| 14:40:42 | <Webuser720452> | The archiveteam_askfm is access-restricted (401) so I can't grep the megawarc CDX myself. |
| 14:40:42 | <Webuser720452> | Just trying to find a hint what I can check next. |
| 14:41:31 | <@arkiver> | Webuser720452: if i remember correctly, that project was not finished in time, their site was too slow |
| 14:41:47 | <@arkiver> | yeah see the numbers here https://tracker.archiveteam.org/askfm/ |
| 14:45:45 | | unknownsrc quits [Quit: Ping timeout (120 seconds)] |
| 14:45:59 | | unknownsrc (unknownsrc) joins |
| 14:49:09 | <Webuser720452> | Yeah, I see. Do you know if the completed part is fully in the Wayback or is some of it only inside the restricted WARCs? Just trying to get whether it's "never crawled" or "crawled but not publicly visible" |
| 14:49:43 | <@arkiver> | Webuser720452: anything that was archived should be in the WARCs, and those WARCs should be indexed in the Wayback Machin |
| 14:49:46 | <@arkiver> | e |
| 14:50:04 | <@arkiver> | there are cases in which a WARC can fall out of the Wayback Machine for some time, but that is unlikely to be the case with these |
| 14:50:59 | <Webuser720452> | Okay, thanks you! |
| 14:51:40 | | Webuser720452 quits [Client Quit] |
| 15:02:37 | <hexagonwin> | arkiver: idk much about that website, and my messages above weren't strictly related to it |
| 15:03:35 | <hexagonwin> | about that specific website, their 'browser check' seems trivial to implement in a custom script so maybe DPoS should be done about it? |
| 15:05:41 | <@JAA> | It has some ugly session ID stuff that needs to be dealt with as well. |
| 15:25:22 | | pepsi2024 quits [Ping timeout: 268 seconds] |
| 15:25:40 | | pepsi2024 joins |
| 15:42:21 | | Ryz quits [Quit: Ping timeout (120 seconds)] |
| 15:42:38 | | dxrt_ (dxrt) joins |
| 15:42:38 | | @ChanServ sets mode: +o dxrt_ |
| 15:42:57 | | @dxrt quits [Ping timeout: 255 seconds] |
| 15:43:04 | | Ryz (Ryz) joins |
| 15:45:42 | | do2 joins |
| 15:47:59 | | croissant` joins |
| 15:49:21 | | croissant quits [Ping timeout: 252 seconds] |
| 15:49:25 | | d02 quits [Ping timeout: 268 seconds] |
| 15:58:36 | | croissant` is now known as croissant |
| 15:59:41 | | do2 quits [Client Quit] |
| 16:10:37 | | PredatorIWD8 quits [Quit: Ping timeout (120 seconds)] |
| 16:10:49 | | PredatorIWD8 joins |
| 16:18:00 | | pabs quits [Read error: Connection reset by peer] |
| 16:18:51 | | pabs (pabs) joins |
| 16:31:55 | | KenSoftTH joins |
| 16:34:58 | <KenSoftTH> | arkiver: small parser bug report on the youtubedislikes public metadata release: description URLs were taken from the display text of description.runs[], which YouTube shortens to 37 chars + "…", instead of the real target in runs[].navigationEndpoint.urlEndpoint.url (the /redirect?q= link). So every URL over 37 characters in every description in |
| 16:34:59 | <KenSoftTH> | the release is truncated, and a ")" right after a link is dropped with it. |
| 16:34:59 | <KenSoftTH> | Example: zyQBnGrT5is has http://www.facebook.com/ClhleeslVIChe... where the raw youtubei/v1/next response carries the full URL in the endpoint. Noticed it while trying to recover download links from a terminated channel's descriptions — every unrecoverable one is exactly a >37‑char URL. It's a one‑line parser fix on the source data if a |
| 16:34:59 | <KenSoftTH> | re‑parse is ever on the table; happy to share it. |
| 16:35:58 | <@arkiver> | KenSoftTH: the metadata release is not the WARCs right? |
| 16:36:08 | <@arkiver> | where is this metadata release? |
| 16:36:20 | <KenSoftTH> | Let me pull the URL, one moment please |
| 16:36:27 | <KenSoftTH> | WARCs should contain the full URL |
| 16:36:53 | <KenSoftTH> | This one https://www.reddit.com/r/DataHoarder/comments/rsu7lf/dislikes_and_other_metadata_for_456_billion/ |
| 16:37:18 | <@arkiver> | KenSoftTH: ah i see |
| 16:37:31 | <@arkiver> | i don't think jopik is here at the moment, maybe someone can get the message to jopik |
| 16:37:44 | <@arkiver> | but that was not me, Archive Team only did the WARCs |
| 16:38:00 | <KenSoftTH> | I sent message to Jopik in April or May-ish and he said he does not have the WARCs anymore, so he cannot reparse it |
| 16:38:18 | <KenSoftTH> | He told me to reach out to you guys because WARCs is not downloadable anymore |
| 16:38:42 | <@arkiver> | ah |
| 16:39:09 | <@arkiver> | all data should be indexed in the Wayback Machine though, but i know it's hard to get it in bulk from there |
| 16:39:31 | <@arkiver> | blame LLMs for WARCs only being available through the Wayback Machine now :/ |
| 16:39:58 | <KenSoftTH> | Yep, LLM sucks, everyone wanna train off the data right now. |
| 16:40:25 | <KenSoftTH> | Problem is, for this project specifically, the WARC is not reachable via Wayback Machine (I tried...) |
| 16:40:53 | <KenSoftTH> | I checked that route first, gated with control captures known to exist. The dislikes grab's youtubei/v1/next records aren't in the Wayback CDX at all — 0 hits for this channel's 151 videos, and enumerating the whole youtubei key space (19.1M keys for /player alone) turned up nothing from that grab for any video. The items look like they never |
| 16:40:53 | <KenSoftTH> | went through CDX indexing. |
| 16:41:29 | <KenSoftTH> | And /next is a POST: even where POST captures are indexed, the CDX key is the __wb_method=POST form with the flattened request body, so there's no URL a person can replay. You need the body to reconstruct the key. IA's own youtubecrawl get_video_info captures are indexed (that's how I know the method works), so it's not that I couldn't find them; |
| 16:41:29 | <KenSoftTH> | the dislikes records genuinely aren't there. |
| 16:41:36 | <@arkiver> | hmm |
| 16:41:48 | <@arkiver> | KenSoftTH: are you sticking around for tomorrow? we can continue then |
| 16:41:55 | <@arkiver> | or perhaps someone else wants to continue the conversation |
| 16:42:01 | <@arkiver> | i need to refresh my brain for the night though |
| 16:42:20 | <KenSoftTH> | I will be sticking around tomorrow. I actually need the missing five links to complete my archive. Sent you an email last week. Thanks! |
| 16:43:15 | <KenSoftTH> | It is 1:42 AM at my place as well, so good night to you too! |
| 16:43:32 | <@JAA> | I'm pretty sure that __wb_method is a pywb(/webrecorder?) thing and not present in IA's CDXs. |
| 16:44:13 | <@JAA> | In some cases, we did POST requests with a suitable query string so they can be found via the URL, but not sure about that particular project. |
| 16:45:24 | <KenSoftTH> | I think I have tried that, but got nothing. Let me pull up my note. |
| 16:46:57 | <KenSoftTH> | The dislike project grab POSTed to /next with video_id in the query string, so if its records had been indexed they'd sit there as plain URL keys. However, they don't, for any video, not just ours. So the records aren't in CDX at all; InnerTube being POST is why they were only ever reachable via the WARCs. |
| 16:50:45 | <@JAA> | No, POST responses do get indexed. |
| 16:51:45 | <@JAA> | (There's no indication of the request method in the response record anyway.) |
| 16:52:44 | <h2ibot> | Nintendofan885 edited Template:Tr (+3, fixing [[Special:DoubleRedirects|double redirect]]): https://wiki.archiveteam.org/?diff=63978&oldid=62868 |
| 16:53:00 | <KenSoftTH> | Right, and that's the form I looked at: plain URL keys. The whole youtubei/v1/next prefix in the CDX API is 4 captures in total, all 2023 probes, none with a video_id= parameter; the 31 __wb_method ones on top are just ArchiveWeb.page uploads. So if the grab's response records had been ingested they'd be sitting there as ordinary keys, and they |
| 16:53:00 | <KenSoftTH> | aren't, unless the collection is excluded from the public CDX rather than never ingested, which you'd know better than I do. |
| 16:53:44 | <KenSoftTH> | I tried query that and got nothing, even with valid video_id parameter. So, I believe it is not there. |
| 16:53:44 | <h2ibot> | Nintendofan885 edited Category:Project with active dedicated IRC channels (+1, fixing [[Special:DoubleRedirects|double redirect]]): https://wiki.archiveteam.org/?diff=63979&oldid=50731 |
| 16:55:03 | <@JAA> | Yeah, something is up, but I'm not sure what. |
| 16:55:48 | <KenSoftTH> | Would appreciate if you could take a look. Thank you! |
| 16:57:01 | <@JAA> | Oh, archiveteam_youtubedislikes is in the mirrortube collection rather than archiveteam. That'd do it. |
| 16:57:04 | <@JAA> | arkiver: ^ |
| 16:58:44 | <KenSoftTH> | JAA: that explains everything I saw, thank you. Our five items all list mirrortube in their collections too. If it gets re-ingested, each record is reachable per video by URL (video_id is in the query string), which solves my case and everyone else's. I'll check back tomorrow. |
| 16:59:35 | <@JAA> | The collection part is fixable quickly, but the records won't appear in the WBM for some time (months). |
| 16:59:36 | <c3manu> | "Frolenkov added that attacks on Ukraine's communications infrastructure are not new, pointing to earlier strikes on energy and telecommunications networks. But the recent escalation of attacks on data centers and communication nodes marks a shift in intensity, he said." - https://kyivindependent.com/russias-latest-target-ukraines-internet/ |
| 17:03:07 | <KenSoftTH> | Months is fine for everyone else, but could I ask a favour for the five that started this? Five response records. I have item + video_id for each; each item's .cdx.gz gives offset/length, so it's five range reads from the .megawarc.warc.zst plus the project dict. Raw frames or JSON bodies, either works, and I have a script that does it if that |
| 17:03:07 | <KenSoftTH> | saves time. |
| 17:03:58 | <KenSoftTH> | They're MediaFire links from 8–12 years ago; the account's other files are still up but MediaFire prunes without notice, and I've been on this since Feb–Mar, so I'm afraid my files will be gone before this is done. |
| 17:06:28 | <@JAA> | Yeah, if you already know the items, it can be done. |
| 17:06:37 | <KenSoftTH> | Can I send you the DM? |
| 17:06:58 | <pokechu22> | Would we already have forwarded those to #mediaonfire? |
| 17:07:06 | <@JAA> | KenSoftTH: Sure |
| 17:07:15 | <@JAA> | pokechu22: No, doesn't look like we did that in that project. |
| 17:07:23 | <KenSoftTH> | pokechu22: I have searched that and nope, my file does not exist. |
| 17:07:40 | <KenSoftTH> | Would be interesting to forward all mediafire link from YouTube description to that project tho |
| 17:09:01 | <@JAA> | Yeah |
| 17:10:07 | <KenSoftTH> | Jopik's filmot also based on the data from broken parser, so all URL there are all truncated |
| 17:20:04 | | Wohlstand quits [Quit: Wohlstand] |
| 17:36:25 | | @imer quits [Ping timeout: 252 seconds] |
| 17:44:10 | | imer (imer) joins |
| 17:44:10 | | @ChanServ sets mode: +o imer |
| 17:53:50 | | @imer quits [Client Quit] |
| 17:54:22 | | imer (imer) joins |
| 17:54:22 | | @ChanServ sets mode: +o imer |
| 17:57:28 | | Wohlstand (Wohlstand) joins |
| 18:24:53 | | Webuser459770 joins |
| 18:25:24 | | sg72 quits [Quit: Leaving] |
| 18:26:14 | <Webuser459770> | Hi! I'm trying to recover my own old Hyves profile from the 2013 ArchiveTeam crawl. My username was xbaby (http://xbaby.hyves.nl/). |
| 18:26:15 | <Webuser459770> | The username index points to: |
| 18:26:15 | <Webuser459770> | 20131121234649/hyves-xbaby-20131122-052403.warc.gz |
| 18:26:15 | <Webuser459770> | CDX shows successful 200 captures of my profile, photos, blog and friends pages, but the Internet Archive item archiveteam_hyves_20131121234649 now says “This item is no longer available”, and Wayback won't replay the captures. |
| 18:26:15 | <Webuser459770> | Does anyone know if hyves-xbaby-20131122-052403.warc.gz is still accessible somewhere, or how I could retrieve the archived contents? |
| 18:26:24 | | sg72 joins |
| 18:30:00 | | msfjarvis quits [Quit: Lurker 2.3.1 (the truth is out there) https://lurker.chat] |
| 18:30:32 | | msfjarvis joins |
| 18:31:38 | | @imer quits [Client Quit] |
| 18:32:10 | | imer (imer) joins |
| 18:32:10 | | @ChanServ sets mode: +o imer |
| 18:45:27 | | Mateon1 quits [Remote host closed the connection] |
| 18:45:43 | | Mateon1 joins |
| 18:50:36 | | mjpowersjr quits [Ping timeout: 255 seconds] |
| 18:53:06 | <AccelArch> | Update on forum.turtlecraft.gg: confirmed hexagonwin guess — the "Checking your browser" page just sets a static cookie via a plain <script> tag (tc_pass=<fixed value>, no nonce/computation), same value confirmed across 3+ independent fetches/IPs. So a normal wpull/grab-site crawl works fine with that cookie as a static header, no browser needed. |
| 18:53:33 | <AccelArch> | Also confirmed: no robots.txt on the domain (404 after the bypass) — no crawl restrictions declared. |
| 18:54:08 | <AccelArch> | One real gotcha for a recursive crawl: phpBB does set a persistent session cookie here and keeps it correctly across requests, but the sid param embedded in every internal link still changes on every single page load regardless — confirmed by walking home→topic both in a real browser and via a session-aware HTTP client, different sid each time. So this will need URL normalization |
| 18:54:12 | <AccelArch> | (stripping sid= before dedup/queueing) to avoid massively duplicating pages in the crawl. |
| 18:56:07 | <@JAA> | Is this LLM output? |
| 18:57:09 | | atphoenix__ quits [Read error: Connection reset by peer] |
| 18:58:21 | <klea> | Sounds like it. |
| 18:58:39 | | atphoenix__ (atphoenix) joins |
| 18:58:39 | | atphoenix__ quits [Remote host closed the connection] |
| 18:59:07 | | atphoenix__ (atphoenix) joins |
| 18:59:48 | | Wohlstand quits [Client Quit] |
| 19:03:27 | <AccelArch> | yeah used an LLM to help test stuff and write it up, but it's not hallucinated — actually ran the requests, checked the cookie across a few different IPs, walked through home->topic both in a real browser and via script to confirm the sid thing. can paste raw output if anyone wants to check |
| 19:07:39 | | atphoenix__ quits [Remote host closed the connection] |
| 19:08:20 | | atphoenix__ (atphoenix) joins |
| 19:13:07 | <@JAA> | Well, the sid part is partially wrong. In any case, I'd rather see your own words than LLM output. |
| 19:40:11 | <cruller> | If it's phpBB 2.x or later, a dedicated scraper might work better. I'll try https://github.com/TUVIMEN/forumscraper with warcprox this weekend, so could you please wait until I finish (or fail) that attempt? |
| 19:40:33 | <cruller> | Even if the initial result is poor, using it as a seed to run grab-site should yield better results than running grab-site alone. |
| 19:41:50 | <AccelArch> | JAA: Alright, I understand. Could you tell me which part is incorrect so I can check it here? |
| 19:45:22 | <AccelArch> | cruller: okay, I'll wait. |
| 19:49:44 | | Pendonym quits [Read error: Connection reset by peer] |
| 19:50:58 | | Pendonym (Pendonym) joins |
| 19:51:12 | | leo60228 quits [Read error: Connection reset by peer] |
| 19:51:16 | | leo60228 (leo60228) joins |
| 20:15:12 | | nine quits [Quit: See ya!] |
| 20:15:24 | | nine joins |
| 20:24:27 | | nine quits [Ping timeout: 268 seconds] |
| 20:25:37 | | nine joins |
| 20:28:28 | | dangersmymidd1ename joins |
| 20:30:43 | | BornOn420 quits [Remote host closed the connection] |
| 20:33:09 | | atphoenix__ quits [Remote host closed the connection] |
| 20:34:39 | | atphoenix__ (atphoenix) joins |
| 20:34:39 | | atphoenix__ quits [Remote host closed the connection] |
| 20:36:39 | | atphoenix_ (atphoenix) joins |
| 20:38:58 | | dangersmymidd1ename quits [Client Quit] |
| 20:40:48 | | BornOn420 (BornOn420) joins |
| 21:06:29 | | Wohlstand (Wohlstand) joins |
| 21:30:05 | <hexagonwin> | duh, why would you paste llm output for such simple stuff |
| 21:30:17 | <hexagonwin> | it's like being lazy to even write |
| 21:38:53 | | Umby joins |
| 21:38:54 | | Umby is now known as Umbire |
| 21:38:54 | | Umbire is now authenticated as Umbire |
| 21:40:27 | <@JAA> | I've archived phpBB with qwarc before. I'll see if I can take a look at the weekend as well. |
| 21:40:50 | <@JAA> | AccelArch: It doesn't always include the sid parameter. That depends on the page. |
| 21:41:08 | <@JAA> | (And whether you already have the cookie, of course.) |
| 21:41:31 | <@JAA> | URL normalisation isn't enough because we want the links to work as well. So it requires reretrieving those pages until there's no sid parameter. |
| 21:51:42 | | atphoenix_ quits [Read error: Connection reset by peer] |
| 21:52:03 | | Webuser459770 quits [Quit: Ooops, wrong browser tab.] |
| 21:52:09 | | atphoenix_ (atphoenix) joins |
| 21:57:07 | | pepsi2024 quits [Ping timeout: 252 seconds] |
| 21:57:52 | | pepsi2024 joins |
| 22:11:45 | | pepsi2024 quits [Ping timeout: 268 seconds] |
| 22:12:02 | | pepsi2024 joins |
| 22:22:34 | | Wohlstand quits [Client Quit] |
| 22:35:43 | | hamouda joins |
| 22:36:49 | | etnguyen03 (etnguyen03) joins |
| 22:44:46 | | thewinwin870494635 joins |
| 22:48:45 | | thewinwin87049463 quits [Ping timeout: 268 seconds] |
| 22:48:53 | | thewinwin870494635 is now known as thewinwin87049463 |
| 22:57:10 | | hamouda quits [Client Quit] |
| 22:58:33 | | pepsi2024 quits [Ping timeout: 255 seconds] |
| 22:59:00 | | pepsi2024 joins |
| 22:59:36 | | hamouda joins |
| 23:19:02 | | nexussfan (nexussfan) joins |
| 23:23:56 | | Webuser177559 joins |
| 23:28:49 | <Webuser177559> | hi, i'm looking for the warc from the archiveteam youtube snapshot at 02;30;32 gmt on oct 5 2020. |
| 23:29:07 | <@JAA> | → #down-the-tube |
| 23:31:06 | | nine quits [Quit: See ya!] |
| 23:31:20 | | nine joins |
| 23:43:57 | | Webuser177559 quits [Client Quit] |
| 23:45:15 | | etnguyen03 quits [Client Quit] |
| 23:50:21 | | thewinwin870494635 joins |
| 23:54:07 | | thewinwin87049463 quits [Ping timeout: 268 seconds] |
| 23:54:11 | | thewinwin870494635 is now known as thewinwin87049463 |