| 00:24:35 | | DogsRNice_ (Webuser299) joins |
| 00:25:58 | | DogsRNice quits [Ping timeout: 265 seconds] |
| 01:17:57 | | prebala-publici quits [Read error: Connection reset by peer] |
| 01:19:45 | | prebala-publici joins |
| 01:22:02 | | prebala-publici quits [Remote host closed the connection] |
| 01:48:18 | <HP_Archivist> | LoC has it's own policy regarding former WH.gov captures for former admins. I assume people here or at IA saved as much as possible before the current admin updated WH.gov in the past 24 hours? |
| 01:53:27 | <@JAA> | The previous sites are preserved anyway. |
| 01:53:45 | <@JAA> | https://trumpwhitehouse.archives.gov/ |
| 01:54:42 | <@JAA> | I'm not sure if whitehouse.gov was archived here before it was moved there and replaced by the Biden administration site, but it was run through ArchiveBot after the move, and I think IA does crawls of .gov in general at the end of term. |
| 01:55:49 | <HP_Archivist> | JAA: Yeah I know there are other means to which it is captured regardless, sounds good though |
| 01:58:09 | <@JAA> | Ah yes, I ran whitehouse.gov through AB pre-election in October. No run between that and the transition. |
| 01:58:57 | <@JAA> | Meh, the WBM calendar for whitehouse.gov doesn't work on the 19th. |
| 01:59:15 | <@JAA> | ('e is undefined', like the other day.) |
| 02:18:06 | | Terbium quits [Ping timeout: 265 seconds] |
| 02:18:18 | | Terbium joins |
| 02:18:35 | | @JAA quits [Ping timeout: 265 seconds] |
| 02:18:39 | | JAA (JAA) joins |
| 02:18:39 | | @ChanServ sets mode: +o JAA |
| 02:20:09 | <HP_Archivist> | JAA: So there were no captures in the WBM on Jan 19th for Whitehouse.gov? |
| 02:23:00 | <@JAA> | HP_Archivist: There definitely were, but you can't list them. Accessing via URL works fine though. E.g. https://web.archive.org/web/20210119235526/https://www.whitehouse.gov/ |
| 03:57:12 | <HP_Archivist> | JAA: I still don't understand. By list what do you mean? |
| 03:58:03 | <@JAA> | Hover over the date on the calendar. It shows an error instead of a list of snapshots on that date. |
| 03:58:37 | <@JAA> | Well, if the calendar loads in the first place instead of throwing a 500 Internal Server Error. |
| 03:59:01 | <@JAA> | Oh, now the 18th and 19th have disappeared from the calendar entirely. |
| 03:59:09 | <@JAA> | That's https://web.archive.org/web/*/whitehouse.gov just to be clear. |
| 03:59:13 | <HP_Archivist> | No, I'm on it now and it doesn't show any captures or dates for 18th or 19th |
| 03:59:16 | <HP_Archivist> | Yeah exactly |
| 03:59:33 | <@JAA> | Yeah, earlier it did but then produced the 'e is undefined' error in the popup. |
| 03:59:47 | <HP_Archivist> | 500 internal error.. |
| 03:59:51 | <HP_Archivist> | I'll try later |
| 04:00:27 | <@JAA> | Safe to say that something's broken at the moment. :-) |
| 04:00:55 | <HP_Archivist> | I've learned not to expect perfection from the WayBack Machine heh |
| 04:09:32 | | fionera quits [Ping timeout: 264 seconds] |
| 04:11:00 | | fionera (Fionera) joins |
| 04:13:27 | | qw3rty_ joins |
| 04:16:44 | | qw3rty quits [Ping timeout: 240 seconds] |
| 04:17:38 | <atphoenix> | yea, time machines are still imperfect. Someday maybe it'll be a solved problem, along with fileformats and emulation. |
| 04:46:55 | | DogsRNice_ quits [Read error: Connection reset by peer] |
| 05:31:36 | | Ajay quits [Remote host closed the connection] |
| 05:43:51 | | Ajay joins |
| 07:47:24 | | prebala-publici joins |
| 07:47:24 | | prebala-publici quits [Remote host closed the connection] |
| 07:47:45 | | prebala-publici joins |
| 07:47:47 | | prebala-publici quits [Remote host closed the connection] |
| 07:48:07 | | prebala-publici joins |
| 08:39:14 | <HP_Archivist> | JAA or atphoenix: Has Archive.org been *actually* capturing YouTube videos? Not sure how long this has been going on - or maybe I just never knew - but I went go save this interview in the WBM and the entire video was already crawled |
| 08:39:21 | <HP_Archivist> | And playable in the WBM along with comments |
| 08:39:23 | <HP_Archivist> | https://web.archive.org/web/20210122083609if_/https://www.youtube.com/watch?v=eYVv4uQjhQE |
| 08:41:48 | <HP_Archivist> | I guess was under the assumption that WBM captures of YT videos never worked/playback wasn't possible |
| 08:47:11 | <atphoenix> | I know that SPN can save twitter videos |
| 08:49:14 | <HP_Archivist> | Well like I said, maybe I just never realized, but the whole damn video along with comments is viewable and loads in the WBM |
| 08:49:21 | <atphoenix> | I wonder if this is related to that mechanism. Maybe it starts downloading it upon actual access, if the video is still up? |
| 08:49:58 | <HP_Archivist> | I don't think so. I've tried saving YT videos in the WBM a few years ago. They were never playable like this |
| 08:50:19 | <HP_Archivist> | Unless this video has been tubeup'ed and somehow the YT video id is matched to the url..? |
| 08:50:54 | <HP_Archivist> | One caveat, I can't seem to load responded-to comments or sub-comments |
| 08:51:40 | <atphoenix> | trying to follow a link to another video got me "Sorry, the Wayback Machine does not have this video (0T4S4nreoKo) archived (or not indexed yet). |
| 08:51:41 | <atphoenix> | " |
| 08:52:06 | <HP_Archivist> | ^^ |
| 08:52:07 | <HP_Archivist> | Hm |
| 08:52:38 | <atphoenix> | but it does load the comments |
| 08:53:57 | <HP_Archivist> | Yeah that's what would happen to me in the past, only comments and the sidebar of suggested videos would be captured |
| 08:54:01 | <atphoenix> | that linked vid was from the 'AD Vids' channel |
| 08:54:03 | <HP_Archivist> | not actual videos |
| 08:54:42 | <HP_Archivist> | Wait, what? That's not the same channel |
| 08:55:12 | <HP_Archivist> | AD Vids is a channel responsible for Tom Snyder interview uploads, I know cause I crawled the channel in #youtubearchive |
| 08:56:38 | <atphoenix> | yes, that channel is what I ended up on when I clicked a related video from your example |
| 08:56:46 | <atphoenix> | the vid id is in the error message |
| 08:57:21 | <atphoenix> | as for your example, I don't see the usual WBM info box at the top |
| 08:57:53 | <HP_Archivist> | Oh, I see what you mean now ^ |
| 08:58:05 | <atphoenix> | oh wait...I did a reload of the Tom Snyder video |
| 08:58:05 | <atphoenix> | https://web.archive.org/web/20210122083609if_/https://www.youtube.com/watch?v=eYVv4uQjhQE |
| 08:58:13 | <atphoenix> | it behaves like your first example now |
| 08:58:38 | <HP_Archivist> | And yes, I noticed that too. But if you search the link in the WBM timeline it will come up |
| 08:59:00 | <atphoenix> | it really appears there is some new behavior going on. It seems to have saved the YT video upon first access! |
| 08:59:16 | <atphoenix> | pretty much how we see twitter do it |
| 08:59:43 | <atphoenix> | I mean how we see IA save twitter |
| 08:59:58 | <HP_Archivist> | Huh. I wonder how long this has been possible?? |
| 09:00:17 | <atphoenix> | I wasn't aware of this until you mentioned it |
| 09:00:26 | <HP_Archivist> | I mean, it's very slow to load that interview but it does load. and it loads the comments (even more valuable..) |
| 09:00:31 | <atphoenix> | i was only aware of twitter vid |
| 09:00:44 | <HP_Archivist> | Hey JAA do you know anything about this? |
| 09:01:02 | <atphoenix> | are you logged in with an IA browser extension? I am. |
| 09:01:19 | <HP_Archivist> | No, I just loaded it in FF, not logged into anything, no extensions |
| 09:01:24 | | qw3rty_ quits [Ping timeout: 240 seconds] |
| 09:01:26 | <atphoenix> | I'm using the github version |
| 09:01:27 | <atphoenix> | ok |
| 09:02:04 | <HP_Archivist> | It seems that somehow the video file is grabbed, I wonder if they're using ytdl, I doubt on the fly video conversion |
| 09:03:05 | <atphoenix> | I bet the missing WBM timeline bar is missing because of that 'if_' tail. I know there are various tails available. |
| 09:03:15 | <atphoenix> | I get the normal timeline view with https://web.archive.org/web/20210122083609/https://www.youtube.com/watch?v=eYVv4uQjhQE |
| 09:03:26 | <HP_Archivist> | It seems this has been available since 2019? https://forum.videohelp.com/threads/391951-How-to-download-a-YouTube-video-archived-by-Wayback-Machine |
| 09:03:47 | <HP_Archivist> | a quick search for 'save youtube videos waybackmachine' brought up that thread |
| 09:04:15 | <atphoenix> | this is in the 'Live Web Proxy Crawls' collection |
| 09:04:56 | <HP_Archivist> | I don't follow, an actual collection on Archive.org? |
| 09:05:20 | <atphoenix> | https://archive.org/details/liveweb |
| 09:05:47 | <atphoenix> | one of many collections on IA. Most are private. AT's collections are mostly public. |
| 09:06:03 | <atphoenix> | but some private collections allow viewing via WBM |
| 09:06:20 | <HP_Archivist> | Hm |
| 09:06:48 | <HP_Archivist> | I wonder if this works on first attempt on any YT video... |
| 09:07:03 | <HP_Archivist> | I tried that interview using the GH chrome extension |
| 09:07:15 | <HP_Archivist> | But it looks like it was crawled in 2018 going by the timeline |
| 09:08:23 | <HP_Archivist> | ivan - though I think your collection would far exceed what's in the WBM, are you aware that YT videos can be saved and playback is possible? |
| 09:08:37 | <atphoenix> | I'd compare this to what has been happening with the URLs project that isn't getting 'requisites'. It only gets the HTML text of a webpage. But if you visit the page in WBM, WBM will try to find the images either in WBM...or on the source server if the source server is still alive. |
| 09:09:03 | <atphoenix> | ivan isn't in here |
| 09:09:29 | <HP_Archivist> | Oh, my bad. Thought he was |
| 09:09:53 | <atphoenix> | so i guess now trying to visit YT pages on WBM might trigger the WBM to save the related videos... |
| 09:10:17 | <HP_Archivist> | But yeah makes sense. I'm still confused how it would grab the video file. YTDL pulls down the highest quality stream possible. I wonder how WBM does it and with what |
| 09:10:46 | <HP_Archivist> | And... that makes archiving very easy... |
| 09:11:20 | <atphoenix> | it looks like ytdl is configurable in terms of the quality it goes for. I haven't tried to tweak those options in my own tests yet. |
| 09:12:03 | <HP_Archivist> | Yeah it is. I've been using YTDL for almost 2 years. And that's what ivan's project uses, but many instances and with a custom script |
| 09:12:11 | <atphoenix> | my guess is IA would use ytdl in the background. I know ytdl works on many other sites besides YT. I use it on twitter vids myself. |
| 09:13:06 | <HP_Archivist> | If they don't use ytdl, how else would they grab the streams, still kind of floored that I can watch YT videos in the WBM heh |
| 09:14:07 | <HP_Archivist> | The entire 'YouTube experience', something that YTDL does not capture, is captured in the WBM is would seem. The html/page, the suggested/recommendation side bar, most comments, the up/down vote on each comment, etc |
| 09:14:42 | <atphoenix> | I don't know what magic they are using to put it all together. Maybe it's just something that happens when warcproxy is used? I don't know enough to say with confidence how they do it. |
| 09:15:32 | <atphoenix> | it's not perfect as you've already noted re:comments. It is still impressive. |
| 09:17:07 | <HP_Archivist> | Yeah, goes beyond my knowledge. Idk if this is a game changer or not when it comes to mass archiving YT. I can say that this adds another tool to our archiving kit as it were, assuming WBM captured videos are not relying on anything externally to still be live on the live page |
| 09:18:09 | <HP_Archivist> | There is a way to test that - I could always upload a test video to my YT account, captured it, then delete the video from my channel |
| 09:18:26 | <HP_Archivist> | See if it still plays in WBM |
| 09:20:21 | <HP_Archivist> | atphoenix: I might try this tomorrow, will let you know what happens. Regardless, cool stuff indeed |
| 09:30:29 | | qw3rty joins |
| 09:48:22 | <mgrandi> | is the WBM having lots of issues lately? it like cannot return a save page request through this library i'm using : "https://web.archive.org:443 "GET /save/https://instagram.com/Merlalias HTTP/1.1" 520 None" |
| 09:58:59 | <HP_Archivist> | Keep trying or try at a later time mgrandi |
| 09:59:16 | <mgrandi> | yeah, my script waits 5 minutes between each attempt |
| 09:59:26 | <mgrandi> | its just been particularly bad this week it seems |
| 09:59:40 | <HP_Archivist> | No surprise given the events this week |
| 09:59:51 | <HP_Archivist> | But yeah, maybe come back and try 30-60 minutes later |
| 09:59:57 | <HP_Archivist> | That's usually what I do and it does work then |
| 10:00:14 | <mgrandi> | yeah, luckily i'm in no rush |
| 10:00:53 | <mgrandi> | the wayback machine could use a real api though, this library claims you just hit archive.org/save/<URL> and then look in the headers and its kinda hacky |
| 10:01:30 | <HP_Archivist> | There is an API |
| 10:01:51 | <HP_Archivist> | I never use it, requires knowledge of python |
| 10:02:32 | <mgrandi> | https://github.com/akamhy/waybackpy/blob/master/waybackpy/utils.py#L195 |
| 10:02:56 | <HP_Archivist> | Oh, nice |
| 10:03:15 | <mgrandi> | well that is the unofficial library, but it has the notes about i guess /save does not return JSON |
| 10:03:34 | <mgrandi> | https://archive.org/help/wayback_api.php says it only really returns json for the querying if a site has been saved i guess |
| 10:04:41 | <mgrandi> | and since it just returns the HTML of the page its hard for it to tell if it succeeded or not, or if its just a html error page i guess |
| 10:05:15 | <HP_Archivist> | Idk if this would help but here's IA's documentation for api use https://archive.org/services/docs/api/ |
| 10:05:31 | <HP_Archivist> | Like I said, I'm pretty novice when it comes to use of it or python |
| 10:06:16 | <mgrandi> | yeah, while confusing, i think that is referring to the Internet Archive itself and interacting with it |
| 10:06:17 | <HP_Archivist> | I did try to use tubeup through the api 6 months back and successfully had it configured for a bit, was nice to use. But never for wbm |
| 10:06:33 | <HP_Archivist> | Right, yeah, what do you mean, use for wbm? |
| 10:06:50 | <mgrandi> | while the WBM seems to be a separate system |
| 10:07:20 | <HP_Archivist> | ^^ good point |
| 10:07:21 | <mgrandi> | my script takes a config file and then tries to do a WBM save of various urls in my configuraiton file |
| 10:08:27 | <HP_Archivist> | So you're trying to do batch URL saves? |
| 10:09:08 | <mgrandi> | i guess, basically i'm looping over them and using that library ^ |
| 10:09:52 | <mgrandi> | or at least automate them, without having me to go to the save page now site every time |
| 12:22:56 | | Jake5 (Jake) joins |
| 12:24:24 | | Jake quits [Ping timeout: 240 seconds] |
| 12:24:24 | | Jake5 is now known as Jake |
| 13:14:00 | <mgrandi> | so apparently doing `web.archive.org/save/https://www.instagram.com/lilsimsie/` is 520ing every time , but putting that url into the SPN page works fine? |
| 13:14:08 | <mgrandi> | ill try again tomorrow i guess |
| 14:15:51 | | Arcorann quits [Ping timeout: 265 seconds] |
| 14:42:51 | | prebala-publici quits [Client Quit] |
| 14:43:10 | | prebala-publici joins |
| 15:03:35 | <@JAA> | HP_Archivist: No idea. I know that there are some videos in the WBM, but I've always seen it as a '1 % of the time, it works every time' situation. |
| 15:04:13 | <@JAA> | mgrandi: There is a proper API for SPN2, but you'd have to ask for access via email I believe. |
| 16:06:27 | <atphoenix> | JAA, was your reply to HP_Archivist missing a word after the comma? |
| 16:25:51 | <@JAA> | atphoenix: Uh, no. |
| 16:26:45 | <atphoenix> | this is the part that was unclear: '1 % of the time, it works every time' |
| 16:27:45 | <@JAA> | https://www.youtube.com/watch?v=k8djas5pIvk |
| 16:39:28 | | rsn_ joins |
| 16:43:46 | | rsn quits [Ping timeout: 252 seconds] |
| 17:09:17 | <atphoenix> | ah, ok. I didn't catch the reference |
| 17:13:29 | <@JAA> | Looks like the WBM has issues with snapshots from around the 18th and 19th currently in general. I'm getting 504s and whatnot. |
| 17:24:10 | | OrIdow6^2 (OrIdow6) joins |
| 17:25:01 | | OrIdow6 quits [Ping timeout: 252 seconds] |
| 18:44:33 | | DogsRNice (Webuser299) joins |
| 20:32:34 | <mgrandi> | I'm fine with not having a full api, but this url is giving me a 500 every time , and then twitter urls have a 30% chance of being corrupted and not saving the actual twitter page |
| 20:54:13 | | Arcorann (Arcorann) joins |
| 21:36:45 | <mgrandi> | yeah, tried it again with that instagram link and it just 520s and says job failed |
| 21:38:28 | <@JAA> | mgrandi: SPN and SPN2 are completely separate systems. When you use /save/URL, that's SPN. When you 'put that URL into the SPN page', that's SPN2. |
| 21:39:08 | <@JAA> | SPN basically just retrieves that URL while writing to WARC. SPN2 runs a browser for the retrieval including scripts and whatnot. |
| 21:39:12 | <mgrandi> | i guess this library is using spn1 then...is there any way to have an 'api like interface' for spn2? like calling web.archive.org/save ? |
| 21:39:16 | <@JAA> | So yeah, they behave very differently. |
| 21:39:32 | <@JAA> | Twitter only works correctly on one of them as well. |
| 21:39:51 | <mgrandi> | (twitter has that issue with spn2 as well, or its only spn2, i'm not sure) |
| 21:40:41 | <@JAA> | Well yeah, the SPN2 API is available but not public. I guess you could reverse-engineer what happens in a browser, but I wouldn't recommend that. |
| 21:40:48 | <mgrandi> | ugh |
| 21:40:58 | <mgrandi> | well i don't really need a full api, the SPN1 one isn't really an api either |
| 21:41:36 | <@JAA> | I guess you could also script the interaction with /save/ with a headless browser or something. |
| 21:42:02 | <@JAA> | But the proper API is definitely easiest. |
| 21:42:38 | <mgrandi> | is it hard to get access to this? i'm like, saving on the order of 10 urls a day maybe |
| 21:46:23 | <@JAA> | No idea. Can't hurt to ask though. I'd assume that if you tell them that, they'd be fine with it. |
| 21:46:51 | <@JAA> | Can't imagine they want to lock that down since it's easier to monitor etc. than people trying to emulate the browser stuff. |
| 22:03:46 | | OrIdow6^2 is now known as OrIdow6 |
| 22:06:37 | | Arcorann quits [Ping timeout: 265 seconds] |
| 22:07:12 | <OrIdow6> | That video was captured by IA's Youtube collection (https://archive.org/details/youtubecrawl?sort=-publicdate) |
| 22:07:58 | <HP_Archivist> | JAA: No worries. Just surprised to have found that YT videos, some, are playable and captured in WBM |
| 22:10:19 | <@arkiver> | yeah they are playable if archived and indexed |
| 22:26:20 | <mgrandi> | who do i ask @JAA ? |
| 22:46:52 | <@JAA> | mgrandi: info@archive.org, cf. https://blog.archive.org/2019/10/23/the-wayback-machines-save-page-now-is-new-and-improved/ |
| 23:21:24 | | Doranwen quits [Ping timeout: 240 seconds] |
| 23:23:40 | | Doranwen joins |
| 23:23:41 | | Doranwen is now authenticated as Doranwen |
| 23:23:41 | | Doranwen quits [Changing host] |
| 23:23:41 | | Doranwen (Doranwen) joins |