| 09:16:26 | | wessel1512 quits [Read error: Connection reset by peer] |
| 09:17:15 | | wessel1512 joins |
| 15:52:14 | | atphoenix_ is now known as atphoenix |
| 18:57:43 | | @AlsoJAA quits [Ping timeout: 258 seconds] |
| 18:57:49 | | AlsoJAA (JAA) joins |
| 18:57:49 | | @ChanServ sets mode: +o AlsoJAA |
| 21:28:32 | | Iki quits [Ping timeout: 244 seconds] |
| 21:48:39 | | Iki joins |
| 22:50:43 | | Ryz (Ryz) joins |
| 22:50:50 | | Wolfin (wolfin) joins |
| 22:51:03 | <Ryz> | Any activity that happened here recently? |
| 22:54:51 | <@JAA> | No, this is a very low-traffic channel. |
| 22:54:57 | <@JAA> | Last messages were in early June. |
| 22:55:40 | <Ryz> | So to bring it up, wolfin was interested in ISP Hosting, provided some stuff from the tools scraped, did a test run with Bell userpages that I took up since there's a small amount |
| 22:55:46 | <Ryz> | Earlier in 2021 August |
| 22:56:13 | <Ryz> | ...Today's around the late 2021 August, both Bell and Sympatico ISP hosting stuff has been downers, now with 403s |
| 22:56:46 | <Ryz> | There was suggestion from JAA to run operations of grabbing these contents separately from #archivebot |
| 23:00:57 | <Ryz> | If there isn't outside infrastructure for wolfin to feed the links they've found using the various tools they have in possession, we may have to revert back to using ArchiveBot all the way~ S: |
| 23:06:17 | <Ryz> | The Sympatico links that wolfin processed are both https://p202.p0.n0.cdn.getcloudapp.com/items/E0uj862k/3113266d-4ada-4483-80a2-a48734f89fb4.txt and https://p202.p0.n0.cdn.getcloudapp.com/items/xQu6ZpRd/bc40a531-21f3-4f39-852c-21e430c17488.txt?v=4f529d164fa659530a2b60c495ef2183 |
| 23:06:24 | <Ryz> | Posting here incase they somehow come back |
| 23:06:48 | <Ryz> | We don't know when exactly the cut off happened as the attempt occured on 2021 August 07 or 2021 August 08 on the Bell ISP hosting websites |
| 23:07:43 | <Ryz> | And here's the one for Rootsweb (although I think it's in a tricky situation right now if hearing commotions about this earlier in August is to be noted): https://p202.p0.n0.cdn.getcloudapp.com/items/L1urx1Nj/a5f1d42c-feb5-482b-a8f3-46594fe84812.txt?v=83459d4f9f3f162fb9708484b7723883 |
| 23:13:32 | <@JAA> | There isn't any infrastructure here. This was used for discussion on discovery etc. in the past, I believe. It's been all but dead for years, similar to #effteepee etc. |
| 23:15:49 | <Ryz> | You did said it would be preferable to run this kind of thing outside of ArchiveBot, which I'm confused oo; |
| 23:18:06 | <@JAA> | Uh, I don't remember what I said, but yeah, it's a question of scale ultimately. ArchiveBot doesn't work well if you have thousands upon thousands of ISP-hosted sites. |
| 23:18:39 | <@JAA> | But short of a distributed project, I'm not sure we currently have anything that could handle that, either. |
| 23:19:19 | <Ryz> | I mean, they're pretty small in comparison to say archiving Blogspot or Wordpress websites~ |
| 23:19:46 | <Ryz> | Here is the quote you said JAA, "Yes please, archiving all of RootsWeb would be nice! Possibly outside of AB though with something that can do a 'recurse through everything on rootsweb.com or any subdomain' crawl starting with a list of a few pages including those mentioned last night." |
| 23:19:48 | <Wolfin> | I do have some access to servers; and can do runs with grab-site (and in fact am doing do with RootsWeb and ComicGenesis) to build WARCs. I'm unsure the amount of usefulness that'll provide, but at least there will be captures |
| 23:20:39 | <@JAA> | Ah right, RootsWeb, yes. |
| 23:20:53 | <Wolfin> | I'm...like a week into scrapping rootsweb at this point, you have to go slow with it. I feel like it's a crusty server in a closet somewhere. |
| 23:20:57 | <@JAA> | We have nothing that can do that at all at the moment, I think, short of plain wpull or wget. |
| 23:21:20 | <Wolfin> | Currently have... let's see, 400GB of WARCs. |
| 23:21:36 | <Wolfin> | About 1/4 to 1/3 of the way through |
| 23:21:46 | <@JAA> | What are you grabbing exactly? |
| 23:22:23 | <Wolfin> | https://p202.p0.n0.cdn.getcloudapp.com/items/L1urx1Nj/a5f1d42c-feb5-482b-a8f3-46594fe84812.txt?source=client -- presently using grab-site. |
| 23:23:09 | <Wolfin> | https://archive.org/details/rootswebwarc / https://archive.org/details/rootswebwarc2 / https://archive.org/details/rootswebwarc3 are the sets of grabs. |
| 23:23:23 | <Wolfin> | *the first sets, rather, bunch more coming. |
| 23:24:26 | <@JAA> | Mhm, I see. Probably won't grab everything, but it's a start. |
| 23:25:53 | <Wolfin> | Yeah, the starting place for this was all known sites already in the WBM, combined with using SERP from Google, Bing, Yahoo and Alexa, finally a 2 level crawl with linkchecker.py to find additional subdirs mentioned from the source set (and also not verify the subdirs still exist) |
| 23:27:17 | <Ryz> | Looking at https://p202.p0.n0.cdn.getcloudapp.com/items/L1urx1Nj/a5f1d42c-feb5-482b-a8f3-46594fe84812.txt further, |
| 23:27:20 | <Ryz> | I see there's two different types of subdomains, https://sites.rootsweb.com/ (like checking https://sites.rootsweb.com/~txcochra/ ) - and http://freepages.rootsweb.com/ (like checking http://freepages.rootsweb.com/~ourpast/history/index.html ) |
| 23:27:30 | <Ryz> | They seem to be separate subdomains |
| 23:27:37 | <Wolfin> | RootWeb/Ancestory had a data breach back a few years ago that would have listed every username (and hence every subdir for testing) but sadly I couldn't find a copy |
| 23:27:49 | <Wolfin> | Yep, they are distinct |
| 23:30:33 | <Ryz> | For http://freepages.rootsweb.com/ - at least sampling a few accounts, I can't seem to check by username like doing it as http://freepages.rootsweb.com/~marchington/ and not http://freepages.rootsweb.com/~marchington/history/index.htm |
| 23:30:49 | <Ryz> | So that's more difficult to get content from... |
| 23:34:28 | <Wolfin> | Yep, that's honestly where SERP is a lifesaver |
| 23:34:42 | <Wolfin> | Also, then tend to link to each other more, so there's an element of self-discovery |
| 23:34:56 | <@JAA> | Wolfin: It's on RaidForums and probably still available, but you need credits etc. |
| 23:35:53 | <Wolfin> | Ah, bleh, I don't want to PAY for illegally obtained good, that's a hard thing to legally explain |
| 23:35:59 | <@JAA> | Yeah |
| 23:39:12 | <Wolfin> | On the upside if any of the major search engines have seen content here, or WBM has found it before, or they linked to each other at some point, this will grab it eventually. So it's likely to be at least fairly comprehensive if not totally definitive. |
| 23:39:56 | <@JAA> | Well, not really though due to the wpull options grab-site uses. |
| 23:40:24 | <@JAA> | If you start a recursive crawl from https://example.org/foo/ and there's a link to https://example.org/bar/, that isn't followed. |
| 23:40:33 | <@JAA> | Or in this case, it won't recurse to other usernames. |
| 23:40:51 | <@JAA> | But it's actually worse because it might not even recurse fully within a username if there are crosslinks. |
| 23:41:48 | <@JAA> | If you start a recursive crawl from https://example.org/ and https://example.net/foo/ with the former linking to https://example.net/foo/bar.html and being retrieved before .net/foo/, bar.html will not be retrieved at all. |
| 23:42:12 | <Wolfin> | I did a little poking around there and I think as long as bare example.com is in the URL list, if /bar in your example was found it /would/ recurse. Unless there's another bit of logic in one of the two code bases stopping it. |
| 23:42:39 | <Wolfin> | Although you're right there's a lot of ways grab-site could muck up it's own to-be-crawled list internally. |
| 23:42:51 | <@JAA> | That's only true of the bare domain is the root of the entire crawl. If you start a recursive crawl from a list, each URL in the list is its own root. |
| 23:43:43 | <@JAA> | s/of/if/ |
| 23:44:24 | <Wolfin> | I was briefly tempted to try and write a more flexible crawler, there are arguably better ways to do this with such a large data set. But I wanted to at least capture what I could with the tools available rather than risk rootsweb going down while I play around on the weekends. |
| 23:45:15 | <Wolfin> | My ideal would use a command server that could spin up multiple digital ocean droplets and co-join a botnet-for-good (insert "muwhahaha" here) |