| 00:02:34 | | jamesatjaminit quits [Ping timeout: 250 seconds] |
| 00:44:50 | | jamesatjaminit (jamesatjaminit) joins |
| 01:43:06 | | jamesatjaminit quits [Ping timeout: 250 seconds] |
| 01:59:21 | | jamesatjaminit (jamesatjaminit) joins |
| 02:05:12 | | jamesatjaminit quits [Ping timeout: 250 seconds] |
| 02:29:42 | | jamesatjaminit (jamesatjaminit) joins |
| 02:58:04 | | jamesatjaminit quits [Ping timeout: 250 seconds] |
| 03:01:48 | | qw3rty_ joins |
| 03:02:17 | | jamesatjaminit (jamesatjaminit) joins |
| 03:05:11 | | qw3rty__ quits [Ping timeout: 244 seconds] |
| 03:59:49 | | qw3rty__ joins |
| 04:03:30 | | qw3rty_ quits [Ping timeout: 250 seconds] |
| 08:00:09 | | britmob256 quits [Quit: britmob256] |
| 08:24:22 | | rsn_ joins |
| 08:26:32 | | rsn quits [Ping timeout: 250 seconds] |
| 12:04:22 | | britmob256 joins |
| 13:01:48 | | rsn joins |
| 13:02:58 | | rsn_ quits [Ping timeout: 244 seconds] |
| 16:19:20 | | Matthww80 joins |
| 16:20:46 | | Matthww8 quits [Ping timeout: 252 seconds] |
| 16:20:46 | | Matthww80 is now known as Matthww8 |
| 17:45:26 | | Larsenv quits [Quit: ZNC 1.8.2+deb1+focal2 - https://znc.in] |
| 17:47:14 | | luckcolors quits [Quit: No Ping reply in 180 seconds.] |
| 17:48:16 | | Larsenv (Larsenv) joins |
| 17:48:32 | | luckcolors (luckcolors) joins |
| 20:46:23 | <spirit> | is the robots.txt parser broken or am i reading it wrong? http://web.archive.org/web/2/https://www.airbnb.de/rooms/12345 is "This URL has been excluded from the Wayback Machine." but https://www.airbnb.de/robots.txt has no such rule |
| 20:47:15 | <OrIdow6> | Exclusions often happen by them sending an email or other notice to the Internet Archive |
| 20:47:49 | <OrIdow6> | I thought I once heard that there was a way to tell where the exclusion came from, but if so I have forgotten |
| 20:47:51 | <spirit> | is there a transparency log of that? |
| 20:47:59 | <OrIdow6> | Not that I know of |
| 20:48:34 | <OrIdow6> | Presumably a lot of the people doing so want to keep information private/non-public, so publicizing it would sort of defeat the purpose |
| 20:54:10 | <spirit> | true |
| 20:54:23 | <spirit> | but for big filters of big sites it might be nice |
| 20:54:37 | <spirit> | in this case i think it just uses the two Allow rules though |
| 20:54:57 | <@JAA> | There's a wiki list for that: https://wiki.archiveteam.org/index.php/List_of_websites_excluded_from_the_Wayback_Machine |
| 20:55:38 | <@JAA> | This error has nothing to do with robots.txt. |
| 20:56:44 | <@JAA> | Looks like https://www.airbnb.de/rooms/ was excluded that way. |
| 21:00:38 | <spirit> | ah |
| 21:00:47 | <spirit> | right the message mentions robots.txt otherwise iirc |
| 23:54:39 | | OrIdow6 quits [Ping timeout: 258 seconds] |