00:01:56dm4v quits [Read error: Connection reset by peer]
00:02:11dm4v joins
00:02:13dm4v quits [Changing host]
00:02:13dm4v (dm4v) joins
00:05:59Arcorann (Arcorann) joins
00:07:34BlueMaxima joins
00:16:46flashmeow quits [Read error: Connection reset by peer]
00:19:10flashmeow (flashmeow) joins
00:42:57tzt quits [Ping timeout: 258 seconds]
01:03:52dm4v quits [Ping timeout: 250 seconds]
01:07:45dm4v joins
01:07:47dm4v quits [Changing host]
01:07:47dm4v (dm4v) joins
02:42:56Ajay_m quits [Read error: Connection reset by peer]
02:44:44Ajay_m joins
02:49:04Ajay_m quits [Read error: Connection reset by peer]
02:57:54TheTechRobo quits [Remote host closed the connection]
03:15:28tzt joins
03:37:22qw3rty__ joins
03:41:12qw3rty_ quits [Ping timeout: 258 seconds]
03:52:19Ajay_m joins
03:56:32wizards_ joins
03:59:48wizards quits [Ping timeout: 250 seconds]
04:33:43Ajay_m quits [Ping timeout: 258 seconds]
04:53:58fuzzy8021 quits [Ping timeout: 250 seconds]
04:57:35DogsRNice_ quits [Read error: Connection reset by peer]
05:13:22Larsenv (Larsenv) joins
05:13:29fuzzy8021 (fuzzy8021) joins
05:16:54Ajay_m joins
05:46:06atphoenix_ is now known as atphoenix
05:47:07sec^nd quits [Ping timeout: 255 seconds]
05:59:19sec^nd (second) joins
06:37:32Ajay_m quits [Ping timeout: 258 seconds]
06:52:13Ajay_m joins
08:02:54Ajay_m quits [Read error: Connection reset by peer]
08:23:00spirit joins
08:44:47Ajay_m joins
08:49:21BlueMaxima quits [Client Quit]
10:08:42Terbium quits [Quit: http://quassel-irc.org - Chat comfortably. Anywhere.]
10:09:08Terbium joins
10:34:03HP_Archivist quits [Ping timeout: 258 seconds]
10:35:00Matthww8 quits [Read error: Connection reset by peer]
10:35:22Matthww8 joins
12:14:42grawity quits [Remote host closed the connection]
14:51:52spirit quits [Client Quit]
15:02:17TheTechRobo joins
15:03:14TheTechRobo quits [Remote host closed the connection]
15:06:42Arcorann quits [Ping timeout: 250 seconds]
15:40:43Iki1 quits [Read error: Connection reset by peer]
15:48:18dm4v quits [Read error: Connection reset by peer]
15:48:31dm4v joins
15:48:33dm4v quits [Changing host]
15:48:33dm4v (dm4v) joins
16:38:13c00k13 quits [Ping timeout: 258 seconds]
16:43:58c00k13 joins
16:51:42spirit joins
17:55:32HP_Archivist (HP_Archivist) joins
19:11:17AlsoHP_Archivist joins
19:15:00HP_Archivist quits [Ping timeout: 258 seconds]
19:15:52AlsoHP_Archivist quits [Ping timeout: 250 seconds]
19:16:42AlsoHP_Archivist joins
19:27:02teej (teej) joins
19:52:23spirit quits [Client Quit]
19:53:54<teej>Hello JAA. If you're available, what are your thoughts about the GitHub Copilot experiment and its use of copyleft source code?
20:09:23<teej>References:
20:09:23<teej>https://github.blog/2021-06-29-introducing-github-copilot-ai-pair-programmer/
20:09:23<teej>https://docs.github.com/en/github/copilot/research-recitation
20:10:01<teej>Discussions:
20:10:01<teej>https://news.ycombinator.com/item?id=27724042
20:10:01<teej>https://drewdevault.com/2021/07/04/Is-GitHub-a-derivative-work.html
20:10:01<teej>https://zephyrtronium.github.io/articles/copilot.html
20:11:14<teej>Anyone can comment on this, not only JAA.
20:14:49<teej>There's also https://news.ycombinator.com/item?id=27676266. That one is pretty long.
20:20:52aleph quits [Read error: Connection reset by peer]
20:20:53aleph joins
20:25:13<teej>Actually...
20:25:38<teej>This is very relevant to archiving... because a lot of people on GitHub will jump ship.
20:28:57<teej>Example: https://twitter.com/nixcraft/status/1411438221918150664
20:37:40<teej>Also watch out for new forks of Audacity.
20:54:17AlsoHP_Archivist quits [Ping timeout: 258 seconds]
20:55:04AlsoHP_Archivist joins
20:56:33<Ryz>https://www.theatlantic.com/science/archive/2021/07/gamers-are-better-scientists-catching-fraud/619324/
20:56:40AlsoHP_Archivist quits [Client Quit]
20:56:41<Ryz>"Why Scientists are Worse Than Gamers" - madeup clickbait title
20:56:59<teej>Thanks Ryz. I'll check it out.
21:12:34<teej>Ryz: That's interesting. I think the comparison between cheating in games and fabricating scientific data is somewhat unjustified. There are just way too many research papers being published.
21:15:06<Ryz>In the article, there have been proposed tools that would automatically or help people find more fraudulent pieces, but because money is often involved, people don't want to dive into that and risk having their own stuff be checked s:
21:18:32<teej>Ryz: I'm not convinced that tools, at least with the current state of technology, can accurately detect fabricated data from scientific research papers, let alone novel research.
21:22:22<Ryz>It would have to start somewhere, even if it's small, since if what the article says is true, the people who are supposed to check the research and findings aren't doing their job properly either s:
21:23:17<@JAA>> Hosts code on GitHub, a commercial platform with terms that grant the company a licence to do countless things with the code.
21:23:21<@JAA>> GitHub uses it.
21:23:22<teej>Hmm, scientific publishers don't make it easy to access paid content.
21:23:24<@JAA>surprised_pikachu.png
21:25:30<Jake>teej: interesting legal perspective here: https://decoded.legal/blog/2021/06/github-copilot-initial-thoughts-from-an-english-law-perspective
21:25:31<teej>JAA: Lol. It's not that simple. A lot of people can take code repositories from other places and easily upload it onto GitHub.
21:26:33<Jake>I imagine when uploading to Github you say it's your code, legally, anyways.
21:26:38<@JAA>Sure, but then the code would be released under a licence that allows for such redistribution (assuming there's nothing illegal going on).
21:29:22<@JAA>Jake: GitHub's terms account for that. 'If you're posting anything you did not create yourself or do not own the rights to, you agree that you are responsible for any Content you post; that you will only submit Content that you have the right to post; and that you will fully comply with any third party licenses relating to Content you post.'
21:29:34<Jake>yup.
21:30:51<teej>Jake: This is a long article. It may take me a little to read through it.
21:40:40<teej>JAA: The article that Jake linked makes an interesting point about whether the GitHub Copilot's code is an act of "copying". GitHub themselves state "We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set."
21:41:24<@JAA>Well, yeah, I guess there are two completely distinct issues here: is it fine for GitHub to use the published code like that, and what happens when someone actually uses the resulting code?
21:42:04<@JAA>I'd say that the answer to the former is a clear yes given the open-source licences and GitHub's terms.
21:42:08<@JAA>The latter is ... complicated.
21:43:05lorwp quits [Quit: ZNC - https://znc.in]
21:43:57lorwp (lorwp) joins
21:45:34t3chler quits [Quit: https://quassel-irc.org - Chat comfortably. Anywhere.]
21:46:02<thuban>another issue, as mentioned briefly in another hn thread (https://news.ycombinator.com/item?id=27710287) is the incorporation of (presumably irremovable) pii
21:47:31<@JAA>Yeah, that's a good point.
21:48:02<@JAA>Or secrets like API keys.
21:48:15t3chler joins
21:49:56<teej>I also don't think GitHub's use copyleft code to train AI/ML algorithms is a concern. I think that the product that Microsoft/GitHub creates that allows them to use parts of copyleft code generated by an algorithm in their own private codebase (internal Microsoft/GitHub code) is controversial, as it incentives a loophole around copyleft licenses.
21:51:12<teej>That builds on what is said on Drew DeVault's blog post (https://drewdevault.com/2021/07/04/Is-GitHub-a-derivative-work.html).
21:51:50<thuban>^^ that's more of a practical problem for users than a legal problem for github
21:54:55<teej>JAA: I remember seeing a Twitter post showing GitHub Copilot autocompleting an SSH private key, but the user commented that it was just a joke.
21:58:11<@JAA>Heh
21:59:18<@JAA>Well, I didn't have to look far: https://twitter.com/pkell7/status/1411058236321681414
22:01:49<teej>JAA: Isn't that probably because someone accidentally committed secrets?
22:01:58<@JAA>Sure, so?
22:02:36<teej>Oh, I was just making sure. I wasn't making a point.
22:02:57<@JAA>Right
22:03:22<@JAA>Yeah, someone (or multiple people) committed code like that, the AI was trained with it, and now it suggests it to other people.
22:03:25<@JAA>Free API keys. :-)
22:03:34<teej>Lol.
22:04:25<@JAA>Some people are saying that any keys or email addresses appearing in code returned by Copilot are generated, but I definitely wouldn't trust that. lol
22:05:26<teej>It's generated (obviously?) by an algorithm that took it from actual code and copied it.
22:06:26<@hook54321>does it use code from private repos too?
22:07:07<@JAA>I don't think so.
22:07:18<teej>hook54321: I would assume not, unless Microsoft really wants to get sued.
22:07:30<thuban>JAA: it definitely generates real github and twitter links (https://twitter.com/kylpeacock/status/1410749018183933952#m), so i'm pretty sure it can handle email addresses :)
22:08:09<@hook54321>part of the problem would be solved by people not committing private API keys
22:08:13<@JAA>I'm wondering whether they filtered the training set for code licences. There's a whole bunch of non-open-source code in public repos on GitHub because people aren't adding a LICENSE file.
22:08:48<@hook54321>that's my biggest concern i think
22:09:34<@hook54321>and attribution, etc.
22:09:47<@JAA>Yeah
22:12:16<teej>I still don't know if Copilot will reproduce binary data.
22:12:18<@JAA>thuban: To clarify, I mean 'I wouldn't trust it to work 100 % of the time'.
22:12:29<@hook54321>also, if it automatically adds an attribution of their github username, i wonder how it'll go if someone changes it.
22:13:21<teej>Binary data encoded in non-binary form.
22:13:56<@JAA>I mean, it's not like they can do that. The code is synthesised, not copied from samples. Most likely, they have no way to even tell which code samples influenced a particular piece of generated code.
22:15:13<teej>I can envision a future license that will prohibits Microsoft from doing this AI/ML code generation.
22:16:12<teej>JAA: Not copied? GitHub states "We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set."
22:16:37<@JAA>Yeah, so 99.9 % of the time, it isn't copied.
22:16:52<teej>That's misleading.
22:17:00wickedplayer494 quits [Ping timeout: 250 seconds]
22:17:20<@JAA>Yeah, sort of, but you know what I mean.
22:17:23<teej>It could be 1000 times...
22:17:37<@JAA>Most of the code generated by it is not simply a copy of an input sample.
22:17:46<@hook54321>teej: afaik, github can kinda legally use the code for this feature, but people can't use the feature without violating some license. so it's already practically not allowed.
22:17:46wickedplayer494 joins
22:18:00<@JAA>And therefore attribution to an input sample is very difficult to impossible.
22:18:05<teej>Jake: The Quake source code output proves my point.
22:18:20<@JAA>hook54321: Try proving that someone used it though. :-/
22:19:01<@hook54321>i don't think i hate the idea itself, but it's not practical.
22:19:06<@hook54321>JAA: true
22:19:49<@hook54321>IF it could attribute people, it could be fine, i guess?
22:20:18<teej>hook54321: And it's legal if they tweaked the algorithm to produce copyleft code more than 0.1% of the time too?
22:21:00<teej>hook54321: Tweaked and used for themselves.*
22:21:21<teej>To bypass GPL restrictions.
22:21:48<@JAA>It would also have to prevent the combination of samples with incompatible licences.
22:21:48<@hook54321>not sure what you mean bypass
22:21:52<@hook54321>yeah
22:22:28<@hook54321>i guess they could train separate APIs for different licenses
22:23:13<@hook54321>excluding any repos without a license
22:23:27<@hook54321>*AIs
22:26:25Mateon1 quits [Remote host closed the connection]
22:26:29Mateon1 joins
22:27:01<teej>hook54321: What I mean to say is that I can imagine a large corporation (Microsoft) abusing an AI/ML algorithm trained on a lot of copyleft code to generate verbatim copies of parts of copyleft code to be used in private repositories without having to release the private source code by using the algorithm as an excuse.
22:28:28<teej>Microsoft is also trying to add a lot of Linux-like features into Windows, for example.
22:33:26superkuh quits [Remote host closed the connection]
22:33:27<teej>hook54321: Yes, I can imagine Microsoft would train the AI/ML algorithm on their own Windows codebases separately. Otherwise, users can potentially leak parts of Windows source code.
22:33:42superkuh joins
22:34:54<teej>To be quite honest, I'm not exactly sure why Visual Studio Code got so popular so quickly.
22:36:05<teej>Is it because of the plugins?
22:36:46<teej>Does anyone here personally use Visual Studio Code?
22:37:29<Jake>I've used VSCode in the past for larger projects, I use gedit for a majority of my 'smaller' projects.
22:38:36<teej>Jake: Did you install plugins?
22:39:07<Jake>I think I installed one for better GitHub integration, and maybe a few for language highlighting, but nothing else
22:42:12<@JAA>Pff... butterflies.
22:44:22<teej>Oh, okay.
22:45:42<teej>I would recommend VSCodium as an alternative if using VS Code is necessary.
22:48:51<teej>JAA: What does "butterflies" refer to?
22:49:04<@JAA>https://xkcd.com/378/
22:50:07<teej>LOL!
22:50:37<teej>That made me laugh.
23:11:22BlueMaxima joins
23:21:29Mateon1 quits [Remote host closed the connection]
23:21:46Mateon1 joins