Mastodon Feed: Post

Mastodon Feed

jonny@neuromatch.social ("jonny (nonvenomous)") wrote:

RE: https://mathstodon.xyz/@andreasthom/117240535270608201

We can assume unreleased models under active development are being trained and retrained constantly, that's what developing them means, so they are exposed to new data that enters the training set. The high capacities of these models mean that they are absolutely capable of reproducing unique training data nearly exactly. In a research setting like this where the model provider is itself doing the research we can also assume they can disable any filters or safeguards that might normally discourage/prevent output that looks like leaking training data in pursuit if a goal.

With all that, the distinction between "user data being entered into the generating context" and "user data being entered into the training data" is not really meaningful.

So it seems like if your research is identifiable/tagged to a named result as some of these proofs are, or really in any condition where you'd expect your work being in training data would meaningfully impact the output for a question, the question of whether an AI company swoops in to steal your work seems like more or less a "will they" as a cost/benefit per the prestige of the problem, rather than a "can they"