Boosted by jonny@neuromatch.social ("jonny (nonvenomous)"):
andreasthom@mathstodon.xyz ("Andreas Thom") wrote:
2/3 OpenAI is drawing a distinction that its answer to me erased, despite the fact that my question explicitly made that distinction.
We are not required to reverse-engineer OpenAI’s internal training pipeline to establish what happened. Only OpenAI has the relevant data for that. If Sellke and Bubeck give a categorical denial, it must disclose the basis: product and privacy settings, relevant datasets and checkpoints, and what “de-identified data derived from usage” means.
I disabled model training on 29 June. That control is still only a promise whose implementation users cannot audit, and it is prospective: it does not answer what happened to earlier conversations or to derivatives already selected. OpenAI’s answer did not mention the setting or any account-specific check.
There is also a mathematical parallel, though a less explicit one. The approach of Kun and myself was not the main line of attack on non-soficity, in fact there were other more promising approaches along the line of quantum games etc. OpenAI’s detailed command of our techniques therefore made me wonder how the model found this route. If nonpublic research supplied by users improved a model and the provider then used that model to race those users to publication - without informed consent, disclosure, or credit - that would be ethically indefensible. De-identification may remove a name; it does not remove the intellectual content of a mathematical idea. Sellke and Bubeck seem to be blind to this simple moral aspect.