A Reddit discussion highlights that OpenAI's opt-out terms for paid users may not protect internal reasoning tokens, potentially allowing the company to use this data for model training. The argument centers on the ambiguity of whether these hidden tokens constitute "Output" under the service agreement, given that users do not receive them directly.
- Paid user opt-out policies typically protect "Output," but reasoning tokens are not visible to the consumer.
- Reasoning tokens contain processed versions of user input and responses, making them valuable for synthetic data generation.
- Unlike OpenAI's current terms, Anthropic's agreement does not require the user to receive the output for opt-out protections to apply.
- This loophole could allow OpenAI to incorporate reasoning data into pretraining or midtraining mixes without violating RL optimization restrictions.
This distinction is significant because it suggests that opting out of training might be ineffective if internal model states are still utilized for improving future versions.