Google has expanded the scope of data it collects to train its artificial intelligence models, now incorporating media uploaded by users across several of its primary search-related services.
The policy change, Engadget reported Monday (July 6), was implemented without much public fanfare and allows the technology giant to use images, audio, video and other files submitted through tools such as Google Lens and Google Translate.
Google’s move highlights the demand for high-quality datasets as generative AI developers confront a scarcity of fresh information to feed their large language models.
Under the updated terms, any photo uploaded to Google Lens for visual identification or audio captured during a voice-activated search may be harvested for training purposes. The data collection also extends to any files processed through Google Translate, encompassing “images, files and audio and video recordings,” according to the report.
For professionals in the digital economy and banking sectors concerned with data privacy or corporate security, it is notable that users are automatically opted into this training program. Engadget, citing earlier findings by TechCrunch, notes that the current policy is restricted to search-related products; personal repositories such as Google Photos are currently excluded from this specific training data sweep.
As generative AI seeks new data sources, Google has provided a manual mechanism for users to restrict their data from being used in this manner. To opt out, users must navigate to their dedicated Search Services History page to uncheck the “Save Media” box. Additionally, users are advised to review their Search Services Personalization settings to ensure no further media is being retained for AI training.
For those seeking to limit their interaction with Google’s AI outputs entirely, the report also highlights a technical workaround: appending “-AI” to a search query will effectively remove AI-generated overview results from the interface.
The shift underscores a broader trend among Big Tech firms seeking to leverage proprietary user interactions to maintain a competitive edge in the AI race, even as questions regarding user permission and data ownership persist. Google itself highlighted this trend earlier this year, when the company pressured news organizations to allow its AI to train on their articles or risk losing the annual payment for being featured in Google News.