Today’s generative AI models require massive quantities of data, as AI ``scaling laws’’ suggest that more data will make models better. The largest, cheapest source of such data is, by far, the internet. This has led to AI model creators deputizing web scrapers to harvest massive amounts of internet data, including image databases, news sites, web forums, and beyond. Unsurprisingly, this use of scraped data - both for training AI models and for real-time retrieval via retrieval-augmented generation (RAG) - is contentious and has led to numerous lawsuits and policy initiatives.
Regulatory approaches to this problem could take years to resolve, so new technical tools are urgently needed to make the use of online data in AI models more transparent, fine-grained, and accountable. Such tools can help alleviate the current tensions between content creators and AI model operators and ground future policy decisions. In light of this, the ARGUS Lab’s research in this space strives (1) to shed light on how data is used in AI systems and (2) build practical tools that empower data hosts and rightsholders to choose how (or if) their data is used in AI models.