Are you in the Weights is a specialized web application that allows individuals to check their inclusion in the training datasets of large language models. Created by Thomas Dimson and Joey Flynn, this tool serves anyone curious about their digital footprint in AI systems, providing a unique window into how personal data influences machine learning. The core value lies in offering transparency about AI training data composition, giving users insight into their representation within these powerful systems that increasingly shape our digital experiences.
Many people are unaware of whether their public information has been used to train AI models that power everything from search engines to creative tools. This lack of transparency creates uncertainty about digital identity and privacy in the age of artificial intelligence. The product addresses this by providing concrete verification of one's presence in LLM training data, helping users understand their relationship with emerging technologies that may be using their information without explicit consent or awareness.
The primary feature is the individual verification system that checks names against LLM training datasets. When users input their name, the system scans through the model weights to determine if their data was included in training. This verification process provides a binary result along with a strength score indicating how prominently the individual appears in the model's knowledge base. The feature works by leveraging the internal representations within the language model to detect patterns associated with specific individuals.
Another key capability is the strength scoring system that quantifies an individual's prominence within the model. Each verified person receives a numerical strength rating between 982-993, as demonstrated by the leaderboard showing Lionel Messi at 993 and various celebrities at 986. This scoring mechanism provides granular insight beyond simple presence/absence, indicating how significantly a person's data contributes to the model's knowledge structure and output generation capabilities.
The platform includes a comprehensive leaderboard feature that ranks individuals by their strength scores, creating a competitive and exploratory element. Users can browse through top-ranked individuals like Hakeem Olajuwon, Greta Thunberg, and Ellen DeGeneres to understand the types of people most prominently represented. This public ranking system allows for comparative analysis and helps users contextualize their own scores within the broader population of individuals captured in LLM training data.
admin
The verification workflow begins with a browser-based security check using Cloudflare Turnstile to prevent automated queries. After successful verification, users can search for any name to check its presence in the weights. The system processes queries by comparing input names against the model's internal representations, returning both a confirmation of presence and a quantitative strength score. This straightforward process makes advanced AI transparency accessible to non-technical users through a simple web interface.
Concrete use cases include journalists verifying source representation in AI systems, public figures monitoring their digital presence, and researchers studying bias in training data. A journalist might use the tool to confirm whether sources from underrepresented communities appear in major LLMs, while a celebrity could check how prominently they're represented compared to peers. Researchers benefit from quantitative data about which demographics and professions receive the most representation in AI training datasets.
The tool targets individuals concerned about AI ethics, digital privacy advocates, journalists, researchers, and public figures. Accessible through any modern web browser, the platform requires no technical expertise to use. While specific pricing details aren't provided, the current implementation appears to be freely accessible. The takeaway is that this tool provides unprecedented transparency into AI training data composition, empowering users to understand their relationship with increasingly influential language models.
Individuals concerned about AI ethics, digital privacy advocates, journalists investigating AI systems, researchers studying training data bias, and public figures monitoring their digital representation. The tool serves anyone curious about their inclusion in large language model training datasets and their relationship with emerging AI technologies.
Updated 2026-06-21