top of page

Data Privacy in an Age of Inference

New resources from LGBT Tech and the Justice Education Project examine why privacy must protect not only the information people disclose, but also what data systems claim to know about them.



In recent years, personal data has increasingly been treated as a resource to be extracted and monetized, rather than as information belonging to the person it describes. Clicks, searches, and location tracking are inputs that companies use to target advertising and generate profit. But what happens to privacy when the data used to describe you is not something you ever chose to share?


Privacy often goes hand-in-hand with consent—the ability to decide what information is shared, how, and with whom. In this model, consent is the barrier to accessing user data, meaning that a company can hold information about you only if you agree to provide it. Many privacy laws are built around this idea, giving people rights to access, correct, delete, or refuse certain collection and uses of their data.


However, data-inference models can bypass this framework. Rather than relying only on disclosed information, inference systems draw on indirect signals, such as behavior, location, app usage, and browsing activity, to predict identities, characteristics, or preferences. These inferences may then be used for tracking, targeting, automated decision-making, or further monetization, often without the user’s awareness.



To examine this growing gap, the Justice Education Project (JEP) and LGBT Tech developed a new policy brief on data inference, privacy, and the heightened risks these practices create for LGBTQ+ people and other marginalized communities.


On July 21, our organizations hosted a webinar on these findings, alongside an expert panel moderated by LGBT Tech fellow Laith Stevenson. Panelists included Eric Null, Director of the Privacy and Data Project at the Center for Democracy & Technology (CDT); Paige Collings, Senior Speech and Privacy Activist at the Electronic Frontier Foundation (EFF); and Sara Geoghegan, Senior Counsel and Director of the Consumer Privacy Program at the Electronic Privacy Information Center (EPIC).



When ordinary data reveals intimate information


Researchers at the University of Cambridge and Microsoft found that patterns of Facebook “Likes” could distinguish gay men from heterosexual men with 88 percent accuracy. The model could draw conclusions from digital behavior, including Likes that were not explicitly related to sexual orientation. In other words, the model did not need users to directly disclose their sexual orientation. Ordinary interaction with a platform could be enough to generate a prediction about one of the most personal aspects of their identity. (PNAS)


This finding can be paired with the fact that coming out is meant to be a process that lets a person disclose their identity on their own terms, at a time and to an audience of their choosing. Inferring someone’s sexual orientation or gender identity removes that choice, revealing or asserting the information regardless of whether the person ever decides to disclose it.


The Federal Trade Commission’s complaint against the data broker Gravy Analytics illustrates how this can play out. According to the complaint, a group used precise mobile geolocation data to identify by name a Catholic priest who had visited LGBTQ+-associated locations, exposing his sexual orientation and forcing him to resign. The complaint describes how location data can be combined with persistent identifiers and patterns of movement to identify individual people and reveal sensitive aspects of their lives. (Federal Trade Commission)


Similar issues arise in automated gender recognition, or AGR, which attempts to classify a person’s gender based on external characteristics. These systems may rely on facial structure and other aspects of gender expression that do not necessarily align with a person’s gender identity. Because gender identity is an internal sense of self, an automated classification can misgender a person while still being treated as an authoritative conclusion. When used for surveillance, screening, or access decisions, that mismatch can expose individuals to discrimination or other real-world harm. (Vrije Universiteit Amsterdam)



What privacy experts told us

During the July 21 webinar, panelists emphasized that inference is not simply a new technical feature. It changes the scale and character of longstanding privacy and surveillance risks.


Paige Collings of EFF described the shift plainly: “We’re in a different era.” Digital systems now capture more of daily life, allowing the same institutions and power structures that have historically engaged in surveillance or discrimination to draw on far larger and more detailed datasets. Collings emphasized that these harms are not identical for everyone: an LGBTQ+ person, immigrant, protester, abortion seeker, or person targeted by law enforcement may each face a different threat model.


Sara Geoghegan of EPIC explained why even seemingly insignificant information can become invasive when combined with other data. While “one data point likely isn’t terribly revealing,” a single location coordinate can take on a different meaning when a person visits the same medical facility repeatedly. Inference systems do not evaluate one click, coordinate, or search in isolation. They aggregate thousands of signals to create profiles that may reveal health conditions, relationships, beliefs, identities, or patterns of life.


Eric Null of CDT focused on how privacy laws can better account for that aggregation. When data is used to reveal a sensitive characteristic, Null argued that “the underlying raw data itself is entitled to the additional sensitive protections.” In other words, a law should not protect sexual orientation, health status, religion, or citizenship only after a company has labeled a person with that trait. It should also address the behavioral and location data used to produce the conclusion.


Together, the panelists identified several related policy needs: stronger data-minimization rules, clearer protection for inferences and profiles, meaningful rights to contest or delete inferred information, safeguards against government purchase of sensitive commercial data, and stronger enforcement when inference leads to outing, stalking, discrimination, or other serious harms. Privacy rights must address the act of inference.



Current privacy laws may allow individuals to learn what information a company holds about them, correct inaccuracies, request deletion, or decline certain future collection. These pathways are important, but they often focus on data that has already been collected or recorded. The act of inference itself—and a system’s continuing ability to regenerate a sensitive conclusion from new or retained data—is addressed inconsistently.


A correction or deletion may change a specific record without changing a system’s ability to draw the same conclusion again. A person could delete a label identifying them as LGBTQ+, for example, while the company retains the browsing, location, purchasing, or social data needed to reproduce it.


The Justice Education Project and LGBT Tech’s brief therefore calls for regulating inference as an act, rather than focusing only on the eventual output or sale. It recommends giving people a meaningful, categorical way to opt out of sensitive inferences and extending deletion rights to cover re-deriving an inference, not only deleting the existing label or the data product in which it appears.


These protections should be paired with data minimization. The less unnecessary information companies collect and retain, the less material they have to build intimate profiles that individuals never agreed to create.


Without stronger protections, inference can decide a person’s coming out before they get the chance to decide for themselves. What is ultimately at stake is the same right coming out has always represented: control over what is known about your identity, who gets to know it, and when it is shared.



The full JEP x LGBT Tech policy brief can be found here.


A recording of the webinar can be viewed below.




 
 
bottom of page