Data governance has its own language. Learn the terms that matter most.
Automated decision-making technology (ADMT) is any system that uses computation to replace or substantially replace human judgment in a decision that affects a person, such as decisions about lending, housing, employment, or healthcare.
AI data lineage is the end-to-end tracking of where data used to train, fine-tune, or prompt an AI model came from, how it was transformed, and what outputs it influenced.
AI governance is the set of policies, controls, and oversight mechanisms an organization puts in place to ensure its artificial intelligence systems are developed and deployed responsibly, legally, and in alignment with business and ethical standards.
AI risk management is the organizational practice of identifying, assessing, and mitigating risks arising from the development and deployment of artificial intelligence systems.
Agentic AI refers to AI systems capable of taking autonomous actions, including browsing the web, executing code, querying databases, sending communications, or chaining multi-step tasks, without requiring human intervention at each step.
Article 28 of the GDPR governs the relationship between data controllers and data processors, requiring that any third party processing personal data on a controller's behalf do so under a written contract with specific data protection obligations.
Article 30 of the GDPR requires organizations to maintain a Record of Processing Activities (ROPA), a documented inventory of every way they process personal data.
Article 32 of the GDPR requires organizations to implement appropriate technical and organizational measures to ensure a level of security appropriate to the risk of processing personal data.
Article 6 of the GDPR establishes the six lawful bases for processing personal data, the legal grounds an organization must rely on to process data lawfully.
Article 9 of the GDPR prohibits the processing of special categories of personal data, including health, biometric, racial, and political information, unless specific conditions are met.
Audience activation is the process of taking customer or prospect data, typically from a first-party database or data platform, and pushing it to advertising, marketing, or personalization systems to execute targeted campaigns and experiences.
An authorized agent is a person or organization that a consumer designates to submit privacy rights requests on their behalf under laws like the CCPA.
Automated decision-making refers to decisions made solely by algorithmic or AI systems, without meaningful human review, that have significant effects on individuals, such as determining credit, employment, insurance, or content access.
The California Consumer Privacy Act (CCPA) is a US state privacy law that gives California residents the right to know, access, delete, and opt out of the sale of their personal information.
The California Privacy Protection Agency (CPPA) is the state body responsible for enforcing the California Consumer Privacy Act and California Privacy Rights Act, with authority to investigate violations and issue fines.
The California Privacy Rights Act (CPRA) is a 2020 ballot measure that expanded and strengthened the CCPA, adding new consumer rights, stricter obligations for businesses, and creating California's first dedicated privacy enforcement agency.
Passed on July 8, 2021, the Colorado Privacy Act (CPA) was the third state-based privacy law in the US with enforcement beginning on July 1, 2023.
In privacy law and data governance, consent is the mechanism by which an individual gives a company permission to collect, process, or share their personal data for a specified purpose.
Consent orchestration is the process of capturing, storing, and automatically propagating individual consent and preference signals to every system that processes data, ensuring that what a person agrees to (or refuses) at the point of collection is enforced everywhere their data travels.
A consent signal is a machine-readable communication of an individual's data privacy preference, indicating whether they have opted in or out of specific data uses, transmitted from the point of capture to the systems that need to act on it.
Consent-based personalization is the practice of delivering individualized customer experiences, including tailored content, product recommendations, and messaging, using only data that the customer has explicitly permitted to be used for that purpose.
Cookieless advertising refers to digital advertising strategies and technologies that reach and measure audiences without relying on third-party cookies, driven by browser restrictions, regulatory requirements, and growing privacy expectations from consumers.
Customer data is any information an organization collects, infers, or receives about the individuals who buy from, use, or interact with it, spanning everything from contact details and purchase history to behavioral data and inferred attributes.
A Customer Data Platform (CDP) is software that unifies customer data from multiple sources into persistent, individual-level profiles that are made available to other systems for marketing, personalization, and analytics.
Dark patterns in privacy are user interface design techniques that manipulate or deceive users into sharing more data, accepting tracking, or waiving privacy rights they would not otherwise choose to waive.
A data processor is an entity that processes personal data on behalf of a data controller, acting under the controller's instructions.
A data catalog is a searchable inventory of an organization's data assets, including where they live, what they contain, how they're structured, and who owns and uses them.
A data clean room is a secure, privacy-preserving environment that allows two or more organizations to match and analyze overlapping datasets without either party exposing their raw data to the other.
A data controller is the entity, typically a company, that determines the purposes and means of processing personal data.
A data inventory is a structured record of all personal or sensitive data an organization holds, covering what types of data exist, where they're stored, how they're collected, how they're used, and who has access.
Data lineage is the tracked history of how data moves through an organization's systems, documenting its origin, the transformations it undergoes, and every system it passes through on the way to its final destination.
Data mapping is the process of documenting where personal or customer data lives across an organization's systems, how it flows between those systems, and what it's used for.
Data minimization is the principle, and in many jurisdictions a legal requirement, that organizations should collect only the personal data that is necessary for a specific, defined purpose, and not retain it longer than that purpose requires.
Data permissions are the rules that govern whether a specific dataset, record, or field can be used for a specific purpose, encoding who can access data, for what reason, and under what conditions.
A data pipeline is an automated sequence of processes that moves, transforms, and loads data from one or more sources into destination systems, powering analytics, AI models, reporting, and operations.
Data provenance is the documented history of a dataset, covering where it originated, who collected it, how it was transformed, and what permissions govern its use.
Data quality refers to the degree to which data is accurate, complete, consistent, timely, and fit for its intended purpose, a foundational requirement for reliable analytics, AI, and regulatory compliance.
A data silo is a repository of data controlled by one department or system that is not accessible or visible to the rest of the organization, creating fragmentation, duplication, and governance blind spots.
A data subject is an identifiable individual whose personal data is being collected or processed, in most commercial contexts a customer, user, or employee.
A data subject access request (DSAR) is a formal request from an individual to a company to receive a copy of the personal data the company holds about them, along with information about how it's being used.
Data subject rights automation is the use of software to receive, verify, process, and fulfill individual privacy rights requests, including access, deletion, correction, and opt-out, without requiring manual handling for each request.
A data use policy is an organization's formal statement of what data it collects, for what purposes, under what conditions, and with what restrictions, governing both how the organization uses data internally and what uses are permissible for third parties receiving that data.
"Do Not Sell" and "Do Not Share" are opt-out rights under California's CCPA and CPRA that give consumers the ability to stop businesses from selling or sharing their personal information with third parties.
Do Not Track (DNT) is a browser signal that users can enable to indicate they don't want their online activity tracked across websites, though most sites are not legally required to honor it.
Do Not Train is a data use control that prevents customer or user data from being used to train, fine-tune, or improve an AI model, a right increasingly demanded by regulators, enterprise contracts, and individuals whose data is in AI pipelines.
Downstream enforcement is the practice of ensuring that data use restrictions, including consent, opt-outs, and policy controls, are applied not just at the point of data collection but in every system that subsequently processes or uses that data.
Facebook Limited Data Use (LDU) is a data processing control that allows businesses to restrict how Meta uses data it receives from website events for users who have opted out of data sale or sharing under California privacy law.
First-party data is information an organization collects directly from its own customers, users, or audience through its own channels and with direct consent, making it the most reliable, privacy-compliant, and commercially durable data asset available.
A foundation model is a large AI model trained on broad, general-purpose datasets that serves as a base to be adapted, through fine-tuning or prompting, for a wide range of downstream tasks.
Under the GDPR, personal data is any information that relates to an identified or identifiable individual, a definition broad enough to cover not just names and contact details but IP addresses, device identifiers, behavioral data, and inferred attributes.
The GDPR's seven data protection principles, set out in Article 5, form the ethical and legal foundation of the regulation, governing how all personal data must be collected, processed, and stored.
The GDPR grants individuals eight data subject rights, legally enforceable entitlements to access, correct, delete, and control how their personal data is used.
The General Data Protection Regulation (GDPR) is the European Union's comprehensive data protection law, in force since May 2018, setting the global standard for how personal data must be handled by organizations operating in or targeting the EU.
Global Privacy Control (GPC) is a browser and device signal that automatically communicates a user's opt-out from the sale or sharing of their personal data, and is legally recognized as a valid opt-out mechanism under California and Colorado privacy law.
The Gramm-Leach-Bliley Act (GLBA) is a US federal law requiring financial institutions to explain their data sharing practices and safeguard the personal financial information of their customers.
The Health Insurance Portability and Accountability Act (HIPAA) is a US federal law establishing national standards for the protection of Protected Health Information (PHI), individually identifiable health data held by healthcare providers, insurers, and their business associates.
Identity resolution is the process of linking data records from multiple sources or systems to a single individual, creating a unified profile that accurately represents one customer across channels, devices, and touchpoints.
The Lei Geral de Proteção de Dados (LGPD) is Brazil's comprehensive data protection law, in effect since 2020, modeled closely on the GDPR and applying to any organization that processes personal data of individuals located in Brazil.
A large language model (LLM) is a type of AI model trained on massive text datasets that can generate, summarize, translate, and reason about language at human-comparable levels.
Master data management (MDM) is the practice of creating and maintaining a single, authoritative record of an organization's core data entities, including customers, products, suppliers, and employees, to ensure consistency across all systems that use that data.
Personally identifiable information (PII) is any data that can be used to identify a specific individual, either directly or in combination with other information.
Preference management is the practice of collecting, storing, and operationalizing individual customer preferences about how they want to be contacted, what data they're willing to share, and what kinds of personalization they'll accept, and ensuring those preferences are respected across every system and channel.
Pseudonymization is the process of replacing personally identifiable fields within a data record with one or more artificial identifiers or pseudonyms.
A Record of Processing Activities (ROPA) is the formal documentation required by GDPR Article 30 that catalogues every way an organization processes personal data, including the purpose, legal basis, data categories, recipients, and retention periods.
Responsible AI is the practice of designing, deploying, and operating artificial intelligence systems in ways that are fair, transparent, accountable, and aligned with legal and ethical standards.
A Retail media network (RMN) is an advertising platform operated by a retailer that allows brands to advertise directly to shoppers using the retailer's first-party purchase and behavioral data, on the retailer's own digital properties and increasingly off-platform.
Retrieval-augmented generation (RAG) is an AI architecture that combines a large language model with a retrieval system, allowing the model to search external knowledge bases and incorporate current or proprietary information into its responses.
Schrems II refers to a 2020 EU Court of Justice ruling that invalidated the EU-US Privacy Shield framework for cross-border data transfers, requiring organizations to rely on alternative mechanisms, principally Standard Contractual Clauses, and to assess whether the destination country offers adequate data protection.
Third-party data is information collected by an entity that has no direct relationship with the individual, typically a data broker, market research firm, or advertising network, and sold or licensed to other organizations.
Tracking cookies are small data files placed on a user's device that enable websites, advertisers, and analytics providers to monitor user behavior across websites and sessions, often without the user's active awareness.
Training data is the dataset used to teach an AI model, the examples from which it learns patterns, relationships, and behaviors.
The Utah Consumer Privacy Act (UCPA) is a state privacy law effective December 31, 2023, granting Utah residents rights over their personal data and applying to businesses that meet specific revenue and data processing thresholds
A verifiable consumer request (VCR) is a privacy rights request submitted under the CCPA or CPRA that a business has taken reasonable steps to confirm was actually made by the consumer to whom the data belongs, or by their authorized agent.
The Virginia Consumer Data Protection Act (VCDPA) is a state privacy law effective January 1, 2023, that grants Virginia residents rights over their personal data and establishes compliance obligations for businesses that process it at scale.