Skip to main contentSkip to navigationSkip to footer
    Data & Analytics

    Data Terms A-Z

    Understand the language of data: From Big Data to ETL to Predictive Analytics – all important terms for data-driven marketing and Business Intelligence.

    Big Data
    Data Lakes
    ETL Processes
    Business Intelligence
    Predictive Analytics
    Data Governance
    225 terms in Data & Analytics

    A

    Accuracy

    A metric in machine learning that measures the proportion of correct predictions made by a model out of all predictions made.

    Ads-to-GMV (A2G)

    Ads-to-GMV (A2G) quantifies advertising revenue as a percentage of the total Gross Merchandise Value (GMV) transacted on a retail platform. This metric is primarily utilized within the retail media sector to assess a platform's monetization strategy through advertising services. It aggregates all advertising earnings, such as those from Sponsored Products or Display Ads, and relates them to the overall sales volume.

    Agent Traffic Analytics

    Agent traffic analytics is the systematic measurement of automated, AI-driven access to a website: which assistants crawl, which content they fetch, which answers generate citations and which sessions actually result in an action.

    Agentic Analytics

    Agentic Analytics combines conversational analytics with proactive AI agents that detect anomalies, opportunities and risks, analyse root causes and deliver recommendations with impact forecasts.

    Agentic Attribution

    Agentic attribution measures the contribution of AI agents (own and third-party) to the purchase journey — from research agents to comparison and buying agents — and attributes revenue accordingly.

    AI Capex Hangover

    The AI capex hangover is the economic aftermath of outsized investment in AI infrastructure: high fixed costs and short depreciation cycles for accelerator hardware meet revenues that materialise later and at thinner margins than expected.

    AI Cost Attribution

    AI Cost Attribution assigns inference, training and ops cost precisely to teams, features, customers or campaigns — the basis for FinOps, prioritisation and ROI proof.

    AI-Powered CDP

    Customer Data Platforms with integrated AI/ML capabilities for automated segmentation, predictions, and activation.

    Analytics

    The systematic analysis of data to gain insights and support decision-making.

    Anomaly Detection

    Identification of unusual patterns or outliers in data.

    Apify

    Apify is a platform for running, scheduling and publishing web scraping and automation programs (Actors) that returns results as structured datasets and can also be connected as an MCP server.

    ARIMA (AutoRegressive Integrated Moving Average)

    A classic statistical model for time series forecasting that combines autoregression, differencing, and moving averages.

    AUC (Area Under the Curve)

    The area under the ROC curve – a single number (0-1) summarizing the overall quality of a binary classifier.

    C

    Causal Inference

    Causal inference is the discipline of estimating cause-and-effect relationships (what would happen if we changed X), not just correlations.

    Chain of Custody

    Chain of custody is the documented trail of how an artifact (data, evidence, content) was collected, handled, stored, and accessed—ensuring integrity and accountability.

    Changepoint Detection

    Detection of time points at which the statistical properties of a time series significantly change.

    Clean-Room AI

    Clean-Room AI combines data clean rooms (AWS, Snowflake, LiveRamp) with AI models to generate insights across brand and platform data — without raw PII ever leaving the room.

    Clickless Attribution

    Clickless attribution refers to measurement approaches that capture content's contribution to business outcomes even without a click to your site — via brand mentions in AI answers, direct visits, branded search queries and surveys.

    Clickstream Data

    A time-ordered record of user interactions (clicks, page views, events) across digital properties such as websites and apps.

    Cohen's Kappa

    A statistic for measuring inter-rater reliability for categorical ratings, corrected for chance agreement.

    Cohort Analysis

    Cohort analysis groups users or entities by a shared starting event/time (e.g., signup week) and tracks behavior over time.

    Compute per Conversion

    Compute per conversion is an efficiency metric attributing the total computational effort spent on a conversion path to a single conversion. It captures model inference, vector search, personalisation computation and evaluation runs — regardless of whether they occur in the frontend or in the background.

    Concept Drift

    Concept drift is the shift in meaning or distribution of concepts over time — e.g. what customers mean by "premium" or what qualifies as "spam".

    Confounding

    A confounder is a variable that influences both the independent and dependent variable, creating a spurious association.

    Confusion Matrix

    A table that summarizes classification performance by counting true positives, false positives, true negatives, and false negatives.

    Consent Orchestration

    Consent orchestration is the central management, distribution and enforcement of user consent across all channels, systems and AI models — including dynamic rules per region and purpose.

    Content Fingerprinting

    Content fingerprinting creates a compact signature (fingerprint) of content to enable identification, deduplication, similarity detection, or provenance tracking.

    Conversational Analytics

    Conversational Analytics lets business users ask complex data and marketing questions in natural language; an AI system translates to SQL, metrics and visual answers — with citations and lineage.

    Cosine Similarity

    A measure of similarity between two vectors that calculates the cosine of the angle between them, independent of their magnitude.

    Customer Data Platform (CDP)

    Central system for unifying customer data from all sources.

    D

    Dark LLM Traffic

    Dark LLM traffic refers to visits and conversions triggered by a brand being mentioned in AI assistants but recorded in analytics as direct, branded search or unknown source.

    Dashboard

    A visual interface that presents key metrics, trends, and alerts to support decision-making.

    Data Catalog

    A searchable inventory of an organization's data assets including metadata, ownership, and documentation.

    Data Clean Room

    A secure environment where multiple parties can combine their data for joint analyses without sharing raw data.

    Data Dictionary

    Documentation that defines the meaning, format, allowed values, and usage of data fields.

    Data Drift

    The change in statistical properties of input data over time, which can degrade model performance.

    Data Enrichment

    Adding additional attributes to existing data—via internal joins or external sources (firmographic providers, geo data).

    Data Governance

    The framework for policies, processes, and responsibilities to manage data assets in an organization.

    Data Labeling

    Process of annotating data with ground truth for supervised learning.

    Data Lake

    Central storage for large amounts of unstructured and structured data.

    Data Layout

    The physical or logical arrangement of data in memory or on storage media, which influences access speed, cache efficiency, and processing performance.

    Data Lineage

    Data lineage describes where data comes from, how it moves through systems, and how it is transformed into downstream datasets and outputs.

    Data Mesh

    Decentralized approach to data architecture with domain-oriented data products.

    Data Mining

    The process of discovering patterns, anomalies, and relationships in large datasets using statistical and machine learning methods.

    Data Pipeline

    A sequence of processes that moves and transforms data from sources to destinations (lake, warehouse, feature store, vector index).

    Data Preprocessing

    Transforming raw data into a form suitable for modeling or analysis (cleaning, normalization, encoding).

    Data Processing Agreement (DPA)

    A legally binding contract between data controller and data processor that governs the terms for processing personal data according to GDPR.

    Data Validation (ML)

    Automated checking of data quality, schema conformity, and statistical properties in ML pipelines.

    Data Visualization

    The graphical representation of data to communicate insights and patterns.

    Data Warehouse

    A system optimized for structured analytics queries over curated, cleaned data—often with strong governance and performance.

    Databricks

    Databricks is a unified analytics platform that combines data engineering, data science, and machine learning on Apache Spark.

    DBSCAN

    DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a clustering algorithm that finds clusters based on density of data points and automatically identifies outliers.

    Decision Support System (DSS)

    A Decision Support System (DSS) helps people make better decisions by combining data, models, and user interfaces.

    Decision Threshold

    The cutoff used to convert a model score/probability into an action (e.g., approve/deny, route/escalate).

    Deduplication

    Deduplication is identifying and removing duplicate (or near-duplicate) items to reduce redundancy and improve quality.

    Demand Forecasting

    Prediction of future demand based on historical data and factors.

    Difference-in-Differences (DiD)

    Quasi-experimental method that estimates causal effects by comparing changes over time between treatment and control groups.

    Differential Privacy

    A mathematically rigorous definition of privacy that guarantees an individual's participation in a dataset is statistically undetectable – even against attackers with arbitrary background knowledge.

    Differential Privacy for Advertising

    Differential privacy for advertising adds calibrated noise to reports, models or aggregates so individuals cannot be identified — at an acceptable accuracy trade-off.

    Dimensionality Reduction

    Techniques for reducing the number of features while preserving important information.

    Double Machine Learning (DML)

    Causal inference method that uses ML models to flexibly control for confounding while enabling valid statistical inference.

    F

    F1 Score

    The harmonic mean of precision and recall, a single metric that balances both aspects of classification performance.

    Feature Engineering

    The process of selecting, transforming, and creating input variables (features) for machine learning models to improve their predictive power.

    Feature Importance

    Feature importance quantifies how much each input feature contributes to a model's predictions (globally or for a specific prediction).

    Federated Marketing

    Federated Marketing uses federated learning and similar techniques to train models across distributed data sources (partners, devices, clean rooms) without moving raw data.

    FinOps for AI

    FinOps for AI applies financial operations practices (cost visibility, optimization, budgeting, accountability) to AI workloads and AI product usage.

    Firecrawl

    Firecrawl is an API that prepares web content for language models. It returns single pages, discovers a site's URLs, traverses them recursively, searches the web and parses documents such as PDF or DOCX — each as Markdown or structured data.

    First-Party Data

    Data collected directly from own customers and users.

    First-Party Data AI

    Strategic approach of using proprietary customer data as a differentiation layer on top of generic foundation models.

    First-Party Graph

    A first-party graph links all owned data sources of a brand (web, app, CRM, support, shop, IoT) into a unified person- or household-centric graph as the basis for AI and personalisation.

    Fraud Detection

    AI-powered detection of fraudulent activities and transactions.

    Fuzzy Matching

    Techniques for finding approximate rather than exact matches in data.

    M

    MAE (Mean Absolute Error)

    The average of absolute differences between prediction and reality – robust to outliers.

    MAP (Mean Average Precision)

    The average of Average Precision across all queries – considers both precision and ranking position of all relevant documents.

    Marketing AI Beyond CDP

    Marketing AI Beyond CDP is an advanced architectural approach for marketing technology where a traditional Customer Data Platform (CDP) is functionally augmented by integrating Digital Twins, custom-trained Small Language Models (SLMs), and autonomous agents. This transforms the CDP from a historical data collection into a proactive, predictive, and simulating system capable of mapping and optimizing complex customer interactions and market scenarios.

    Master Data Management (MDM)

    Master Data Management (MDM) is an approach to ensure critical enterprise data (e.g., customers, products, locations) is consistent, accurate, and governed across systems—often aiming for a "single source/version of truth."

    MinHash

    MinHash is a technique to efficiently estimate similarity between sets (especially Jaccard similarity), commonly used for near-duplicate detection.

    Minimum Detectable Effect (MDE)

    MDE is the smallest true effect size an experiment can reliably detect given traffic, variance, significance level, and power.

    Misattributed Citation

    Misattributed Citation refers to the issue where AI systems erroneously assign a recommendation, an information source, or a user conversion to a brand or channel that was not the actual origin. This distorts marketing attribution and consequently impacts the performance evaluation of campaigns and channels.

    MRR (Mean Reciprocal Rank)

    The average of the reciprocal ranks of the first relevant result across all queries – MRR = 1/n × Σ(1/rank_i).

    MSE (Mean Squared Error)

    The average of squared differences between predicted and actual values – standard loss for regression.

    N

    NaN (Not a Number)

    NaN is a special floating-point value meaning "Not a Number," used to represent undefined or unrepresentable numeric results (e.g., 0/0).

    Natural Experiment

    A natural experiment uses real-world events or operational changes (not randomized by you) that approximate random assignment, enabling causal inference under assumptions.

    NDCG (Normalized Discounted Cumulative Gain)

    A ranking metric that considers both relevance grades and positions in the ranking – higher-ranked relevant items are weighted more heavily.

    NDJSON (Newline-Delimited JSON)

    NDJSON is a format where each line is a valid JSON object—making it easy to stream, append, and process logs/events at scale.

    Negative Binomial Regression

    Negative binomial regression is a statistical model for count data (e.g., clicks, conversions) that handles overdispersion (variance > mean), unlike Poisson regression.

    Negative Control

    A negative control is a variable, outcome, or test condition that should not be affected by an intervention—used to detect bias, confounding, or measurement artifacts.

    Neon (Serverless Postgres)

    Neon is a Postgres platform launched in 2021 in which compute nodes and storage are decoupled. This allows database branches to be created in seconds, compute to scale automatically and idle instances to scale to zero.

    NHST (Null Hypothesis Significance Testing)

    NHST is the traditional statistical testing framework where you test whether observed data is unlikely under a null hypothesis (often "no effect"), typically using p-values.

    NMI (Normalized Mutual Information)

    NMI is a metric used to compare clustering assignments by measuring how much information one clustering shares with another, normalized to be scale-friendly.

    Noise-to-Signal Ratio

    Noise-to-signal ratio measures how much random variation (noise) exists relative to the meaningful pattern (signal) you want to detect.

    Non-Negative Matrix Factorization (NMF)

    NMF factorizes a non-negative matrix into two smaller non-negative matrices, often used for interpretable topic-like decompositions.

    Non-Production Data Masking

    Non-production data masking is the practice of anonymizing, tokenizing, or synthesizing sensitive data before it is used in dev/staging/test environments.

    Normal Form (Database)

    In databases, normal forms (1NF, 2NF, 3NF, BCNF) describe levels of normalization that reduce redundancy and improve data integrity.

    Normalized Cost per Answer

    Normalized cost per answer is the cost of generating an AI answer adjusted for comparability (e.g., normalized by answer length, tokens, difficulty tier, or traffic segment).

    Normalized RMSE (NRMSE)

    NRMSE is RMSE normalized by a scale factor (e.g., range, mean, or standard deviation) to make errors comparable across datasets.

    Nowcasting

    Forecasting the current or imminent state using high-frequency real-time data.

    Null Value

    A null value represents missing or unknown data (distinct from zero, empty string, or false).

    P

    p-Hacking

    Manipulating analysis choices (stopping rules, segmentation, metrics, exclusions) to obtain statistically significant results.

    p-Value

    The probability of observing results at least as extreme as what you observed if the null hypothesis were true.

    PII (Personally Identifiable Information)

    Information that can identify a person directly or indirectly (e.g., name, email, phone number, government IDs).

    Pinecone

    Pinecone is a commercial, fully managed service for storing and similarity-searching vectors (embeddings) that takes over index management, scaling and query optimisation.

    PostHog

    PostHog is a platform for product and web analytics with session replay, feature flags, experiments, error tracking, surveys and observability for AI applications. Model calls are captured with prompt, response, tokens, cost and duration, and linked to usage data.

    Power Analysis

    Calculation of the necessary sample size to detect an effect of a given size with desired probability (power).

    Precision

    The proportion of correctly classified positive cases out of all cases classified as positive.

    Precision and Recall

    Two complementary metrics for evaluating classification models on imbalanced data.

    Precision@k

    Measures how many of the top-k retrieved items are relevant (relevant items in top-k ÷ k).

    Privacy Budget

    A quantitative measure (epsilon, ε) of the total privacy loss accumulated through repeated queries on privacy-protected data.

    Prophet (Facebook/Meta)

    An open-source forecasting tool developed by Meta that automatically models trend, seasonality, and holiday effects.

    Provenance

    Provenance is metadata that describes the origin, history, and transformation path of data or content—where it came from, how it changed, and who/what changed it.

    Pseudonymization

    Replaces identifiers with pseudonyms so data can't be directly attributed to a person without additional information kept separately.

    S

    Sampling

    Sampling is selecting a subset of data (or outcomes) from a larger population/process to estimate properties, reduce cost, or enable exploration.

    Scenario Analysis

    Scenario analysis evaluates outcomes under a set of coherent, plausible future conditions (scenarios), rather than changing one variable at a time.

    Schema

    A Schema defines the structure, organization, and constraints of data – whether in databases, APIs, or structured data formats.

    Schema-on-Read

    Schema-on-Read is a data management approach where the structure of data is applied only at query time, not when storing.

    Scrapfly

    Scrapfly is a commercial service for fetching web pages via an API that provides proxy rotation, cloud browsers for JavaScript-heavy pages and measures against bot detection.

    Seasonality

    Regularly recurring patterns in time series that repeat at fixed intervals.

    Segment Analysis

    Segment analysis breaks metrics down by meaningful groups (segments) such as channel, device, region, customer tier, or intent.

    Sensitivity Analysis

    Sensitivity analysis evaluates how changes in inputs affect outputs, to understand robustness and key drivers.

    Sentiment Score

    Numerical value that quantifies the emotional polarity of a text.

    Session

    Period of user interaction with a website or app.

    Sessionization

    Sessionization groups user events into sessions to analyze behavior over time (page flows, search sequences, conversions).

    Shadow Eval

    Shadow eval is an evaluation method in which real production requests are additionally sent to a candidate version. Its outputs are logged and scored but never served, enabling quality comparison under real conditions without exposing users to risk.

    Silicon Sample

    A silicon sample is a set of model-generated persona responses that, based on demographic and psychographic specifications, aims to approximate the response distribution of a real sample.

    SimHash

    SimHash is a fingerprinting method that produces a compact hash where similar documents tend to have similar hashes (small Hamming distance).

    Simpson's Paradox

    Simpson's paradox is when a trend appears in multiple groups but reverses or disappears when the groups are combined, due to confounding and aggregation.

    Snorkel

    Snorkel is a framework for programmatic data labeling that uses labeling functions instead of manual annotation to efficiently create large training datasets.

    Snowflake

    Snowflake is a cloud-native data warehouse platform that separates storage and compute, enabling scalable data analysis with SQL.

    Specificity

    The proportion of correctly classified negative cases out of all actual negative cases.

    Stationarity

    A time series is stationary when its statistical properties remain constant over time.

    Statistical Significance

    Statistical significance describes the probability that an observed effect did not arise by chance — measured via the p-value against a defined threshold (usually 0.05).

    Streaming Data

    Continuous data flow that is processed in real-time.

    Survival Analysis

    Statistical method for analyzing time until an event occurs (e.g., churn, conversion, failure), accounting for censored data.

    Synthetic Customers

    Synthetic Customers are AI-simulated customer profiles, often referred to as Digital Twins, generated from extensive first-party data and panel information. Their primary function is to pre-test purchasing behavior and responses to marketing campaigns in a controlled virtual environment.

    Synthetic Data

    Artificially generated data that replicates statistical properties of real data – used for training, testing, and privacy protection when real data is scarce, sensitive, or expensive.

    Synthetic Marketing Data

    Synthetic marketing data are machine-generated, statistically realistic datasets (customers, purchases, interactions) that complement or replace real first-party data — without GDPR risk.

    V

    Validation Set

    A validation set is a held-out dataset used during model development to tune hyperparameters and select model versions without touching the final test set.

    Variance

    Variance is the degree to which a model's performance changes across different datasets/samples; high variance often indicates sensitivity to training data (overfitting risk).

    Vector Database

    A vector database stores embeddings and supports fast similarity search (nearest neighbors), often with metadata filtering and indexing for scale.

    Vector Embedding

    A vector embedding is a numerical representation (array of floats) of text, images, or other data that encodes semantic meaning in a high-dimensional space.

    Vector Index

    A vector index is the data structure/algorithm used to speed up nearest-neighbor search over embeddings at scale.

    Vector Quantization

    Vector quantization (VQ) compresses continuous vectors by mapping them to a finite set of representative vectors (a codebook).

    Vector Search

    Vector search retrieves items by similarity in an embedding space rather than exact keyword match.

    Vector Similarity

    Vector similarity is a measure of how close two embeddings are (commonly cosine similarity or dot product).

    Vector Store

    A vector store is the storage layer (database or service) that holds embeddings plus metadata for retrieval and similarity search.

    Vector Store Hygiene

    Vector store hygiene is the operational discipline of keeping a vector store accurate, secure, performant, and up-to-date (dedupe, versioning, ACL correctness, drift monitoring, purge workflows).

    Term not found?

    Browse the full glossary with over 2166 terms from all categories.

    View Full Glossary