Orchestrating the Future of AdTech and MarTech
Building the Data and AI Foundation for Agentic Advertising, Real-Time Decisioning, and Privacy-Safe Measurement
Introduction
AdTech and MarTech vendors enter this cycle with a growing top line and a squeezed cost structure. WPP Media's midyear 2026 forecast projects global ad revenue reaching $1.3 trillion in 2026, up 8.9% excluding US political spend, with AI investment named as the primary growth driver. Growth is available. Margin is not, and every incremental dollar now depends on data and AI infrastructure that most platforms have not yet built.
Across the AdTech and MarTech ecosystem, DSPs, SSPs, ad servers, identity and measurement providers, CDPs, and customer engagement platforms are converging on five data and AI investment priorities:
- Agentic AI and campaign automation
- Data platform consolidation and cost control
- Real-time decisioning and activation
- Identity, consent, and privacy-safe measurement
- Retail media and clean-room interoperability
Each priority rests on the same foundation: clean, timely, governed data moving through reliable, observable pipelines that span bid streams, impression logs, customer profiles, clean rooms, and consent systems. When that foundation is fragile, models train on stale data, decisioning falls back to defaults, billing misreports impressions, and consent handling becomes a compliance exposure.
This guide defines what each priority requires and shows how Apache Airflow® and Astro make them executable.
Apache Airflow has grown to become the industry's most widely used system for orchestrating data workflows, as well as being one of the world's most active open source projects. Astro, Astronomer's unified orchestration platform, elevates Airflow into an enterprise-grade control plane purpose-built for high-scale AI and data-driven workflows.WHY AIRFLOW AND ASTRO?
INITIATIVE ONE Agentic AI and campaign automation
Agentic AI has moved from roadmap item to buying criterion. In the IAB 2026 Outlook Study of more than 200 US brands and agency buyers, five of the six top areas of increased advertiser focus tie directly to AI, and two-thirds of buyers are now focused on agentic AI for ad buying and campaign execution.
Supply is not keeping pace, and the blocker is trust rather than model quality. IAB UK's State of AI in Advertising found 74% of members at least experimenting with agentic AI, split between 58% piloting and 16% scaling, while only 4% describe themselves as fully agent-first. Among those members, 67% say they do not trust AI agents in advertising because decision-making lacks transparency. The gap is not model capability. It is the inability to show how an agent reached a decision.
Priority use cases
- LLM-driven creative generation and dynamic creative optimization that produce and score thousands of variants per campaign, then feed winners back into delivery.
- Agentic campaign planning and budget reallocation that forecast outcomes, shift spend across channels, and adjust pacing and frequency without a human in the loop.
- Audience and lookalike modeling agents that build, expand, and validate segments from first-party and commerce signals.
Why this is hard today
Agents need clean, current, well-described data and most platforms cannot supply it. Campaign, creative, identity, and outcome data sit in separate systems on different grains and refresh cadences, so agents run on stale inputs or in isolation.
The consequences hit the P&L and trust. An agent acting on a bad signal spends real advertiser money. A creative pipeline that cannot explain why a variant was served creates disclosure exposure under EU AI transparency rules. Pilots that cannot be governed never ship.
From pilot to product: the pipeline layer that makes it work
| Required capability | How Astro helps |
|---|---|
| Orchestrate multi-step LLM and agent workflows in production | The Airflow Common AI Provider runs end-to-end LLM and agent pipelines with branching logic, tool calls, and production-grade retries, so campaign-planning and creative-generation agents behave predictably under load. |
| Keep advertiser PII, customer profiles, and proprietary models in your environment | Remote Execution separates orchestration from execution, so PII, first-party profiles, proprietary bidding and propensity models, and code never leave your VPC. |
| Run real-time, parallel inference on campaign and user events | Event-driven scheduling and parallel task execution trigger inference on impression, conversion, and in-product events; autoscaling absorbs campaign-launch and seasonal spikes. |
| Trace a bad model output back to its source data | Astro Observe and native pipeline tasks link data quality checks, anomalies, and SLA breaches to AI pipelines with end-to-end lineage, so a mis-targeted segment or bad creative score traces to the task that produced it. |
| Diagnose AI pipeline failures in minutes, not hours | Otto, the data engineering agent for Astro, pulls the logs, analyzes the failure, and proposes a fix, reaching root cause without manually digging through code and logs. |
| Recover long-running agent and fine-tuning jobs without restarting from scratch | The task state store persists external job IDs and checkpoints across task retries, so a failed agent step or model fine-tuning run reconnects to its still-active job instead of resubmitting a duplicate and burning expensive GPUs. |
| Ship and roll back AI changes safely | Astro Runtime, the Astro IDE, and CI/CD provide a hardened Airflow distribution, browser-based Dag development with AI-pair programming, and Git-driven deployment to ship AI changes quickly with rollbacks. |
| Integrate any model, harness, or inference platform without re-architecting | Building on Apache Airflow, teams plug in any model provider, training platform, or serving stack, ensuring long-term flexibility as agentic patterns and vendors change. |
Airflow and Astro in Action
Airflow and Astro is already used by some of the most demanding AI companies and agentic workloads on the planet:
- OpenAI has standardized on Airflow across its business with over 7,000 data pipelines spanning research, operations, and finance, all while providing a foundation for 10x growth. Read more.
- GitHub Copilot relies on Airflow to process billions of developer events per day, orchestrating the feedback loops used to continuously improve the company's assistant and agents. Read more.
The AdTech and MarTech industry is following suit. For example, a leading North American contextual ad tech platform ran its AI, MLOps, and analytics workloads across four fragmented systems, so shipping a new data product took days and SLA coverage depended on a handful of senior engineers. The team standardized on Astro and added Astro Observe for unified observability. The results: new AI and ML data products roll out in 10 minutes instead of days, faster root cause analysis across model pipelines, and the retirement of a separate data catalog and homegrown SLA tooling.
INITIATIVE TWO Data platform consolidation and cost control
Infrastructure cost is now a strategic constraint on AdTech and MarTech margin. Log-level processing, egress, and GPU spend scale with volume while pricing pressure holds take rates flat. Industry forecasts show cloud infrastructure spend outpacing ad revenue growth by 3x. Consolidation is how platforms keep cost of goods below revenue growth.
Engineering capacity is the second constraint. In the State of Analytics Engineering report, 57% of data professionals said they spend most of their time maintaining or organizing existing datasets rather than building new capability, a figure flat year over year despite widespread AI-assisted coding. Every hour spent nursing brittle pipelines is an hour not spent on the product roadmap.
Priority use cases
- Warehouse-native and composable customer data architectures that activate profiles in place instead of copying them into a separate CDP store.
- Migration off cron sprawl, in-house schedulers, and self-managed Airflow onto a single governed control plane with uptime commitments.
- Consolidation of DSP logs, impression and bid data, profiles, clean-room outputs, and event streams into one pipeline layer feeding lakes, warehouses, models, and BI.
- dbt-based transformation at scale for audience, reporting, and billing models with task-level visibility.
- Reliable metering, aggregation, and rating feeding platform billing, take-rate revenue, and partner reconciliation.
Why this is hard today
Most platforms grew by acquisition and product line, so orchestration fragmented alongside. Separate teams run separate schedulers against overlapping copies of the same impression and profile data. Nobody owns the total picture, cost attribution is impossible, and idle workers burn spend around the clock.
A failed pipeline here means misreported impressions, missed billing runs, make-good exposure, and delayed optimization. Fragmented orchestration also blocks the consolidation story investors expect, because infrastructure cost cannot be shown to bend.
From fragmented to consolidated: the pipeline layer that makes it work
| Required capability | How Astro helps |
|---|---|
| Accelerate, derisk, and consolidate migrations from legacy schedulers | Otto converts legacy scheduler definitions into production-ready Dags, mapping job dependencies as it goes and producing deterministic output 10x faster than mechanical translation. Astro then consolidates those workflows, alongside scattered cron and Airflow instances, into a single managed control plane backed by uptime SLAs. |
| Bridge legacy ad servers, DMPs, and modern cloud platforms | Astro connects to legacy systems via JDBC/ODBC or custom hooks and orchestrates phased migrations with synchronized ETL, enabling stepwise modernization without big-bang cutover risk. |
| One pipeline layer across logs, profiles, warehouses, and streams | With 2,100+ connectors and flexible orchestration, Astro integrates siloed systems into a single pipeline layer feeding lakes, warehouses, models, and BI tools without redundant copies. |
| Cost-aware, scalable execution | Autoscaling and high availability scale workers up at peak and down when idle, delivering 2x faster execution versus self-managed Airflow while cutting infrastructure waste. |
| Cost visibility across data and AI workloads | Astro links pipeline execution to compute usage, so platform teams see which workloads drive cost spikes and optimize the ones that matter. |
| Unify orchestration and transformation | Orchestrate, run, and observe dbt workflows with Cosmos, the open-source standard for dbt orchestration and task-level visibility in Apache Airflow. |
| Reliable metering for usage-based and take-rate billing | Airflow Dags on Astro run metering, aggregation, and rating pipelines that feed billing systems with accurate, timely consumption data. |
| Multi-tenant, governed environments for platform teams | Workspace isolation and RBAC let a central platform team offer shared, governed orchestration to product, data science, and analytics teams with clear boundaries. |
Astro in Action
Foursquare processes billions of records daily, including GPS pings and ad impressions, feeding both internal workflows and customer-facing location products. That firehose ran through a fragmented mix of self-hosted Airflow, Luigi, and homegrown systems across roughly 50 engineers, with no central view of which datasets existed or what depended on them. Foursquare consolidated onto Astro as a single control plane. The results: 9,300+ data assets orchestrated under one platform, 5x faster pipeline development, and a 90% reduction in data discovery and access time, cutting days to minutes. Read more in the case study.
System1, a high-scale digital advertising and data business, ran Airflow across multiple teams and tools, with point solutions layered on top and FTE hours going into managing the platform rather than building on it. Public-company reporting put mission-critical pipelines on tight SLAs. System1 replaced self-managed open source Airflow with Astro and Astro Observe. The results: 10B+ rows per day orchestrated across 40+ web properties, point solutions retired in favor of unified orchestration and observability, and issues diagnosed before they escalate into SLA breaches.
INITIATIVE THREE Real-time decisioning and activation
Decisioning speed converts directly into revenue. On the AdTech side, the bidder answers inside a 50 to 100 millisecond window. On the MarTech side, activation latency shows up as abandoned journeys and unconverted intent. The commercial upside is measurable. In Twilio's State of Customer Engagement Report, 75% of companies said personalization increases customer spending, with an average 32% lift per purchase, while 84% of businesses believed they delivered good personalization against only 54% of consumers who agreed.
Where orchestration fits
The bid engine and the message engine make the decision in milliseconds, reading from in-memory stores and loaded model artifacts. Orchestration sits upstream from that process. It computes the features those stores serve, trains and deploys the models the engine loads, builds the audiences and suppression lists pushed to destinations, and aggregates the outcomes that inform pacing. A decision is only as good as the state behind it, and that state is a data product produced by an orchestrated data pipeline.
From stale state to fresh signal: the pipeline layer that makes it work
| Required capability | How Astro helps |
|---|---|
| Keep the features decisioning systems read current | Event-driven scheduling triggers feature pipelines the moment source data arrives, without polling or batch windows, so the stores the serving layer reads reflect recent behavior. |
| Train, evaluate, and deploy the models the engine loads | Astro orchestrates the full model lifecycle for propensity, churn, and valuation scoring, with scheduled and event-triggered retraining and Git-driven deployment. |
| Build and push audiences and suppression lists continuously | Orchestrated workflows assemble segments and suppression state and sync them to activation destinations on a cadence the campaign can rely on. |
| Trigger lifecycle workflows on behavioral change | Event-driven scheduling starts onboarding, upsell, and retention workflows when usage thresholds or behavioral patterns shift. |
| Absorb volume spikes in the pipelines behind decisioning | Astro orchestrates distributed workflows such as Spark jobs and batch scoring, and scales automatically to handle event-stream volume without manual capacity planning. |
| Crash-safe recovery for long-running distributed jobs | The task state store checkpoints external job IDs and progress markers, so a retry reconnects to a running job or resumes from its last checkpoint instead of reprocessing the full dataset. |
| Trust the signals behind decisioning models | Astro Observe and native pipeline tasks enforce schema, volume, and freshness checks on the pipelines powering scores, segments, and pacing models. |
| Production-grade reliability for the pipelines decisioning depends on | Autoscaling, cross-region DR, and zero-downtime updates deliver a 99.9% uptime SLA, replacing the overhead of self-managing Airflow through peak traffic. |
Astro in Action
A North American AI-driven advertising technology platform trains and serves models for individual advertisers. One business unit already ran on Astro; the other still ran self-managed Airflow, a second orchestration stack and a blocker to the upgrade the company needed.
It consolidated both onto Astro, migrating customer-facing workloads with zero downtime. Astro workers now run the full ML lifecycle behind 1,000 production ML pipelines including per-advertiser model training, at 13.07M monthly tasks across 128 deployments.
INITIATIVE FOUR Identity, consent, and privacy-safe measurement
Addressability and measurement credibility are the two assets buyers pay for, and both are eroding. Google's reversal on third-party cookies did not restore addressability; identity fragmented across hashed emails, alternative IDs, and platform graphs, and match rates decay at every hop.
Budget is moving to methods that survive without user-level tracking. eMarketer reports that media mix modeling now tops the incrementality measurement stack for retail brands, with 61% of US decision-makers using it, because in-platform attribution is no longer trusted. Vendors that prove incrementality with auditable data will hold pricing power.
Why this is hard today
Consent and deletion are treated as checklists outside the data flow, so a suppression flag set in one system does not reliably reach the pipeline that builds the audience. Identity pipelines span multiple providers with no single lineage, which makes match-rate decay impossible to attribute and measurement impossible to defend. The exposure is regulatory and commercial at once. GDPR fines reach up to 4% of global turnover and the EU AI Act up to 7%. When an advertiser asks how an attributed conversion was calculated and the answer takes three weeks, the contract is already at risk.
From exposure to evidence: the pipeline layer that makes it work
| Required capability | How Astro helps |
|---|---|
| Enforce consent, deletion, and residency inside execution | Astro pipelines enforce consent flags, deletion requests, and regional data rules as part of execution, not as an external checklist, so identity resolution and audience builds respect them by default. |
| Comprehensive data lineage and cataloging | Astro logs every task execution and data movement, providing a traceable path from source to output that supports audit readiness and simplifies impact analysis when rules or providers change. |
| Policy-as-code governance | Pipelines are defined in code and deployed via CI/CD, so teams embed masking, validation, and logging as enforced steps, codifying governance directly into identity and measurement operations. |
| Strong access control and identity management | Astro enforces RBAC, integrates with enterprise SSO and IAM, and supports isolated environments so access to identity graphs and consent signals is tightly scoped and auditable. |
| Automated compliance monitoring | Centralized metadata and usage dashboards surface failures, SLA breaches, and anomalies in the pipelines that feed regulated identity, consent, and measurement processes. |
| Orchestration-aware data quality on identity and measurement pipelines | Data quality checks such as row volume, null percentage, and custom SQL run as native pipeline tasks, and failures trace to the exact task that produced them and trigger standard alerting. |
| Keep raw identifiers and identity graphs in your environment | Remote Execution separates orchestration from execution, so raw identifiers, consent records, and match keys never leave your VPC or region. |
| A hardened, current runtime with expert support | Astro Runtime delivers a production-hardened Airflow distribution with timely security patches, Otto plans upgrades across your Dag fleet, and 24x7 support from the engineers who build Airflow backs mission-critical pipelines. |
Astro in Action
Kleinanzeigen, Germany's largest classifieds marketplace, rebuilt its data platform to power advertising insights for 30 million monthly users under EU data rules. The team used Astro to embed governance in the platform itself: standardized schema versioning, permission patterns across catalogs and secrets, CI/CD-enforced code review, and monitoring on by default. Guardrails are executed, not documented.
The results: 600+ production tables and 2 PB orchestrated, migrated in under six months, with domain teams contributing independently inside enforced standards. Read more in our case study.
INITIATIVE FIVE Retail media and clean-room interoperability
Retail and commerce media is now the third structural pillar of advertising alongside search and social, and for AdTech and MarTech vendors it is a direct revenue line rather than a demand-side trend. WPP Media reported commerce media reaching $178.2 billion globally in 2025. Vendors that supply the technology, measurement, and data plumbing capture margin on every network launched.
The operating reality lags the market size. eMarketer research on clean rooms finds fewer than half of US retail media networks offer clean-room capability today, and buyers consistently cite the absence of standards as a barrier to shifting more budget. The gap between what networks sell and what they can prove is the opportunity.
Why this is hard today
Clean rooms do not interoperate. Each retailer runs its own on different platforms and taxonomies, so data is copied between environments, eroding match rates at every hop. Purchase, campaign, and audience data arrive on different grains and have to be reconciled before anything can be measured.
The result is a channel selling faster than it can substantiate. Networks bill against numbers they cannot fully trace, brands discount the results, and measurement disputes stall renewals.
From network launch to proof: the pipeline layer that makes it work
| Required capability | How Astro helps |
|---|---|
| Event-driven data services and API integration for multi-party collaboration | Astro supports event-driven orchestration and native API integration, enabling the real-time data services and microservices patterns that retailer, brand, and clean-room collaboration require. |
| Integrate first-party purchase, SKU, campaign, and clean-room data | With 2,100+ connectors and flexible orchestration, Astro pulls retailer and brand systems into a single pipeline layer feeding measurement, activation, and reporting. |
| Auditable lineage from ad exposure to purchase | Astro logs every task execution and data movement, giving networks a traceable path from exposure to transaction that stands up to brand and auditor scrutiny. |
| Consent and residency enforced inside collaborative workflows | Astro pipelines enforce consent flags, deletion requests, and regional data rules as part of execution, so cross-party collaboration respects GDPR and clean-room constraints by design. |
| Reliable metering, aggregation, and rating for take-rate revenue | Airflow Dags on Astro run metering, aggregation, and rating pipelines that feed billing systems with accurate, timely consumption data. |
| Advanced analytics at commerce scale | Astro orchestrates complex distributed workflows such as Spark jobs, SKU-level aggregation, and model scoring, and scales automatically to handle large datasets without bottlenecks. |
| Let commercial and analytics teams build without breaking core data | Blueprint lets analysts and data scientists build and test governed Dags via a drag-and-drop canvas that fits existing CI/CD and deployment processes, with rollbacks keeping experiments off production-critical pipelines. |
Conclusion Next steps
From agentic campaign automation to clean-room measurement, each of the investment priorities profiled in this guide shares the same foundational requirements:
- Clean, timely, governed data across bid streams, impression logs, customer profiles, consent systems, and clean rooms.
- Reliable, observable pipelines with lineage that survives an audit and a buyer asking how a number was produced.
- Scalability and cost efficiency that absorbs bid-stream and campaign spikes while holding infrastructure cost below revenue growth.
That is the role of orchestration. The AdTech and MarTech platforms that win the next three years will treat orchestration as the control plane for AI, decisioning, identity, and commerce media revenue, and they will operationalize it with Astro.
Build a trusted, future-ready data stack today
Run an Astro TCO analysis and get in touch with our experts today to get results faster.
GET THE FULL GUIDE
Enter your work email to keep reading.
By proceeding you agree to our Privacy Policy, our Website Terms and to receive emails from Astronomer.