September 29, 2026
An AI agent I was using to inform marketing allocation identified a major real-estate development as a candidate for increased advertising. The recommendation looked persuasive: recent sales value had risen sharply and the latest month appeared to show renewed buyer momentum.
Then I challenged the recommendation. Transaction activity across the last three months had moved from 23 to 9 to 12. Quarter-on-quarter, residential sales value was down 33.8% and transaction volume was down 14.3%. The apparent monthly improvement could have reflected a concentrated release or registration event rather than broad-based demand.
Nothing the agent had reported was fabricated. It had simply produced a defensible interpretation from one analytical frame. Had I accepted it at face value, I could have acted on a recommendation grounded in real data and still insufficiently tested for the decision being proposed. That is a harder governance problem than an obviously wrong AI answer.
Correct data is not enough
Much of the current conversation around AI governance focuses, reasonably, on familiar risks: poor data quality, hallucination, bias, explainability and human accountability. Those controls matter. (For more on this, see the CMO Council-BPI Network report The Pathway to GenAI Competitive Advantage.)
But they do not fully address what happens when an autonomous analytical system reaches a plausible recommendation from legitimate data through a series of individually defensible analytical choices.
An agent analyzing a large dataset may decide which population matters, which period to examine, which comparison base to use, where to set thresholds, how to aggregate observations and which contradictory findings deserve further investigation. None of those decisions needs to be objectively "incorrect." Yet collectively they can determine whether the final recommendation looks compelling, weak or even points in the opposite direction.
That creates a different problem for CMOs increasingly relying on AI not only to generate content but to interpret performance, identify opportunities, recommend resource allocation and support commercial decisions. The executive may receive the conclusion without seeing enough of the analytical path to judge how much confidence that conclusion deserves.
I call this the analytical approval gap: the decision-maker remains accountable for the action while much of the substantive analytical process that shaped it has become invisible.
Bias can enter before the data is analyzed
One of the least obvious sources of analytical bias is the instruction itself. Prompt bias does not require telling a system what conclusion you want, it can be much subtler than that.
I encountered this while asking an agent to analyze approximately 150,000 rows of market transaction data to inform a market report. The instruction explicitly asked for a neutral analysis. But it also established that the output was being produced for a marketing purpose, and that contextual cue was enough to matter.
Faced with several equally plausible ways of assessing market direction, the agent favored a time horizon that showed substantially more positive movement than alternative periods. The selected analysis was not factually wrong, and the time horizon was not inherently unreasonable. But other defensible comparison periods produced a materially less positive interpretation.
Nothing in the prompt had explicitly asked the agent to make the market look stronger. The direction entered more subtly: the commercial purpose of the task appeared to influence which analytical view the system treated as most useful.
That is precisely why prompt bias is difficult to govern. A request to "find growth opportunities" obviously privileges positive movement. But even apparently neutral instructions can carry contextual signals about the intended use of the analysis, and those signals can influence which evidence is surfaced, which time horizon receives emphasis and how uncertainty is resolved.
Commercial analysis will always have a purpose. The answer is not to remove context from the prompt. It is to recognize that the prompt itself is part of the analytical methodology. That means analytical provenance should preserve not only the data and calculations behind a recommendation, but also the original instruction that framed the task. If a materially different but equally reasonable framing could have produced a different conclusion, the decision-maker should be able to see that.
From data provenance to analytical provenance
Marketing organizations are becoming more sophisticated about data provenance: where information came from, who owns it, when it was collected and whether it can be trusted. AI decision support requires another layer.
Analytical provenance is the record of the substantive choices that sit between the data and the recommendation. At minimum, for consequential analysis, that means preserving:
This is different from keeping technical logs. Token usage, tool calls, model versions and execution histories may help teams operate the system. But they do not necessarily tell a CMO why or whether the system selected one population over another, privileged one time horizon or comparison base, excluded a particular subset, applied one threshold rather than another, or abandoned one analytical path in favor of a different one.
Those are the choices that can materially shape the recommendation while remaining largely invisible in the final output. The useful record is the one that allows the organization to answer three questions later: Why did the system reach this conclusion? What would make that conclusion weaker? What does the decision-maker need to know before acting on it?
But preserving that record creates a second governance problem: who is realistically going to inspect it?
Executives cannot audit every analytical decision
A CMO cannot inspect hundreds of queries, alternative specifications, calculations and intermediate analytical branches every time an AI system produces a recommendation. If responsible AI governance requires executives to become forensic analysts, it will fail.
The full provenance record therefore needs to exist in the background, while the executive sees only what could materially change the decision. That executive layer should surface five things.
These are not questions about whether AI is "good" or "bad." They are questions about whether the recommendation is decision-ready.
Separate the analysis from the challenge
One way to operationalize this is to separate the analytical objective from the challenge objective. The first agent solves the business problem. A second agent — or another independent challenge mechanism, which I will call the challenge layer — interrogates the retained analytical provenance before the recommendation reaches the executive.
Its job is to approach the analysis with a different objective: not to strengthen or optimize the original recommendation, but to identify the assumptions, analytical dependencies, contradictory evidence and alternative interpretations that could materially weaken it. It might ask whether the original task was framed in a way that narrowed the answer, whether another reasonable comparison would materially change the result, whether contradictory evidence was sufficiently explored, or the conclusion depends heavily on a particular subset, threshold or analytical choice.
The challenge layer's own analytical choices are not immune to the dynamics described above. But its task is narrower and more falsifiable than the first agent's: showing that a specific alternative materially changes the result is a smaller, more checkable claim than synthesizing an open-ended recommendation, which is what makes it easier to spot-check, not what makes it infallible.
But identifying those weaknesses is only half the job. The challenge also has to be translated into something the decision-maker can use. The executive does not need the entire analytical trail. The translated output could simply read:
ACCEPT WITH QUALIFICATION. The evidence supports testing additional investment, but not the proposed 30% reallocation. The result is concentrated in two campaigns and is less pronounced over a longer period. Recommend a smaller controlled increase with a defined review point before further budget is moved.
That is the second function of the challenge layer: translation. It does not remove analytical complexity. It identifies which complexity is consequential and compresses it into decision-ready information.
Depending on the evidence and the consequence of the proposed action, the output might be ACCEPT, ACCEPT WITH QUALIFICATION, REVIEW or REJECT. The status should be accompanied only by the small number of analytical conditions that explain it.
The value of the challenge layer is therefore not that another model is inherently more intelligent than the first. It is that it has a different objective: challenge the recommendation, identify what materially changes its strength, and translate those findings into a form that makes human judgement practicable.
Governance before failure
AI governance often becomes visible after something has gone wrong. Analytical governance needs to operate earlier.
The most dangerous recommendation may not be the one containing an obvious hallucination. That will often be challenged immediately. It may be the recommendation that contains real numbers, answers the question it was given and arrives with enough apparent analytical sophistication that nobody thinks to ask what else could have been concluded from the same evidence.
As AI takes on more of the analytical work behind marketing decisions, provenance cannot end at the source of the data. CMOs also need visibility into the analytical choices that transformed that data into a recommendation.
That does not mean supervising every step. It means designing systems that preserve the analytical trail, challenge consequential conclusions and translate the material weaknesses into something a decision-maker can realistically use before somebody acts.
The recommendation is only the first output. The second gives the decision-maker the context needed to decide whether to act on it.
Nikolett Vilmos is an award-winning marketer and Head of Marketing & AI Transformation at Christie's International Real Estate UAE, where she leads branding, PR, digital strategy, lead generation, partnerships, events, and content initiatives for a brokerage recognized as Affiliate of the Year across Europe, the Middle East, and Asia-Pacific. Her work spans marketing strategy, brand, performance, analytics, and the design and deployment of AI-enabled systems across marketing and commercial operations.
Under her leadership, the brokerage expanded into Ras Al Khaimah, becoming the first international luxury brokerage to enter the emirate, where it has since transacted over AED 1.3 billion and been contracted to exclusively represent multiple branded developments. Nikolett's approach blends financial literacy with narrative-driven marketing, positioning luxury properties as both lifestyle purchases and long-term investment decisions, and her perspective on branded residences as an emerging asset class has been featured in international finance press.
Before Christie's, Nikolett founded her own advertising consultancy and held senior in-house and agency leadership roles, with responsibility for media budgets of up to £1.5 million per month at brands across Europe, the US, and the Middle East, including Solutions Leisure Group, LDX Digital, edyn, and Re_Set. She holds a Master's degree in Digital Marketing & Data Analytics.
No comments yet.