Skip to main content
New Introducing Mosaic Learn more
Conversion Labs
Labs Engineering

One-to-One: Evaluating AI Personalization in Marketing Systems

One-to-One: a benchmark for AI agents building personalized marketing content.

September 28, 2026 6 min read

Marketing systems let a team communicate with millions of people. Making those messages relevant means accounting for the differences between them: what someone cares about, where they are in a buying process, what their company needs, and what has already happened in the relationship. Personalization is how that context becomes part of the message.

We see this as an important frontier for AI in go-to-market work, especially marketing. An agent working inside these systems needs to be good at customization. It has to understand the available data, choose the right fields, and translate a marketer’s intent into behavior that works across many different contacts. Its ability to do that reliably is one practical measure of how useful it can be.

At Conversion, we care about this deeply because we want our agents to help customers create copy that is relevant to the people they contact. That requires evaluation at the point where customer data becomes content. A message that looks right for one fully populated contact may fall apart when a name is missing, a company has several opportunities, or a different event starts the workflow.

We built One-to-One to measure one part of that problem: how well can an AI generator turn a personalization request into a template that behaves correctly across different records? We started with 150 development tasks and eight models, then evaluated three models on a fresh set of requests with the generator held fixed.

From a request to a rendered message

Liquid is the template language behind personalization in Conversion. A template such as Hi {{ contact.first_name | default: "there" }}, produces a different greeting for each recipient. More involved templates read company fields, loop over opportunities, or use the event that started a workflow.

Here is one of the renewal templates generated during the benchmark:

{%- assign parts = "" -%}
{%- if company.company_name.size > 0 -%}
  {%- assign parts = company.company_name -%}
{%- endif -%}
{%- if company.tier.size > 0 -%}
  {%- if parts.size > 0 -%}
    {%- assign parts = parts | append: " · " -%}
  {%- endif -%}
  {%- assign parts = parts | append: company.tier -%}
{%- endif -%}
{%- if company.renewal_date -%}
  {%- capture renews -%}renews {{ company.renewal_date | date: "%b %-d" }}{%- endcapture -%}
  {%- if parts.size > 0 -%}
    {%- assign parts = parts | append: " · " -%}
  {%- endif -%}
  {%- assign parts = parts | append: renews -%}
{%- endif -%}
{%- if contact.first_name.size > 0 -%}
  {{ contact.first_name }}{% if parts.size > 0 %}: {{ parts }}{% endif %}
{%- else -%}
  {{ parts }}
{%- endif -%}

In plain English: write a renewal line using the person’s name, company, plan, and renewal date. Leave out anything missing, and keep the punctuation tidy.

For one test contact, that becomes Sarah: Acme Corp · Enterprise · renews Dec 15. For a contact with no first name or renewal date, it becomes Beta Labs · Starter. For someone with only a first name available, it is simply Miguel.

The template has to handle each combination of available fields, including when to insert a colon, add a separator, or format a date. You may notice that this looks a lot like a code generation problem.

The model gets the user’s request and the fields available in their workspace. It can look up additional information, write a template, and check its work. If a check fails, it gets the error back and can revise the template.

Comparing eight models

We started with 150 cases, from greetings with missing names to renewal messages, lists of opportunities, and nested JSON. We ran eight models through the same generator and used these cases while developing the feature.

Opus and Kimi had the highest pass rates in this comparison. The chart shows the original scores alongside the effect of the corrections identified in our scoring audit.

Development-set pass rates for eight models, showing original scoring and scores after targeted audit corrections.

Figure 1. Results from the development suite. Orange dots apply targeted scoring corrections; this was not a complete regrade. These scores are separate from the fresh-request results below. Full methods are in the research report.

Speed varied too. GLM Flash had the lowest median generation time at 2.57 seconds. Opus took 3.52 seconds, while Gemini Flash took 9.34 seconds. The slower end of the response-time distribution matters as well: a feature that usually feels quick can still leave users waiting on harder requests.

Median and 95th-percentile generation times for all eight models in the original development run.

Figure 2. Median and 95th-percentile generation times from eligible original attempts. Inference failures are excluded, and later replacement runs are not mixed into these timings.

Testing on fresh requests

Next, we held the generator fixed and wrote 36 new requests. We chose Opus and Kimi, the strongest models in the initial comparison, and GLM Flash, the fastest. All three used the same generator, tools, and workspace data.

Each model attempted every request three times. We rendered each template for 24 synthetic contacts, including people with missing names, missing companies, and multiple opportunities. A template passed only if it produced the requested output for every contact.

ModelTests passedMedian generation timeModel cost per 100 attempts
Claude Opus 597.2%3.49 s$1.61
Kimi K395.4%7.31 s$1.67
GLM 5.3 Flash87.0%2.75 s$0.12

Generation time includes the agent’s checks and revisions. Costs are gateway-reported charges from this run, which reused workspace context and benefited from caching. Provider connection failures are excluded from model quality; none occurred in these runs. Full methods and limitations are in the research report.

Opus passed the most tests, though this small evaluation doesn’t establish a clear advantage over Kimi. GLM Flash was faster and much cheaper, but its templates failed more often. The right tradeoff depends on how much an incorrect message costs you, as well as how much it costs to generate the template.

The agent’s revisions helped. Across all three models, 17 attempts started with an incorrect draft and finished with a correct template after feedback from the checker. Overall, the pass rate rose from 88.0% on first drafts to 93.2% on final answers.

A template can run and still get the message wrong

The generator accepted every final template. Yet 22 failed our output checks. All but one ran successfully, making the mistakes easy to miss if you only checked whether the template worked.

One request asked for a link containing tier=growth. An Opus template encoded the equals sign along with the text, producing tier%3Dgrowth. That looks like a small change, but the destination no longer receives the requested tier parameter.

Another request asked for a JSON array with separate customer and region tags. A Kimi template put both tags inside a single string. The JSON was valid. A system reading it would receive one tag instead of two.

These mistakes matter because the template will repeat them for every affected recipient. A preview that looks plausible is a useful start. Checking the actual output against the request tells us much more.

Making complex personalization dependable

Liquid lets a marketer express the details that make a message relevant: which opportunity to mention, which renewal date to use, what to say for a particular plan, and what to leave out when data is missing. Those rules live in a template that can be read, edited, tested, and reused across a campaign.

For us, Liquid is the right foundation for this kind of personalization. A marketer defines the behavior once, and the system applies it to each recipient’s data. A single campaign can account for differences that would otherwise require many manually maintained versions.

Natural-language generation makes that power available to people who would never write a nested Liquid template themselves. It also puts more responsibility on the agent. Every condition in the request has to survive the translation into code. A mistake in a fallback, a link, or a rule for selecting an opportunity can repeat across thousands of messages. As the personalization gets more ambitious, a plausible preview becomes less reassuring.

We want customers to be able to describe the message they have in mind, including the complicated parts, and trust our agent to carry that intent through to the people receiving it. One-to-One gives us a way to measure progress toward that goal. The standard is whether the generated template keeps working as the recipient’s circumstances change, across the many people a marketing team needs to reach.

Read the One-to-One paper for the full methodology and results.

Share this article