Evaluating Synthetic Consumers with Semantic Similarity Rating

Published on September 4, 2026

I’ve always believed that the best way to prove you understand a concept is to build it yourself.

I was introduced to a fascinating research paper by PyMC Labs about Semantic Similarity Ratings (SSR). The paper introduces a method for evaluating synthetic consumers generated by LLMs, particularly in the realm of consumer research. Both this post and the accompanying demo are heavily based on their work:

B. F. Maier, U. Aslak, L. Fiaschi, N. Rismal, K. Fletcher, C. C. Luhmann, R. Dow, K. Pappas, and T. V. Wiecki. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, 2025. https://arxiv.org/abs/2510.08338.

Reading the theory is one thing, but rebuilding it from scratch taught me very different things. I only trusted that I really understood it once I’d built a small, local SSR model myself to test it out.

You can find the demo I created for it here: ollama-ssr-pipeline.

Why We Need SSR

In a typical scenario, a corporation might send out a survey (say, via WhatsApp) to calculate the mean purchase intent (PI) for a new product. We collect the human ratings to build a probability mass function (PMF), which gives us a solid, empirical number.

But what happens when we use LLMs to simulate these consumers?

Instead of asking the LLM for a direct, arbitrary number (which can be wildly inconsistent), we use the SSR method. We input the synthetic customer’s demographics and the product description into the LLM. The expected output isn’t a score; it’s a rating text.

The Magic of Cosine Similarity

The question then becomes: how do we turn that text back into a reliable Likert rating score?

This is where the magic happens. We take the LLM’s output text and a set of predefined reference statements (like “I would definitely buy this” or “I would never buy this”). We pass both through a text-embedding model (in the paper’s case OpenAI’s text-embedding-3-small) to convert them into embedded vectors.

By calculating the cosine similarity between the generated text and our reference anchors, we can map the LLM’s raw thoughts onto a valid PMF vector. We aren’t just trusting a random number; we are mathematically measuring how closely the LLM’s simulated sentiment aligns with specific anchor points.

Measuring Success and Trustworthiness

To know if this actually works, the paper relies on two primary success metrics:

  1. Distributional Similarity: We calculate the Kolmogorov-Smirnov (KS) distance between the cumulative distributions of the human ratings and the synthetic ratings. The closer this score is to 1, the more accurately our synthetic model mirrors human distribution.
  2. Correlation Attainment: We compute the Pearson correlation between the purchase intents of the human data and our LLM data. To measure our model’s trustworthiness, we calculate the correlation attainment ratio (ρ)—estimating how close our empirical correlation is to the maximum theoretical correlation possible.

Optimising for Alignment

If our correlation attainment ratio isn’t close to 1, we don’t just give up. We optimise.

Instead of calculating the probability using standard linear indicators, we subtract the minimum similarity baseline across our reference statements. This adjustment corrects for the naturally narrow bandwidth of cosine similarities (where even wildly different sentences might still have scores like 0.81 and 0.85).

We also introduce a temperature parameter (T). Setting T < 1 sharpens the probability distribution to increase confidence peaks, while T > 1 smears it out to convey uncertainty. Finally, we normalise everything so our probabilities sum perfectly to 100%.

Closing Thoughts

Implementing this pipeline from scratch was incredibly rewarding. It gave me a deep, hands-on appreciation for how we can robustly evaluate synthetic data without relying on arbitrary numbers. The process of translating academic theory into a functioning local model proved that these complex methodologies can be made accessible and verifiable.

Until next time!