Simba documentation

Bayesian Modeling — The Statistical Foundation Behind Simba

In brief: Bayesian modeling gives you a range of likely outcomes (not just a single number), so you can make decisions with known uncertainty. In Simba, this runs automatically — no statistics knowledge needed.

Simba is built on Bayesian statistics. This page explains what that means, why it matters for marketing measurement, and how it works in practice — all without assuming a statistics background.


What Is Bayesian Inference?

Bayesian inference is a method of learning from data. It starts with a belief about how something works, observes new evidence, and then updates that belief to produce a more informed conclusion.

Here is the intuition: imagine you are evaluating a new paid social campaign. Before seeing any results, you have some expectation about how effective it will be — maybe based on past campaigns, industry benchmarks, or your own experience. That initial expectation is your prior belief. Once the campaign runs and you observe actual performance data, you combine that data with your prior belief to form an updated belief — your posterior.

This is exactly what Bayesian inference does, but with mathematical precision.


The Three-Step Workflow: Prior, Likelihood, Posterior

Every Bayesian model follows the same logical structure:

Step 1: Define the Prior

The prior represents what you believe about a parameter before seeing the data. In Simba, priors describe your expectations about things like:

Priors can be informative (expressing strong expectations based on domain knowledge) or weakly informative (expressing only broad constraints, like “this parameter should be positive”). Simba provides smart default priors that work well for most use cases, and you can adjust them when you have domain knowledge. See Priors and Distributions for a full guide.

Step 2: Define the Likelihood

The likelihood describes how the observed data relates to the model parameters. In an MMM context, the likelihood encodes the relationship between channel spend (after saturation and adstock transformations), control variables, seasonality, and the outcome variable. It answers the question: “Given a specific set of parameter values, how probable is the data we actually observed?”

You do not need to define the likelihood manually in Simba. The platform constructs it automatically from your data and model configuration.

Step 3: Compute the Posterior

The posterior is the result of combining the prior and the likelihood using Bayes’ theorem:

Posterior is proportional to Prior multiplied by Likelihood

In plain language: your updated belief equals your initial belief, adjusted by the evidence. If the data strongly supports a particular parameter value, the posterior will concentrate there regardless of the prior. If the data is weak or ambiguous, the posterior will stay closer to the prior.

The posterior is not a single number — it is a distribution. It tells you the full range of plausible values for each parameter, along with how probable each value is. This is the foundation of Bayesian uncertainty quantification.


Why Bayesian Beats Frequentist for Marketing Measurement

Traditional (frequentist) regression — the approach used in legacy MMM tools — produces point estimates and p-values. Bayesian modeling produces full probability distributions. Here is why that difference matters in marketing:

Uncertainty You Can Actually Use

A frequentist model might tell you that TV drives $2.1M in incremental revenue with a confidence interval of $1.5M to $2.7M. But what does that interval actually mean? In frequentist statistics, it means “if we repeated this analysis infinitely many times, 95% of the intervals would contain the true value.” That is a statement about a hypothetical procedure, not about your specific estimate.

A Bayesian credible interval is far more intuitive: “Given our data and assumptions, there is a 94% probability that TV’s true incremental contribution is between $1.5M and $2.7M.” This is the statement marketers actually want to make.

Incorporating What You Already Know

Frequentist regression treats every analysis as if you have zero prior knowledge. But in marketing, you almost always know something:

Bayesian modeling lets you encode this knowledge directly through priors, leading to more accurate and stable estimates — especially when data is limited.

Graceful Handling of Sparse Data

Marketing datasets are often short. You may have only 52 to 104 weekly observations, with some channels active for only part of that period. Frequentist regression can produce wildly unstable estimates in these conditions. Bayesian priors act as principled regularization, keeping estimates reasonable even when data is scarce.

No p-Value Theater

Frequentist analysis encourages a binary “significant or not” framing that is poorly suited to marketing decisions. A channel with a p-value of 0.06 is not meaningfully different from one with a p-value of 0.04, yet the significance threshold treats them completely differently. Bayesian analysis replaces this with continuous probabilities: “There is an 87% probability that this channel has a positive effect.” This supports nuanced, risk-aware decision-making.


How Simba Uses PyMC for Bayesian Computation

Simba’s statistical engine is powered by PyMC-Marketing, an open-source probabilistic programming library maintained by the PyMC development team. PyMC-Marketing is purpose-built for marketing science applications, providing pre-built components for saturation functions, adstock transformations, and time-varying effects.

Under the hood, PyMC uses advanced sampling algorithms — specifically Markov Chain Monte Carlo (MCMC) methods (algorithms that draw thousands of random samples to map out the posterior distribution) — to explore the posterior distribution. The core sampler is the No-U-Turn Sampler (NUTS) (an efficient MCMC algorithm that automatically tunes its step size), a state-of-the-art variant of Hamiltonian Monte Carlo that efficiently handles high-dimensional parameter spaces.

You do not need to understand MCMC to use Simba. The platform manages sampling configuration, convergence diagnostics, and chain validation automatically. But if you are a data scientist who wants to go deeper, Simba exposes convergence metrics (R-hat, effective sample size, trace plots) for full diagnostic transparency.


Uncertainty Quantification and Credible Intervals

One of the most powerful features of Bayesian modeling is uncertainty quantification — the ability to assign probabilities to ranges of outcomes.

What Is a Credible Interval?

A credible interval (sometimes called a Bayesian confidence interval) is a range of values that contains the true parameter with a stated probability. For example:

Simba uses 94% credible intervals by default, which is a common convention in Bayesian analysis (chosen for technical reasons related to the tails of the posterior distribution).

Why Uncertainty Matters for Decisions

Imagine two channels:

A naive analysis would favor Channel B because its point estimate is higher. But a risk-aware marketer would recognize that Channel B’s wide interval means its true ROAS could easily be below Channel A’s. Bayesian analysis gives you the information you need to make this distinction.

Propagating Uncertainty Through Predictions

When Simba forecasts outcomes or optimizes budgets, it does not plug in point estimates. It samples from the full posterior distribution and propagates uncertainty through every calculation. The result is a predictive distribution — a range of likely outcomes — rather than a single number. This means budget recommendations come with honest assessments of upside and downside risk.


Benefits of Bayesian Modeling in Simba

Incorporate Domain Knowledge

Encode industry benchmarks and expert intuition as informative priors. Lift test results are integrated separately as likelihood observations in the Model Details step. The model treats this knowledge as evidence, weighting it against the observed data.

Handle Sparse and Noisy Data

Short time series, channels with limited history, and noisy signals are all handled more robustly than with frequentist regression.

Full Posterior Access

Every parameter in the model has a full posterior distribution, not just a point estimate. This enables richer analysis, scenario planning, and risk assessment.

Principled Model Comparison

Bayesian methods offer natural tools for comparing model specifications (e.g., with or without a channel, different saturation functions) without relying on fragile hypothesis tests.

Fully Transparent

Simba’s UI lets you inspect every prior, posterior, and diagnostic. You can see exactly what the model assumed, how the data updated those assumptions, and how confident the final estimates are.


Key Takeaways


See this in action: Start your free 28-day trial — no credit card required.


Next Steps